
In scope: token pricing trends, inference gross margins, the AI-lab P&L question, and how memory, optics and foundry cost structure feed through to the price of a token.
Inference & Token-Pricing Economics
This rung is about what a token costs to make and what it sells for — token pricing trends, inference gross margins, the AI-lab P&L question, and how memory, energy and hardware cost structure feed through to the price of a token. Its sharpest observation is that the two numbers are only loosely attached to each other: Quanyan Zhu's tokenomics framework argues that token expenditure and economic value are different quantities, so pricing by tokens processed "prices the input, not the output" (AI Tokenomics, read at abstract level only, and explicitly theoretical — it carries no measured figures at all).
What this rung prices
A quoted price per million tokens is the start of the calculation and not the end of it, and this rung holds the three multipliers that sit between it and the price of a finished job. First, the seller's side: the same hardware serving the same model delivers wildly different cost per token depending on how busy it is, so a cost figure without a utilisation figure is not a cost figure. Second, the buyer's side: a task is only done when it is done correctly, so the unit that settles a purchase is cost per correct answer, not cost per token — a cheap model that is often wrong can cost more per finished job than an expensive one. Third, the decline: prices per token have fallen fast but very unevenly, and the fall lands on some tiers and some tasks and not others, so "inference is getting cheaper" is not a statement anyone can act on until it is asked about a particular tier. Get those three right and a token price converts into a task price; skip any one and the conversion is off by an order of magnitude or more.
The load-bearing findings
- Identical hardware, identical model, a 73× spread in cost — from load alone. Effective cost ran from $0.21 to $15.25 per million output tokens on the same H100 setup, driven purely by offered request rate; the underutilisation penalty is 2.5–24× across a normal 1–10 requests-per- second enterprise load and up to 36.3× near idle. The dollar level is anchored to a named basis — $6.98/GPU-hour, Azure H100 NVL on-demand list — and the analytic point is that a utilisation-naive estimate understates true cost by exactly 1/U, which systematically over-sells self-hosting (Beyond Per-Token Pricing). This is the cost-measurement problem in one number.
- The right unit is cost per correct answer, and it halves every few months. Cost-of-pass is per-attempt cost divided by the probability the attempt is right; measured that way, the frontier cost-of-pass for complex quantitative tasks "roughly halved every few months", and the gains come from model-level innovation rather than inference-time tricks (Cost-of-Pass, read at abstract level, so this rests on the abstract). It also explains why an expensive reasoning model gets bought: higher accuracy lowers cost per pass even as cost per token rises. See cost per task.
- Real agentic work costs single-digit dollars a task at the frontier and cents at the cheap end. Across 40 state-based personal-computing tasks and 11 frontier models, measured API cost per task ran roughly $0.17 to $6.55 — a ~38× spread on the same tasks, at 3–15 minutes per task (STAGE-Claw). Read at abstract level, and the cost table itself was surfaced by search rather than found in the fetched abstract, so full-text confirmation is still owed; the figures are API list prices, not what anyone actually paid.
- The price decline is real, fast and unequal — the median hides the number you need. Across six benchmarks the price to hold a fixed performance level fell at a median 50× a year, with a range of 9× to 900×; the widely-quoted GPT-4-level PhD-science milestone specifically fell 40× a year, below that median, and restricting to models released after January 2024 lifts the median to 200× (Epoch AI). Decomposed by tier, the half-life is 1.10 years at the economy tier and 1.55 years at mid, while the flagship tier barely fits a trend at all (R² = 0.031) because reasoning models carry a ~31.5× premium (Tiered Super-Moore's Law, read at abstract level). That is the price-decline distribution.
- Where the decline lands is the open-weight tier, not the frontier. Re-running a fixed workplace-agent benchmark two years on, best-agent task completion went 43% → 98% and unintended harmful actions 26% → 1.9%; the cost finding is that open-weight models collapsed the cost of a performance level that had been proprietary-only while frontier serving costs stayed flat (WorkBench Revisited, abstract-grade — the per-tier dollar tables have not been read). The frontier list price behaved the same way in the one provider observation held here: on 2026-07-16 two of three GPT-5.6 tiers shipped at the prior generation's exact per-token price, $5/$30 and $2.50/$15, while the new mid variant came in at $1/$6, above the previous mid rung (Token Price Index — a MenFem-collected dataset, and one whose index was retired on 2026-08-19, so it is a historical observation, not a live instrument). See price-rung persistence.
- Nobody on this rung can tell you what a token costs a lab to make. Benedict Evans puts inference at 40–50% gross margin but is explicit that this excludes training cost, "which is currently far larger than revenue" — so reported inference profitability is not lab profitability, and his own conclusion is that the question of whether foundation models become low-margin commodity infrastructure is unresolved (Ways to Think About Token Pricing). That is the margin question, still open.
- At the thin-margin layer, price is still a product decision rather than a cost floor. On one model (Llama 4 70B output) observed Q2-2026 list prices spread ~6×, from $0.65 on Together's batch tier to $4.20 on Cerebras, with throughput spreading 5–7× (Groq ~750 tokens/sec, Cerebras ~620, H100 endpoints ~115) — the spread tracks latency tolerance, SLA and hardware specialisation rather than raw cost to serve (Q2 2026 provider pricing matrix, an analysis-grade blog reading list prices on a single model).
What we do not know yet
- Almost nothing here was measured first-hand. The one own-dataset source is the Token Price
Index observation, which is a read of a published pricing page, not a measurement of cost — and
that index has since been retired. No cost curve on this rung has been reproduced locally, and
the instrument that would allow it,
vllm-cost-meter, is named in Beyond Per-Token Pricing and has not been run here. - Five of the fourteen sources were read at abstract level, and two of them carry the rung's sharpest numbers. The $0.17–$6.55 per-task range (STAGE-Claw) is search-surfaced from a cost table nobody has opened, and the tier half-lives and the 103.7%-TFP / −0.9%-hardware decomposition (Tiered Super-Moore's Law) come from an abstract page of a 26-page paper.
- Two sources disagree about what drives the price collapse, and the rung cannot settle it. The tiered-price paper attributes ~103.7% of cost reduction to total-factor productivity and −0.9% to GPU hardware — software, not silicon. Matsuoka's scenario analysis builds its whole edifice on memory and hardware vintage instead (Memory Scarcity, which this desk has not read directly — it is reconciled to the hardware desk's full-PDF close read, and its scenario probabilities are iterated expert judgment, not computed output). Both are on the record; neither has been checked against the other on shared data.
- List prices are not paid prices, anywhere on this rung. Every dollar figure held here is a published list rate — the provider matrix says so explicitly, the per-task benchmark costs are API list-based, and the Token Price Index notes that caching, batch discounts and negotiated enterprise rates are not captured. The realised-price gap is unmeasured.
- The macro channel rests on a stipulated number. The Inference-Cost Phillips Curve estimates a pass-through coefficient of 0.087 and attributes 0.18–0.41 percentage points of headline inflation to inference-cost shocks, but the size term that scales the entire channel, λ̄ = 0.18, is a calibrated Table I parameter with no stated source — an assertion that ~18% of economy-wide firm marginal cost was AI inference over 2022–2026. Its cost series is a GPU-list-price plus electricity composite, i.e. utilisation-naive by construction (Phillips Curve). See the macro channel.
- The energy leg is top-down only. The thermodynamic framing gives a TWh-to-token-count bridge — 326 TWh of projected 2028 US AI energy supporting ~6.5×10^17 tokens a year, ~225,000 tokens per person per day, more than three orders of magnitude above mid-2024 use (Photons = Tokens, abstract grade and explicitly order-of-magnitude policy framing). It says energy is not the near-term binding constraint on token volume; it does not give a bottom-up $/kWh to $/token conversion for any real SKU. See the energy cost of tokens.
- The measurement vocabulary is agreed but unused. Goodput — requests completing per second while meeting a stated service-level objective — is the metric that stops a cost-per-token figure flattering a system that is fast and unusable, and the definitions are settled (LLM Inference Handbook, a vendor-published reference from a company selling an inference stack, whose every numerical example is illustrative and which contains no measurement of any kind). No cost figure held on this rung is quoted against a goodput denominator.
Read next
- beyond-per-token-pricing-concurrency-aware-cost-methodology — the only fully-read, measured cost study here, and the one that shows why every list-price cost claim on this rung is soft.
- cost-of-pass-economic-framework-language-models — the framework that converts a token price into a task price.
- epoch-ai-llm-inference-price-trends — the price-decline baseline every other source on the rung argues with.