Inference & Token-Pricing Economics

The economics of serving AI models — token pricing trends, inference gross margins, the AI-lab P&L question, and how memory/optics/foundry cost structure feeds through to the price of a token

16 sources·8 concepts·4 entities

Inference & Token-Pricing Economics

The economics of serving AI models: what a token actually costs to produce, how fast that cost is falling, whether the margin on top of it is durable or transitory, and how upstream cost structure (memory, optics, foundry) feeds through into the price a buyer pays. Created 2026-07-14 because no existing KB topic tracked token-pricing/inference-margin economics as a first-class subject.

As of the 2026-07-23 compile, the topic has four layers. At the top, the umbrella question — the inference-margin question: durable-margin business or commodity infrastructure? Beneath it, two mirror-image price empirics: the price-decline distribution (buy-side — how fast the price of fixed capability falls: Epoch's median 50x/yr, 9x-900x range; now with Du's tier half-lives, economy 1.10yr / mid 1.55yr) and price-rung persistence (sell-side — providers hold nominal per-token rungs while capability rises: GPT-5.6 held two of three prior-generation rungs to the cent). Underneath both, the cost layer: the cost-measurement problem (now holding the topic's first measured cost-to-serve series) and the depreciation conveyor (whose cost is lower, why, and what the buildout needs to stay solvent). And a unit layer that says the token is the wrong denominator: cost per task (price a correct answer, not a token). A seventh concept, inference cost as a macro input, holds the speculative branch: what if inference cost passes through to the general price level.

The 2026-07-23 compile's headline is that the topic's defining gap is partly closed — and the remaining gap is now sharply stated. DigitalOcean measured a 70B open model (Llama-3.3-70B FP8) on a single H200 at $20.32/M at batch=1 → $0.45/M at batch=128 (~44x from traffic shape alone; 2026-07-08). That is the topic's first measured cost-to-serve series, and it independently reproduces Patil's utilization thesis in dollars. But it is a POINT, not a series (one SKU / one open model / one framework / one date), and — the sharpest remaining gap — there is still no measured cost floor for any frontier/proprietary model, so the Evans-40% vs Patel-70% lab-margin debate stays unfalsifiable. The cheapest independent-provider list price (Together $0.65/M batch, 70B) sits on that measured floor → thin batch-tier margin.

A new first-order conflict is on the record, unresolved. Tiered Super-Moore's Law attributes ~103.7% of the token-price decline to software/total-factor productivity and –0.9% to GPU hardware, cutting directly against the memory/hardware-scarcity-as-destiny framing of Matsuoka and Patel. Memory may gate capacity while software drives price; the sources locate the cost lever in different layers. Recorded on the price-decline and depreciation-conveyor pages and as frontier Conflict 7 — not adjudicated.

Evidence discipline. Epoch AI = strong (documented methodology). The Token Price Index rung observation = strong, own dataset, primary against OpenAI's own pricing page (2026-07-16). DigitalOcean's cost floor = measured throughput × list rate, read in full, but analysis-grade and single-config. Epoch's serving-capacity model = modelled but calibrated to 111 measured InferenceX runs, read in full. Ben Evans = moderate (single analyst essay). The six arXiv sources = preprints, single-author or small-team, not peer-reviewed; the four newest (Tiered Super-Moore's, Cost-of-Pass, AI Tokenomics, and the DigitalApplied matrix's arXiv siblings) were read from abstract/article pages only — their measured ratios/trajectories are moderate, their modelled projections are weak, and that split is marked on every page. Dylan Patel's Anthropic-margin claim = weak (attributed, disclosed conflict, paywalled). Every cost or price figure carries its as-of date, and every number is labelled measured / modelled / projected / estimated / derived / observed. No directional calls.

Concept Map

Concepts

ConceptKey ClaimEvidenceLast Updated
Token Pricing & the Inference-Margin QuestionUmbrella: commodity infrastructure or durable margin? Evans 40-50% vs Patel's attributed 70%+; a measured open-model cost floor now exists but no frontier floor, so the debate stays unfalsifiable; the first quantitative model ties the two poles at 25% eachMixed (strong/moderate/weak by claim)2026-07-23
The Cost-Measurement ProblemGap PARTLY CLOSED: one H200 measures $0.45-$20.32/M (70B open, ~44x from traffic); Epoch models the prefill/decode physics; no frontier-model floor exists — now the sharpest gapModerate (measured floor + calibrated model; single-config)2026-07-23
Cost Per Task (Cost-of-Pass)The unit that actually matters: cost-of-pass = per-token cost ÷ accuracy; token expenditure ≠ value; frontier cost-of-pass halving every few months; no absolute dated level heldModerate (frameworks; two abstract-page preprints)2026-07-23
The Depreciation Conveyor & Vintage EconomicsEntrant-incumbent cost gap never closes (3.2x→1.9x→3-4x, modelled); solvency needs ~2x/yr token demand for 4 yrs AND sticky premium pricing; NOW carries the software-vs-hardware conflictWeak-moderate (modelled, single-author preprint)2026-07-23
The Price-Decline DistributionMedian 50x/yr, 9x-900x; GPT-4-level 40x (below median); Du's tier half-lives (economy 1.10yr/mid 1.55yr) + May-2024 competition break; software-vs-hardware conflict (~103.7% TFP / -0.9% hardware)Strong (Epoch) + moderate (Du, abstract-page)2026-07-23
Price-Rung PersistenceGPT-5.6 held 2 of 3 prior-gen rungs exactly (Sol $5/$30, Terra $2.50/$15); a named solvency variable; Epoch's 3.4x-supply/10x-demand gap is a cost-side reason rungs stay stickyStrong (own dataset, primary)2026-07-23
Inference Cost as a Macro InputInference-Cost Phillips Curve: US slope 0.087, near-unit-elasticity pass-through claimed. Filed speculative — three named reasons for low confidence (λ̄=0.18 implausible)Weak (unreplicated preprint)2026-07-23

Entities

Topic-lensed stubs only — canonical company profiles live in the ai/ topic and are cross-linked, not duplicated.

EntityTypeLensKey Claim
OpenAI (pricing)Companypricing behaviourHolds capability-tier price rungs across generations (GPT-5.6). Primary, own dataset. List prices only.
Anthropic (economics)Companyunit economicsFCF-positive / $50B+ ARR / 70%+ GM — attributed (Patel), conflict disclosed, no cost denominator.
Independent Inference ProvidersMarket segmentindependent-provider pricingTogether/Fireworks/Groq/Cerebras cohort; ~6x same-model list spread; cheapest list ($0.65/M batch) sits on the measured cost floor → thin batch margin. Observed list, one model, Q2 2026.
vllm-cost-meterProductmeasurement instrumentOpen-source meter reporting real $/M-tokens off a live vLLM server against the operator's own traffic. Repo not inspected.

Sources

SlugTitleTypeInstitutionPublishedStatus
cost-of-pass-economic-framework-language-modelsCost-of-Pass: An Economic Framework for Evaluating Language ModelspreprintarXiv 2504.13359 (Stanford, inferred)2025-04-17 (rev 2026-02-26)compiled
epoch-ai-llm-inference-price-trendsLLM Inference Prices Have Fallen Rapidly but Unequally Across Taskstechnical-reportEpoch AI2025-03-12compiled
tiered-super-moores-law-inference-price-evolutionTiered Super-Moore's Law: Price Evolution … in LLM Inference ServicespreprintarXiv 2603.285762026-03-30compiled
independent-inference-provider-pricing-matrix-q2-2026AI Inference Providers Compared: Q2 2026 Pricing MatrixanalysisDigital Applied2026-04-24compiled
economics-of-ai-inference-phillips-curveThe Economics of AI Inference: … Inference-Cost Phillips CurvepreprintarXiv 2605.20281 (Stockholm Univ.)2026-05-19compiled
epoch-inference-compute-crunch-serving-capacityIs a Compute Crunch Coming? (serving-capacity model, InferenceX-calibrated)technical-reportEpoch AI2026-05-25compiled
beyond-per-token-pricing-concurrency-aware-cost-methodologyBeyond Per-Token Pricing: A Concurrency-Aware Methodology …preprintarXiv 2606.116902026-06-10compiled
ai-tokenomics-economics-tokens-computation-pricingAI Tokenomics: The Economics of Tokens, Computation, and Pricing …preprintarXiv 2606.246162026-06-10compiled
measured-h200-token-serving-cost-digitaloceanToken Economics Across Traffic Profiles on Dedicated GPUs (measured H200)analysisDigitalOcean2026-07-08compiled
memory-scarcity-open-models-ai-industry-restructuring-2026-2030Memory Scarcity, Open Models, and the Restructuring of the AI IndustrypreprintarXiv 2607.07207 (RIKEN R-CCS)2026-07-08compiled
ben-evans-ways-to-think-about-token-pricingWays to Think About Token Pricinganalysisben-evans.com2026-07-09compiled
dylan-patel-podcast-alpha-2026-07-10Dylan Patel of SemiAnalysis: The $11M Bill, the Memory Shortage …opinion (attributed)Podcast Alpha / SemiAnalysis2026-07-10compiled
tpi-gpt-5-6-rung-hold-2026-07-16Token Price Index — GPT-5.6 holds the prior-generation price rungstechnical-report (own dataset)MenFem TPI2026-07-16compiled

All 13 registered sources are ingested and compiled; staleSources is 0. DigitalOcean and Epoch (compute-crunch) were read in full; the four newest arXiv/analysis economics sources (Tiered Super-Moore's, Cost-of-Pass, AI Tokenomics, DigitalApplied matrix) were read from abstract/article pages only. The three older arXiv preprints (Matsuoka, Patil, ICPC) have had full-text reconciliation/touch-ups. Every raw file states its read-depth at the top. Zero peer-reviewed papers.

Cross-Topic Links

Timeline

See timeline.md for chronological developments.

Research Frontier

See frontier.md.