Inference & Token-Pricing Economics
The economics of serving AI models — token pricing trends, inference gross margins, the AI-lab P&L question, and how memory/optics/foundry cost structure feeds through to the price of a token
Inference & Token-Pricing Economics
The economics of serving AI models: what a token actually costs to produce, how fast that cost is falling, whether the margin on top of it is durable or transitory, and how upstream cost structure (memory, optics, foundry) feeds through into the price a buyer pays. Created 2026-07-14 because no existing KB topic tracked token-pricing/inference-margin economics as a first-class subject.
As of the 2026-07-23 compile, the topic has four layers. At the top, the umbrella question — the inference-margin question: durable-margin business or commodity infrastructure? Beneath it, two mirror-image price empirics: the price-decline distribution (buy-side — how fast the price of fixed capability falls: Epoch's median 50x/yr, 9x-900x range; now with Du's tier half-lives, economy 1.10yr / mid 1.55yr) and price-rung persistence (sell-side — providers hold nominal per-token rungs while capability rises: GPT-5.6 held two of three prior-generation rungs to the cent). Underneath both, the cost layer: the cost-measurement problem (now holding the topic's first measured cost-to-serve series) and the depreciation conveyor (whose cost is lower, why, and what the buildout needs to stay solvent). And a unit layer that says the token is the wrong denominator: cost per task (price a correct answer, not a token). A seventh concept, inference cost as a macro input, holds the speculative branch: what if inference cost passes through to the general price level.
The 2026-07-23 compile's headline is that the topic's defining gap is partly closed — and the remaining gap is now sharply stated. DigitalOcean measured a 70B open model (Llama-3.3-70B FP8) on a single H200 at $20.32/M at batch=1 → $0.45/M at batch=128 (~44x from traffic shape alone; 2026-07-08). That is the topic's first measured cost-to-serve series, and it independently reproduces Patil's utilization thesis in dollars. But it is a POINT, not a series (one SKU / one open model / one framework / one date), and — the sharpest remaining gap — there is still no measured cost floor for any frontier/proprietary model, so the Evans-40% vs Patel-70% lab-margin debate stays unfalsifiable. The cheapest independent-provider list price (Together $0.65/M batch, 70B) sits on that measured floor → thin batch-tier margin.
A new first-order conflict is on the record, unresolved. Tiered Super-Moore's Law attributes ~103.7% of the token-price decline to software/total-factor productivity and –0.9% to GPU hardware, cutting directly against the memory/hardware-scarcity-as-destiny framing of Matsuoka and Patel. Memory may gate capacity while software drives price; the sources locate the cost lever in different layers. Recorded on the price-decline and depreciation-conveyor pages and as frontier Conflict 7 — not adjudicated.
Evidence discipline. Epoch AI = strong (documented methodology). The Token Price Index rung observation = strong, own dataset, primary against OpenAI's own pricing page (2026-07-16). DigitalOcean's cost floor = measured throughput × list rate, read in full, but analysis-grade and single-config. Epoch's serving-capacity model = modelled but calibrated to 111 measured InferenceX runs, read in full. Ben Evans = moderate (single analyst essay). The six arXiv sources = preprints, single-author or small-team, not peer-reviewed; the four newest (Tiered Super-Moore's, Cost-of-Pass, AI Tokenomics, and the DigitalApplied matrix's arXiv siblings) were read from abstract/article pages only — their measured ratios/trajectories are moderate, their modelled projections are weak, and that split is marked on every page. Dylan Patel's Anthropic-margin claim = weak (attributed, disclosed conflict, paywalled). Every cost or price figure carries its as-of date, and every number is labelled measured / modelled / projected / estimated / derived / observed. No directional calls.
Concept Map
Concepts
| Concept | Key Claim | Evidence | Last Updated |
|---|---|---|---|
| Token Pricing & the Inference-Margin Question | Umbrella: commodity infrastructure or durable margin? Evans 40-50% vs Patel's attributed 70%+; a measured open-model cost floor now exists but no frontier floor, so the debate stays unfalsifiable; the first quantitative model ties the two poles at 25% each | Mixed (strong/moderate/weak by claim) | 2026-07-23 |
| The Cost-Measurement Problem | Gap PARTLY CLOSED: one H200 measures $0.45-$20.32/M (70B open, ~44x from traffic); Epoch models the prefill/decode physics; no frontier-model floor exists — now the sharpest gap | Moderate (measured floor + calibrated model; single-config) | 2026-07-23 |
| Cost Per Task (Cost-of-Pass) | The unit that actually matters: cost-of-pass = per-token cost ÷ accuracy; token expenditure ≠ value; frontier cost-of-pass halving every few months; no absolute dated level held | Moderate (frameworks; two abstract-page preprints) | 2026-07-23 |
| The Depreciation Conveyor & Vintage Economics | Entrant-incumbent cost gap never closes (3.2x→1.9x→3-4x, modelled); solvency needs ~2x/yr token demand for 4 yrs AND sticky premium pricing; NOW carries the software-vs-hardware conflict | Weak-moderate (modelled, single-author preprint) | 2026-07-23 |
| The Price-Decline Distribution | Median 50x/yr, 9x-900x; GPT-4-level 40x (below median); Du's tier half-lives (economy 1.10yr/mid 1.55yr) + May-2024 competition break; software-vs-hardware conflict (~103.7% TFP / -0.9% hardware) | Strong (Epoch) + moderate (Du, abstract-page) | 2026-07-23 |
| Price-Rung Persistence | GPT-5.6 held 2 of 3 prior-gen rungs exactly (Sol $5/$30, Terra $2.50/$15); a named solvency variable; Epoch's 3.4x-supply/10x-demand gap is a cost-side reason rungs stay sticky | Strong (own dataset, primary) | 2026-07-23 |
| Inference Cost as a Macro Input | Inference-Cost Phillips Curve: US slope 0.087, near-unit-elasticity pass-through claimed. Filed speculative — three named reasons for low confidence (λ̄=0.18 implausible) | Weak (unreplicated preprint) | 2026-07-23 |
Entities
Topic-lensed stubs only — canonical company profiles live in the ai/ topic and are cross-linked, not duplicated.
| Entity | Type | Lens | Key Claim |
|---|---|---|---|
| OpenAI (pricing) | Company | pricing behaviour | Holds capability-tier price rungs across generations (GPT-5.6). Primary, own dataset. List prices only. |
| Anthropic (economics) | Company | unit economics | FCF-positive / $50B+ ARR / 70%+ GM — attributed (Patel), conflict disclosed, no cost denominator. |
| Independent Inference Providers | Market segment | independent-provider pricing | Together/Fireworks/Groq/Cerebras cohort; ~6x same-model list spread; cheapest list ($0.65/M batch) sits on the measured cost floor → thin batch margin. Observed list, one model, Q2 2026. |
| vllm-cost-meter | Product | measurement instrument | Open-source meter reporting real $/M-tokens off a live vLLM server against the operator's own traffic. Repo not inspected. |
Sources
| Slug | Title | Type | Institution | Published | Status |
|---|---|---|---|---|---|
| cost-of-pass-economic-framework-language-models | Cost-of-Pass: An Economic Framework for Evaluating Language Models | preprint | arXiv 2504.13359 (Stanford, inferred) | 2025-04-17 (rev 2026-02-26) | compiled |
| epoch-ai-llm-inference-price-trends | LLM Inference Prices Have Fallen Rapidly but Unequally Across Tasks | technical-report | Epoch AI | 2025-03-12 | compiled |
| tiered-super-moores-law-inference-price-evolution | Tiered Super-Moore's Law: Price Evolution … in LLM Inference Services | preprint | arXiv 2603.28576 | 2026-03-30 | compiled |
| independent-inference-provider-pricing-matrix-q2-2026 | AI Inference Providers Compared: Q2 2026 Pricing Matrix | analysis | Digital Applied | 2026-04-24 | compiled |
| economics-of-ai-inference-phillips-curve | The Economics of AI Inference: … Inference-Cost Phillips Curve | preprint | arXiv 2605.20281 (Stockholm Univ.) | 2026-05-19 | compiled |
| epoch-inference-compute-crunch-serving-capacity | Is a Compute Crunch Coming? (serving-capacity model, InferenceX-calibrated) | technical-report | Epoch AI | 2026-05-25 | compiled |
| beyond-per-token-pricing-concurrency-aware-cost-methodology | Beyond Per-Token Pricing: A Concurrency-Aware Methodology … | preprint | arXiv 2606.11690 | 2026-06-10 | compiled |
| ai-tokenomics-economics-tokens-computation-pricing | AI Tokenomics: The Economics of Tokens, Computation, and Pricing … | preprint | arXiv 2606.24616 | 2026-06-10 | compiled |
| measured-h200-token-serving-cost-digitalocean | Token Economics Across Traffic Profiles on Dedicated GPUs (measured H200) | analysis | DigitalOcean | 2026-07-08 | compiled |
| memory-scarcity-open-models-ai-industry-restructuring-2026-2030 | Memory Scarcity, Open Models, and the Restructuring of the AI Industry | preprint | arXiv 2607.07207 (RIKEN R-CCS) | 2026-07-08 | compiled |
| ben-evans-ways-to-think-about-token-pricing | Ways to Think About Token Pricing | analysis | ben-evans.com | 2026-07-09 | compiled |
| dylan-patel-podcast-alpha-2026-07-10 | Dylan Patel of SemiAnalysis: The $11M Bill, the Memory Shortage … | opinion (attributed) | Podcast Alpha / SemiAnalysis | 2026-07-10 | compiled |
| tpi-gpt-5-6-rung-hold-2026-07-16 | Token Price Index — GPT-5.6 holds the prior-generation price rungs | technical-report (own dataset) | MenFem TPI | 2026-07-16 | compiled |
All 13 registered sources are ingested and compiled; staleSources is 0. DigitalOcean and Epoch (compute-crunch) were read in full; the four newest arXiv/analysis economics sources (Tiered Super-Moore's, Cost-of-Pass, AI Tokenomics, DigitalApplied matrix) were read from abstract/article pages only. The three older arXiv preprints (Matsuoka, Patil, ICPC) have had full-text reconciliation/touch-ups. Every raw file states its read-depth at the top. Zero peer-reviewed papers.
Cross-Topic Links
- The Matsuoka preprint (arXiv 2607.07207) is also held in
hardware/, with a desk close-read at hardware/raw/memory-scarcity-2607.07207-closeread.md. This topic holds the inference-economics lens on it ($/PB as a cost unit, the depreciation conveyor, the solvency corridor);hardware/owns the memory-supply read. Deliberately not duplicated. - Memory-shortage claim (Patel, attributed) folded into hardware/wiki/concepts/hbm4-memory-architecture.md
- CPO-ramp-timing claim (Patel, attributed) folded into optical-computing/wiki/concepts/co-packaged-optics.md
- Anthropic financial claim (Patel, attributed) canonical on ai/wiki/entities/anthropic.md; economics-lensed stub here at entities/anthropic-economics.md
- OpenAI canonical profile at ai/wiki/entities/openai.md; pricing-lensed stub here at entities/openai-pricing.md
Timeline
See timeline.md for chronological developments.
Research Frontier
See frontier.md.