Inference & Token-Pricing Economics

The economics of serving AI models — token pricing trends, inference gross margins, the AI-lab P&L question, and how memory/optics/foundry cost structure feeds through to the price of a token

16 sources·8 concepts·4 entities
PAPER
2026-07-08

Memory Scarcity, Open Models, and the Restructuring of the AI Industry, 2026-2030

Satoshi Matsuoka · RIKEN Center for Computational Science (R-CCS), Japan

Quantitative 5-scenario model (Rotating Landlord Oligopoly 25% / Commoditization Crash 25% / Jevons Absorption 20% / System-Layer Re-differentiation 18% / Geopolitical Bifurcation 12%) for how DRAM/HBM pricing, open-weight frontier models, and compute-resale entrants (Meta, xAI) restructure AI-industry economics 2026-2030; introduces the 'depreciation conveyor' mechanic explaining why incumbent fleets stay cost-advantaged even as hardware prices normalize. Directly engages this topic's open 'commodity infrastructure?' question.

PAPER
2026-06-10

Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation

Chitral Patil · Independent Researcher (unnamed employer; research stated independent of it)

Shows identical H100 hardware yields $0.21-$15.25 per million output tokens (2.5x-36.3x variation) purely from request-rate/utilization assumptions — utilization-as-fixed-parameter is the dominant error source in cost calculators. Introduces C_eff = f(H,M,Q,λ,L) framework + open-sourced vllm-cost-meter tool, validated on 42 benchmarks. Directly relevant to grounding any 'cost per token' claim cited in the Evans/Patel margin debate this topic's frontier.md already tracks.

PAPER
2026-05-19

The Economics of AI Inference: Inflation Dynamics, Welfare Costs, and Optimal Monetary Policy under the Inference-Cost Phillips Curve

Gustav Olaf Yunus Laitinen-Fredriksson Lundström-Imanov · Department of Economics, Stockholm University, Sweden

First attempt to model AI inference cost as a first-order marginal-cost input to the macroeconomy: an augmented New Keynesian 'Inference-Cost Phillips Curve' (ICPC), estimated via GMM against 2022-2026 US + G7 panel data (slope coefficient 0.087, near-unit-elasticity pass-through). Speculative but a genuinely novel lens for where token-pricing economics could matter beyond the AI-lab P&L question this topic otherwise tracks.

PAPER
2026-03-30

Tiered Super-Moore's Law: Price Evolution, Production Frontiers, and Market Competition in LLM Inference Services

Mingdeng Du · Not stated on abstract page

Decomposes the ~600x token-price decline into TIER-SPECIFIC half-lives (economy 1.10yr, mid 1.55yr, flagship near-zero fit R2=0.031 from 31.5x reasoning premium); dates a May-2024 structural break (Chow F=5.74, p=0.005) from technology- to competition-driven decline; attributes ~103.7% of cost reduction to TFP vs -0.9% from GPU hardware — a direct empirical counter to the memory/hardware-centric cost story. MEASURED (OpenRouter 318 + Epoch 3,237 models + 62 milestones); econometric decomposition ESTIMATED.

PAPER
2025-04-17

Cost-of-Pass: An Economic Framework for Evaluating Language Models

Mehmet Hamza Erol, Batu El, Mirac Suzgun, Mert Yuksekgonul, James Zou · Stanford University (inferred from author roster; not stated on abstract page)

Formalizes cost-per-TASK: 'cost-of-pass' = expected monetary cost of a CORRECT solution (per-token cost / accuracy); 'frontier cost-of-pass' = min across models or human experts. Frontier cost-of-pass for complex quantitative tasks 'roughly halved every few months'; model-level (not inference-time) innovation drives gains. The canonical framework for the unit this topic said actually matters. Original 2025-04-17, revised 2026-02-26.

PAPER
2026-06-10

AI Tokenomics: The Economics of Tokens, Computation, and Pricing in Foundation Models

Quanyan Zhu · Not stated on abstract page (Zhu is NYU Tandon; unconfirmed from source)

Proposes 'AI tokenomics' as a formal field; load-bearing claim: token EXPENDITURE and economic VALUE are distinct — value depends on marginal productivity, workflow position, hidden reasoning activity, risk, downstream propagation, NOT tokens burned. So per-token pricing prices the input not the output. Names 'hidden-token measurement' and 'empirical calibration' as open — conceding the theory has no measured cost series.

PAPER
2026-06-09

STAGE-Claw: Automated State-based Agent Benchmarking for Realistic Scenarios

Liang, Yu, Wang, Guo, Hu, Cao, Zhao, Liu, Zeng, Cai, Liu · Not stated in abstract (Meituan-affiliated author list)

State-based personal-agent benchmark (40 tasks, 11 frontier models) scored by final SYSTEM-STATE correctness rather than text; reports a measured per-task API cost ~$0.17 (cheapest, MiniMax-tier) to ~$6.55 (Claude-Opus-4.7-tier) — a ~38x spread — the KB's FIRST dated absolute cost-per-task levels, partly closing the topic's sharpest cost-per-task gap. The $0.17-$6.55 figures are search-surfaced from the paper's cost table, NOT in the fetched abstract; full-text confirmation owed. API list-price-based, personal-computing scenarios only.

PAPER
2026-06-10

WorkBench Revisited: Workplace Agents Two Years On

Olly Styles, Sam Miller · Not stated in abstract

Two-year re-run of the WorkBench workplace-agent benchmark: best-agent task completion 43% (GPT-4, Mar-2024) -> 98% (Claude Fable 5, mid-2026); unintended harmful actions 26% -> 1.9% (capability and safety moved TOGETHER). KB-relevant finding: open-weight models collapsed the cost of a given performance level while FRONTIER serving costs stayed flat — a direct benchmark datapoint that the price decline shows up in the open-weight rung, not the frontier rung (open-weight-driven, ~100x since 2024 per the body). Per-tier $ tables not in abstract; full-text read owed.

PAPER
2026-02-23

Photons = Tokens: The Physics of AI and the Economics of Knowledge

Alec Litowitz, Nick Polson, Vadim Sokolov · Not stated in abstract

Treats the token as a physical quantity with a measurable thermodynamic cost (Landauer's principle + Shannon channel capacity) and builds a supply/demand balance sheet for global token production — deriving the energy->token bridge the KB's 'no energy pass-through leg' gap was missing: the projected 2028 US AI energy allocation of 326 TWh could support ~6.5x10^17 tokens/yr ~225,000 tokens/person/day, >3 orders of magnitude above mid-2024 utilization (i.e. energy is not the near-term binding constraint on token VOLUME; direction — which questions are worth asking — is). Order-of-magnitude policy-framing estimates, not a measured per-SKU cost series.

REPORT
2025-03-12

LLM Inference Prices Have Fallen Rapidly but Unequally Across Tasks

Ben Cottier, Ben Snodin, David Owen, Tom Adamczewski · Epoch AI

Median LLM inference price decline 50x/yr across 6 benchmarks (range 9x-900x/yr); GPT-4-level PhD-science (GPQA Diamond) milestone specifically 40x/yr, below the all-benchmark median; post-Jan-2024-only median rises to 200x/yr.

REPORT
2026-07-16

Token Price Index — GPT-5.6 holds the prior-generation price rungs

MenFem Token Price Index (own dataset) · MenFem — research/inference/token-prices.csv

GPT-5.6 (2026-07-16) holds the exact prior-generation per-token price at two of three tiers: Sol (frontier) $5/$30 = GPT-5.5's rung, Terra (strong) $2.50/$15 = GPT-5.4's rung; Luna (mid) $1/$6 is a new point. Capability up, price rungs held.

REPORT
2026-05-25

Is a Compute Crunch Coming? (inference serving-capacity model, calibrated to SemiAnalysis InferenceX)

Luke Emberson, Jaime Sevilla · Epoch AI

First-principles prefill(compute-bound)/decode(bandwidth-bound) serving-throughput model CALIBRATED to 111 measured SemiAnalysis InferenceX Kimi K2.5 runs (fitted: compute eff 65%, bandwidth eff 30%, 5ms/step). GB200 NVL72 ~400k tok/s; global capacity 500M-20B tok/s; capacity growth ~3.4x/yr vs demand ~10x/yr => crunch. Denominated in tokens/sec, NO $/token. As-of 2026-05-25.

ANALYSIS
2026-07-09

Ways to Think About Token Pricing

Benedict Evans · ben-evans.com (independent)

Frames token pricing as unresolved between sellers' marginal cost and buyers' ROI; current 40-50% inference gross margin excludes training cost which exceeds revenue; mobile-data-network analogy for commodity-infrastructure risk.

ANALYSIS
2026-07-08

Token Economics Across Traffic Profiles on Dedicated GPUs (measured H200 serving-cost benchmark)

Vinayak Baranwal · DigitalOcean

First MEASURED cost-per-token series in this KB: single H200 ($3.44/GPU-hr, Llama-3.3-70B FP8, vLLM 0.24.0), swept batch 1->256, yields $20.32 -> $0.45 per M output tokens (~44x spread on ONE SKU from traffic shape alone), independently reproducing Patil's utilization thesis with dollar-anchored levels. As-of 2026-07-08; MEASURED throughput, cost derived at list rate.

ANALYSIS
2026-04-24

AI Inference Providers Compared: Q2 2026 Pricing Matrix (independent-provider observed pricing)

Digital Applied Team · Digital Applied

First independent-provider pricing in the KB: on ONE model (Llama 4 70B output), observed Q2-2026 LIST prices spread 6x — Together $0.65 (batch) to Cerebras $4.20 (wafer-scale); throughput spread 5-7x (Groq 750 tps, Cerebras 620 vs H100 ~115). Cheapest list ($0.65/M batch) sits at DigitalOcean's measured ~$0.45-0.64/M cost floor => thin-to-zero margin at batch tier. Opens the layer where marginal-cost pricing shows up first.

ANALYSIS
2026-07-10

Dylan Patel of SemiAnalysis: The $11M Bill, the Memory Shortage, and Why CPO Is Two Years Late

Dylan Patel (interviewed) · Podcast Alpha / SemiAnalysis

ATTRIBUTED/UNVERIFIED: Anthropic FCF-positive, $50B+ ARR, 70%+ gross margin; memory multi-year structural shortage (capacity +20-30%/yr vs demand doubling); CPO real ramp 2029 vs Street 2027 (Rubin/Feynman all-copper; Amphenol named beneficiary). Disclosed conflict: Patel/SemiAnalysis is a reported Anthropic enterprise customer.

Inference & Token-Pricing Economics | Knowledge Base | MenFem