Inference & Token-Pricing Economics — Timeline

Timeline

Dates are publication dates of the underlying source, not ingest dates. Every cost/price figure carries its as-of date because these numbers decay fast.

2026

July

  • Jul 23 — [Compile] Six new sources compiled (13 total): DigitalOcean's measured H200 cost series, Epoch's serving-capacity model, Tiered Super-Moore's Law, Cost-of-Pass, AI Tokenomics, and the DigitalApplied Q2-2026 provider matrix. The topic's defining gap partly closes — a measured cost-to-serve series now exists (for one open 70B model). Two new pages: cost per task and the independent inference providers entity. A new first-order conflict recorded: software/TFP (~103.7%) vs GPU hardware (–0.9%) as the source of falling prices, cutting against memory-as-destiny. The sharpest remaining gap: no measured cost floor for any frontier/proprietary model.
  • Jul 22 — [Compile] Three arXiv preprints ingested and compiled (abstract pages only): Matsuoka 2607.07207, Patil 2606.11690, ICPC 2605.20281. Topic gains its first research-grade sources and three concepts — the cost-measurement problem, the depreciation conveyor, inference cost as a macro input. Net effect subtractive: both standing margin figures (Evans 40-50%, Patel 70%+) recorded as unfalsifiable as stated for want of a cost denominator.
  • Jul 16 — [Data] OpenAI releases GPT-5.6 (Sol / Terra / Luna). Observed into the MenFem Token Price Index the same day: Sol (frontier) and Terra (strong) hold the exact prior-generation price rungs ($5/$30, $2.50/$15); Luna is a new mid point ($1/$6). Own-dataset primary, anchors price-rung persistence. (TPI)
  • Jul 10 — [Opinion/attributed] Dylan Patel (SemiAnalysis) claims on Podcast Alpha that Anthropic is FCF-positive, past $50B ARR, at 70%+ gross margin — well above Evans' 40-50%. Unverified, disclosed conflict. (podcast)
  • Jul 09 — [Analysis] Ben Evans publishes "Ways to Think About Token Pricing" — the four-question framework and the mobile-carrier commodity-infrastructure analogy. Inference gross margin put at 40-50%, training cost excluded. (essay)
  • Jul 08 — [Measured] DigitalOcean (Baranwal) publishes the topic's first measured cost-to-serve series: one H200 ($3.44/GPU-hr, Llama-3.3-70B FP8, vLLM 0.24.0) at $20.32/M output tokens at batch=1 → $0.45/M at batch=128 (~44x from traffic shape alone); own-GPU-vs-serverless crossover ≈73% utilization. Throughput measured; $/M derived at list rate. One SKU / one open model / one date — a point, not a series. (measured)
  • Jul 08 — [Preprint] Satoshi Matsuoka (RIKEN R-CCS) posts "Memory Scarcity, Open Models, and the Restructuring of the AI Industry, 2026-2030" (arXiv 2607.07207). Reformulates inference economics in $/PB of bandwidth delivered; names the depreciation conveyor (entrant-incumbent gap 3.2x→1.9x→3-4x, modelled); prices buildout solvency as a two-condition corridor; scenario probabilities put Commoditization Crash and Rotating Landlord Oligopoly tied at 25%. (preprint)

June

  • Jun 10 — [Preprint] Chitral Patil posts "Beyond Per-Token Pricing" (arXiv 2606.11690). Measures $0.21-$15.25 per million output tokens on identical H100 hardware (up to 36.3x near idle), shows every surveyed public cost calculator treats utilization as a fixed input, and releases vllm-cost-meter. (preprint)
  • Jun 10 — [Preprint] Quanyan Zhu posts "AI Tokenomics" (arXiv 2606.24616). Proposes AI tokenomics as a formal field; load-bearing claim: token expenditure ≠ economic value — per-token pricing prices the input, not the output. Theoretical, no measured figures. Anchors cost per task. (preprint)

May

  • May 25 — [Report] Epoch AI (Emberson & Sevilla) publishes "Is a Compute Crunch Coming?" — a first-principles prefill(compute-bound)/decode(bandwidth-bound) serving-capacity model calibrated to 111 SemiAnalysis InferenceX runs (compute eff 65% / bandwidth eff 30% / 5ms/step); ~400k tok/s on GB200 NVL72; capacity growth ~3.4x/yr vs demand ~10x/yr → crunch. Tokens/sec, no $. (report)
  • May 19 — [Preprint] "The Economics of AI Inference … Inference-Cost Phillips Curve" (arXiv 2605.20281, Stockholm Univ.) proposes inference cost as a first-order macro marginal-cost input; US GMM slope 0.087 (HAC s.e. 0.021) on 2022:M01-2026:M04 data, G7 panel 0.094. Filed speculative. (preprint)

April

  • Apr 24 — [Analysis] Digital Applied publishes the Q2-2026 independent-provider pricing matrix. On one model (Llama 4 70B output), observed list prices spread ~6x — Together $0.65 (batch) to Cerebras $4.20 (wafer-scale); throughput spread 5-7x. Cheapest list sits on the measured cost floor → thin batch margin. Anchors independent inference providers. (analysis)

March

  • Mar 30 — [Preprint] Mingdeng Du posts "Tiered Super-Moore's Law" (arXiv 2603.28576). Decomposes the ~600x token-price decline into tier half-lives (economy 1.10yr / mid 1.55yr / flagship near-zero, 31.5x reasoning premium); dates a May-2024 structural break (tech → competition, HHI 4,558→2,086); attributes ~103.7% of the cost fall to software/TFP and –0.9% to GPU hardware — the topic's software-vs-hardware conflict. Abstract-page read. (preprint)

2025

April

  • Apr 17 — [Preprint] Erol, El, Suzgun, Yuksekgonul & Zou post "Cost-of-Pass: An Economic Framework for Evaluating Language Models" (arXiv 2504.13359; revised 2026-02-26). Formalizes cost per task — cost-of-pass = per-token cost ÷ accuracy — and finds frontier cost-of-pass on hard quantitative tasks "roughly halved every few months," model-level innovation the driver. The earliest-dated source in the topic; anchors cost per task. (preprint)

March

  • Mar 12 — [Data] Epoch AI publishes "LLM Inference Prices Have Fallen Rapidly but Unequally Across Tasks" — median 50x/yr decline (9x-900x range), GPT-4-level milestone at 40x/yr (below median), post-Jan-2024 median 200x/yr. Ingested to this KB 2026-07-14; grounds the price-decline distribution. (Epoch AI)
Timeline — Inference & Token-Pricing Economics | KB | MenFem