The Price-Decline Distribution
Active FrontierThe Price-Decline Distribution
There is no single rate at which inference prices fall. It is a distribution, and the shape of that distribution — not the headline number — is what the margin-durability question turns on. Epoch AI's cross-benchmark study is the load-bearing evidence: it measures how fast the price to hit a fixed performance bar has fallen, and finds the answer depends enormously on which capability you are buying.
The central figures (Epoch AI, methodology-documented): a median 50x/year decline across six benchmarks, with the rate ranging from 9x to 900x per year depending on the benchmark and threshold chosen. The often-quoted "GPT-4-level" milestone — GPQA Diamond, PhD-level science — specifically falls at 40x/year, below the all-benchmark median, not above it. Restricting the sample to models released after January 2024 pushes the median up to 200x/year, suggesting the pace has accelerated recently — but that faster regime, and the fastest ~900x tails, are drawn entirely from a short post-2024 window, so their persistence is unknown.
Why the distribution shape matters more than the median. The buyer's real question is not "how fast do prices fall on average" but "how fast do they fall for the capability I depend on." A task whose price collapses at 900x/year is commoditizing under your feet — value there migrates to whoever sits above the model. A task whose price falls at only 9x/year is scarce and defensible: the slow curves are the moat signal. Reading the median alone ("prices fall 50x/year, so inference is a commodity") throws away exactly the information — the spread — that tells you which slices of inference retain pricing power and which do not.
The spread resolves into tier half-lives (added 2026-07-23). Tiered Super-Moore's Law (Mingdeng Du, arXiv 2603.28576, 2026-03-30 — abstract-page read only, 26pp PDF unread) takes the same "prices fall fast and unequally" phenomenon and gives it structure Epoch does not, built on integrated datasets (OpenRouter 318 models + Epoch AI 3,237 models + 62 validated 2020-2026 milestones):
- Economy tier: ~1.10-year price half-life.
- Mid tier: ~1.55-year half-life.
- Flagship tier: near-zero exponential fit (R² = 0.031) — flagship prices don't fall cleanly, because reasoning models carry a ~31.5x price premium over non-reasoning, scrambling the trend. This is the measured mechanism behind Epoch's "flagship resists the median."
- Aggregate: ~600-fold relative decline (no absolute dated $/M in the abstract).
- A structural break dated to May 2024 (Chow F=5.74, p=0.005): before it, decline was technology-driven; after, competition-driven — with market concentration falling (HHI 4,558 → 2,086; Malmquist productivity index peaking at 4.11 in 2024). All measured price half-lives; the break and decomposition are estimated econometrics.
This is a more structured version of the KB's "slow tails are the moat signal" reading: the economy tier commoditizes on a ~1.1-year clock, the flagship tier resists because the reasoning premium dominates it. The 31.5x reasoning premium is why cost per task, not price per token, is the right unit for reasoning workloads — the premium is bought because accuracy lowers cost-per-correct-answer.
The decline is bifurcated by tier — a benchmark datapoint (added 2026-07-24). WorkBench Revisited (Styles & Miller, arXiv 2606.13715, 2026-06-10 — abstract-grade) re-runs a fixed workplace-agent benchmark two years on and observes the distribution's shape directly in cost-of-performance terms: the cost of a performance level that was once proprietary-only collapsed (~100x since 2024, per the body), driven by open-weight models — while frontier serving costs stayed flat. Over the same window best-agent task completion rose 43% (GPT-4, Mar-2024) → 98% (Claude Fable 5, mid-2026) and unintended harmful actions fell 26% → 1.9% (capability and safety moved together). The load-bearing reading for this page: the price decline shows up in the open-weight rung, not the frontier rung — the fast tails Epoch measures and the near-zero flagship fit Du measures are the same asymmetry seen from a benchmark, and it locates the commoditization where an open-weight model can reach a once-proprietary bar. It is a direct, if abstract-grade, datapoint on the open-weight-vs-frontier question the frontier gaps track; the exact per-tier $ levels need the full text. (Absolute per-task levels for this workload class — $0.17–$6.55/task on STAGE-Claw — live on cost per task.)
⚠ The software-vs-hardware conflict (added 2026-07-23). Du's headline econometric result is that ~103.7% of the cost reduction comes from total-factor productivity / software-architectural innovation and ~–0.9% from GPU hardware. In plain terms: hardware is not where the falling prices come from. This cuts directly against the memory/hardware-scarcity-as-destiny framing of Matsuoka (the depreciation conveyor) and Patel, in which hardware cost structure is the load-bearing variable. The two can be partly reconciled — memory scarcity may gate supply/capacity while software drives price — but as stated they locate the cost lever in different layers, and this KB keeps both without resolving it. Provenance to weigh: Du's decomposition is abstract-page only (the >100% TFP residual and Malmquist internals need the full PDF), a single-author preprint, not peer-reviewed. Logged as frontier Conflict 7; the same conflict is recorded on the depreciation-conveyor page.
This page measures price, not cost — and the two have come apart (added 2026-07-22). Epoch's series is built from published prices, which is what makes it strong evidence. But a falling price tells you nothing about a falling cost: Patil's measurement (arXiv 2606.11690, 2026-06-10) finds effective cost on identical H100 hardware spanning $0.21-$15.25 per million output tokens (2.5-24x at 1-10 rps, up to 36.3x near idle) purely from offered request rate. A price series declining 50x/yr against a cost basis that can move 36x with traffic shape cannot, on its own, support a statement about margins. Keep this page's claims strictly on the price side; the cost side is the cost-measurement problem.
Two mechanisms behind the decline, one of them newly named. Matsuoka (arXiv 2607.07207, 2026-07-08) projects a training-cost bifurcation — a luxury tier at $18-38B per frontier run by 2030 against a mass tier reaching previous-frontier parity via RL/distillation for a cost falling toward $5M (both modelled). If that holds, the price of yesterday's capability collapses not because anyone cut a price but because reproducing it becomes ~3-4 orders of magnitude cheaper, and frontier-capable open-weight models (Matsuoka cites GLM-5.2) put a floor-setting free option under every paid tier. That is a supply-side mechanism for exactly the fast tails Epoch measures.
A caution now attaches to the demand side of any commodity argument built on this page. Matsuoka additionally argues that public token trackers overstate monetizable demand, and that all pre-Q2-2026 projections predate the industry's shift from token maximization to token minimization. Epoch's price series is not a token-volume tracker and is not directly hit by that critique — but the demand figures this KB pairs with it (Goldman's 24x-by-2030, Gartner's 5-30x agentic per-task multipliers, both logged 2026-07-17) largely predate that line. Recorded as a live conflict: preprint critique vs bank/analyst forecast, both moderate, unresolved.
The relationship to the rung. This concept measures price falling for fixed capability (the buy-side view). Its sell-side mirror is price-rung persistence: providers hold nominal per-token price points constant while capability rises. They are the same phenomenon seen from opposite ends — if the frontier rung stays at $5/$30 while the model improves, then the price to buy yesterday's capability must collapse (you buy it a rung down, or from open source). The Epoch distribution is what that collapse looks like when measured across tasks.
Key Claims
- Median inference price decline: 50x/year, measured across six benchmarks and multiple thresholds. Evidence: strong (methodology-documented benchmark study) (Epoch AI)
- Range: 9x to 900x per year — the rate varies enormously by benchmark/threshold; there is no single "true" rate. Evidence: strong (Epoch AI)
- GPT-4-level PhD-science (GPQA Diamond): 40x/year — below the all-benchmark median, despite being the headline example most often quoted. Evidence: strong (Epoch AI)
- Post-Jan-2024-only median: 200x/year — the pace appears to have accelerated recently, but the fastest tails come entirely from a short window, so durability is unknown. Evidence: moderate (short track record) (Epoch AI)
- The slow curves (9x end) are the moat signal — where price falls slowly, capability is scarce and pricing power persists; where it falls fast (900x), commoditization is racing. Evidence: moderate (interpretation of the Epoch spread; not an Epoch claim)
- Tier half-lives: economy ~1.10yr, mid ~1.55yr, flagship near-zero fit (R²=0.031) scrambled by a ~31.5x reasoning premium; ~600x aggregate decline. Evidence: moderate-strong — measured price half-lives, abstract-page read (26pp PDF unread) (Du / Tiered Super-Moore's)
- A May-2024 structural break shifts the decline from technology-driven to competition-driven (Chow F=5.74, p=0.005); HHI falls 4,558 → 2,086. Evidence: moderate — estimated econometrics, abstract-page (Du)
- CONFLICT (unresolved): ~103.7% of the price decline attributed to software/TFP, ~–0.9% to GPU hardware — hardware is not where falling prices come from, cutting against Matsuoka/Patel memory-as-destiny. Memory may gate capacity while software drives price. Evidence: estimated econometric decomposition, abstract-page, single-author preprint; the >100% residual needs the full text (Du)
- The decline is bifurcated open-weight vs frontier: on a fixed workplace-agent benchmark the cost of a once-proprietary performance level collapsed (~100x since 2024, open-weight-driven) while frontier serving costs stayed flat; completion 43%→98% and harmful actions 26%→1.9% (Mar-2024→mid-2026). The price decline shows up in the open-weight rung, not the frontier rung. Evidence: moderate — measured longitudinal benchmark, abstract-grade; per-tier $ levels need the full text (WorkBench Revisited)
- A falling price series cannot support a margin claim on its own — effective cost on identical H100 hardware spans $0.21-$15.25/M output tokens (as of 2026-06) with traffic shape alone. Evidence: moderate (measured ratios; absolute level rests on an undisclosed $/GPU-hour input) (Patil)
- Training-cost bifurcation is a supply-side mechanism for the fast tails: $18-38B per frontier run by 2030 vs mass-tier previous-frontier parity falling toward $5M. Evidence: weak-moderate (modelled projection to 2030) (Matsuoka)
- Public token trackers overstate monetizable demand; pre-Q2-2026 demand projections predate the token-maximization → token-minimization shift. Evidence: moderate (preprint claim, magnitude not stated, not reproduced here); conflicts with the Goldman/Gartner demand figures logged 2026-07-17 (Matsuoka)
Benchmarks & Data
| Metric | Value | Source |
|---|---|---|
| Median decline (all benchmarks) | 50x / year | Epoch AI |
| Range across benchmarks/thresholds | 9x – 900x / year | Epoch AI |
| GPT-4-level PhD-science (GPQA Diamond) | 40x / year (below median) | Epoch AI |
| Median, post-Jan-2024 models only | 200x / year | Epoch AI |
| Economy-tier price half-life | ~1.10 year | Du / Tiered Super-Moore's |
| Mid-tier price half-life | ~1.55 year | Du |
| Flagship-tier fit | near-zero (R² = 0.031) | Du |
| Reasoning per-token premium | ~31.5x | Du |
| Aggregate decline (relative) | ~600x | Du |
| Structural break (tech → competition) | May 2024 (Chow F=5.74, p=0.005) | Du |
| Source of decline (estimated) | ~103.7% TFP / –0.9% GPU hardware | Du |
| Cost of a fixed performance level, open-weight vs frontier | open-weight ~100x cheaper since 2024 / frontier flat | WorkBench Revisited |
| Best-agent workplace-task completion | 43% (Mar-2024) → 98% (mid-2026) | WorkBench Revisited |
Epoch benchmarks in the study: MMLU (general knowledge), GPQA Diamond (PhD science), MATH-500 / MATH Level 5 (math), HumanEval (coding), Chatbot Arena ELO (chat). Du's figures are relative/econometric — no absolute dated $/M-token levels appear in the abstract.
Open Questions
- Does the post-2024 200x/year pace persist, or was 2024-2025 a one-off efficiency unlock (MoE, quantization, speculative decoding) that won't repeat?
- Which specific tasks sit at the slow (9x) end of the distribution — the defensible slices — and does that set shrink over time?
- Does the distribution's shape change (does the spread widen or narrow) as the frontier advances, or only its median?
- Is there a matching cost-decline distribution? This page has a price series and no cost series. Until one exists, "prices fell 50x/yr" and "margins compressed" are unconnected statements.
- Does the Epoch series hold up under the token-minimization shift? If providers and buyers both optimise for fewer tokens per task, a per-token price series and a per-task cost diverge — and per-task is the figure that matters commercially.
- Is Epoch's quality-held 50x/yr the right denominator for a commodity argument? The blended-spot price a buyer actually pays fell only ~3x/yr (2025→2026) — quality-per-dollar and unit price are different metrics that differ by >1 order of magnitude (see frontier.md → Evidence Located).
Related Concepts
- Token Pricing & the Inference-Margin Question — the umbrella question this distribution feeds: fast price decline is the mechanism by which inference could become commodity infrastructure.
- Price-Rung Persistence — the sell-side mirror: providers hold the rung while capability rises, which is what forces the buy-side price collapse this page measures.
- The Cost-Measurement Problem — this page is a price series; that page is why it cannot be read as a cost or margin series.
- The Depreciation Conveyor & Vintage Economics — training-cost bifurcation and open-weight parity are the supply-side mechanisms behind the fastest tails here; also carries the same software-vs-hardware conflict from the cost-structure side.
- Cost Per Task (Cost-of-Pass) — the reasoning premium that scrambles the flagship tier is why cost-per-task, not price-per-token, is the right unit for reasoning workloads.
Backlinks
Pages that reference this concept:
Changelog
- 2026-07-24 — Compiled WorkBench Revisited (2606.13715, abstract-grade). Added the bifurcated-decline datapoint: on a fixed workplace-agent benchmark the cost of a once-proprietary performance level collapsed ~100x since 2024 (open-weight-driven) while frontier serving costs stayed flat — the price decline lands in the open-weight rung, not the frontier rung (completion 43%→98%, harmful actions 26%→1.9%). New section, one Key Claim, two data rows; per-tier $ levels owed to the full text. Absolute per-task levels ($0.17–$6.55/task, STAGE-Claw) routed to cost per task. No change to the Epoch/Du price series.
- 2026-07-23 — Compiled Tiered Super-Moore's Law (Du, 2603.28576, abstract-page). Added tier half-lives (economy 1.10yr / mid 1.55yr / flagship near-zero, R²=0.031, 31.5x reasoning premium), the May-2024 tech→competition structural break (HHI 4,558→2,086), and — prominently — the software-vs-hardware conflict (~103.7% TFP / –0.9% GPU hardware), recorded unresolved against Matsuoka/Patel memory-as-destiny. Linked the reasoning premium to cost-per-task. No change to the Epoch price series itself.
- 2026-07-22 — Compiled Patil (2606.11690) and Matsuoka (2607.07207) against this page: added the explicit price-is-not-cost boundary, training-cost bifurcation as a supply-side mechanism for the fast tails, and Matsuoka's demand-tracker critique as a recorded conflict with the Goldman/Gartner figures logged 2026-07-17.
- 2026-07-16 — Created from the Epoch AI source (already on the shelf, ingested 2026-07-14). Split out from the founding umbrella concept to make the price-decline distribution a first-class idea; paired with the new price-rung-persistence concept as its sell-side mirror.