# The Price-Decline Distribution

Canonical URL: https://menfem.com/kb/inference-economics/concepts/price-decline-distribution
Knowledge base topic: [Inference & Token-Pricing Economics](https://menfem.com/kb/inference-economics)
Frontier status: active
Tags: token-pricing, inference-economics, price-performance, benchmarks

---

There is no single rate at which inference prices fall. It is a **distribution**, and the shape of that distribution — not the headline number — is what the margin-durability question turns on. Epoch AI's cross-benchmark study is the load-bearing evidence: it measures how fast the price to hit a *fixed* performance bar has fallen, and finds the answer depends enormously on which capability you are buying.

**The central figures (Epoch AI, methodology-documented):** a **median 50x/year** decline across six benchmarks, with the rate ranging from **9x to 900x per year** depending on the benchmark and threshold chosen. The often-quoted "GPT-4-level" milestone — GPQA Diamond, PhD-level science — specifically falls at **40x/year, *below* the all-benchmark median**, not above it. Restricting the sample to models released after January 2024 pushes the median up to **200x/year**, suggesting the pace has accelerated recently — but that faster regime, and the fastest ~900x tails, are drawn entirely from a short post-2024 window, so their persistence is unknown.

**Why the distribution shape matters more than the median.** The buyer's real question is not "how fast do prices fall on average" but "how fast do they fall *for the capability I depend on*." A task whose price collapses at 900x/year is commoditizing under your feet — value there migrates to whoever sits above the model. A task whose price falls at only 9x/year is *scarce and defensible*: the slow curves are the moat signal. Reading the median alone ("prices fall 50x/year, so inference is a commodity") throws away exactly the information — the spread — that tells you which slices of inference retain pricing power and which do not.

**The spread resolves into tier half-lives (added 2026-07-23).** Tiered Super-Moore's Law (Mingdeng Du, arXiv 2603.28576, 2026-03-30 — abstract-page read only, 26pp PDF unread) takes the same "prices fall fast and unequally" phenomenon and gives it structure Epoch does not, built on integrated datasets (OpenRouter 318 models + Epoch AI 3,237 models + 62 validated 2020-2026 milestones):

- **Economy tier: ~1.10-year price half-life.**
- **Mid tier: ~1.55-year half-life.**
- **Flagship tier: near-zero exponential fit (R² = 0.031)** — flagship prices *don't* fall cleanly, because **reasoning models carry a ~31.5x price premium over non-reasoning**, scrambling the trend. This is the measured mechanism behind Epoch's "flagship resists the median."
- Aggregate: ~**600-fold** relative decline (no absolute dated $/M in the abstract).
- **A structural break dated to May 2024** (Chow F=5.74, p=0.005): before it, decline was *technology-driven*; after, *competition-driven* — with market concentration falling (HHI 4,558 → 2,086; Malmquist productivity index peaking at 4.11 in 2024). *All measured price half-lives; the break and decomposition are estimated econometrics.*

This is a more structured version of the KB's "slow tails are the moat signal" reading: the economy tier commoditizes on a ~1.1-year clock, the flagship tier resists because the reasoning premium dominates it. The 31.5x reasoning premium is *why* [cost per task](./cost-per-task.md), not price per token, is the right unit for reasoning workloads — the premium is bought because accuracy lowers cost-per-correct-answer.

**The decline is bifurcated by tier — a benchmark datapoint (added 2026-07-24).** WorkBench Revisited (Styles & Miller, arXiv 2606.13715, 2026-06-10 — abstract-grade) re-runs a fixed workplace-agent benchmark two years on and observes the distribution's shape directly in *cost-of-performance* terms: **the cost of a performance level that was once proprietary-only collapsed (~100x since 2024, per the body), driven by open-weight models — while frontier serving costs stayed flat.** Over the same window best-agent task completion rose **43% (GPT-4, Mar-2024) → 98% (Claude Fable 5, mid-2026)** and unintended harmful actions fell **26% → 1.9%** (capability and safety moved together). The load-bearing reading for this page: **the price decline shows up in the *open-weight* rung, not the *frontier* rung** — the fast tails Epoch measures and the near-zero flagship fit Du measures are the *same* asymmetry seen from a benchmark, and it locates the commoditization where an open-weight model can reach a once-proprietary bar. It is a direct, if abstract-grade, datapoint on the open-weight-vs-frontier question the [frontier gaps](../frontier.md) track; the exact per-tier $ levels need the full text. (Absolute per-task levels for this workload class — $0.17–$6.55/task on STAGE-Claw — live on [cost per task](./cost-per-task.md).)

**⚠ The software-vs-hardware conflict (added 2026-07-23).** Du's headline econometric result is that **~103.7% of the cost reduction comes from total-factor productivity / software-architectural innovation and ~–0.9% from GPU hardware.** In plain terms: *hardware is not where the falling prices come from.* This cuts directly against the memory/hardware-scarcity-as-destiny framing of Matsuoka ([the depreciation conveyor](./depreciation-conveyor.md)) and Patel, in which hardware cost structure is the load-bearing variable. The two can be *partly* reconciled — memory scarcity may gate *supply/capacity* while software drives *price* — but **as stated they locate the cost lever in different layers, and this KB keeps both without resolving it.** Provenance to weigh: Du's decomposition is abstract-page only (the >100% TFP residual and Malmquist internals need the full PDF), a single-author preprint, not peer-reviewed. Logged as [frontier Conflict 7](../frontier.md#conflicts-on-the-record); the same conflict is recorded on the depreciation-conveyor page.

**This page measures price, not cost — and the two have come apart (added 2026-07-22).** Epoch's series is built from *published* prices, which is what makes it strong evidence. But a falling price tells you nothing about a falling cost: Patil's measurement (arXiv 2606.11690, 2026-06-10) finds effective cost on *identical H100 hardware* spanning **$0.21-$15.25 per million output tokens** (2.5-24x at 1-10 rps, up to 36.3x near idle) purely from offered request rate. A price series declining 50x/yr against a cost basis that can move 36x with traffic shape cannot, on its own, support a statement about margins. Keep this page's claims strictly on the price side; the cost side is [the cost-measurement problem](./cost-measurement-problem.md).

**Two mechanisms behind the decline, one of them newly named.** Matsuoka (arXiv 2607.07207, 2026-07-08) projects a **training-cost bifurcation** — a luxury tier at **$18-38B per frontier run by 2030** against a mass tier reaching previous-frontier parity via RL/distillation for a cost **falling toward $5M** (both modelled). If that holds, the price of *yesterday's* capability collapses not because anyone cut a price but because reproducing it becomes ~3-4 orders of magnitude cheaper, and frontier-capable open-weight models (Matsuoka cites **GLM-5.2**) put a floor-setting free option under every paid tier. That is a supply-side mechanism for exactly the fast tails Epoch measures.

**A caution now attaches to the demand side of any commodity argument built on this page.** Matsuoka additionally argues that **public token trackers overstate monetizable demand**, and that **all pre-Q2-2026 projections predate the industry's shift from token maximization to token minimization**. Epoch's price series is not a token-volume tracker and is not directly hit by that critique — but the demand figures this KB pairs with it (Goldman's 24x-by-2030, Gartner's 5-30x agentic per-task multipliers, both logged 2026-07-17) largely predate that line. Recorded as a live conflict: preprint critique vs bank/analyst forecast, both moderate, unresolved.

**The relationship to the rung.** This concept measures price *falling for fixed capability* (the buy-side view). Its sell-side mirror is [price-rung persistence](./price-rung-persistence.md): providers hold nominal per-token price points constant while capability *rises*. They are the same phenomenon seen from opposite ends — if the frontier rung stays at $5/$30 while the model improves, then the price to buy *yesterday's* capability must collapse (you buy it a rung down, or from open source). The Epoch distribution is what that collapse looks like when measured across tasks.

## Key Claims

- **Median inference price decline: 50x/year**, measured across six benchmarks and multiple thresholds. *Evidence: strong (methodology-documented benchmark study)* ([Epoch AI](../../raw/epoch-ai-llm-inference-price-trends.md))
- **Range: 9x to 900x per year** — the rate varies enormously by benchmark/threshold; there is no single "true" rate. *Evidence: strong* ([Epoch AI](../../raw/epoch-ai-llm-inference-price-trends.md))
- **GPT-4-level PhD-science (GPQA Diamond): 40x/year — below the all-benchmark median**, despite being the headline example most often quoted. *Evidence: strong* ([Epoch AI](../../raw/epoch-ai-llm-inference-price-trends.md))
- **Post-Jan-2024-only median: 200x/year** — the pace appears to have accelerated recently, but the fastest tails come entirely from a short window, so durability is unknown. *Evidence: moderate (short track record)* ([Epoch AI](../../raw/epoch-ai-llm-inference-price-trends.md))
- **The slow curves (9x end) are the moat signal** — where price falls slowly, capability is scarce and pricing power persists; where it falls fast (900x), commoditization is racing. *Evidence: moderate (interpretation of the Epoch spread; not an Epoch claim)*
- **Tier half-lives: economy ~1.10yr, mid ~1.55yr, flagship near-zero fit (R²=0.031) scrambled by a ~31.5x reasoning premium; ~600x aggregate decline.** *Evidence: moderate-strong — measured price half-lives, abstract-page read (26pp PDF unread)* ([Du / Tiered Super-Moore's](../../raw/tiered-super-moores-law-inference-price-evolution.md))
- **A May-2024 structural break shifts the decline from technology-driven to competition-driven (Chow F=5.74, p=0.005); HHI falls 4,558 → 2,086.** *Evidence: moderate — estimated econometrics, abstract-page* ([Du](../../raw/tiered-super-moores-law-inference-price-evolution.md))
- **CONFLICT (unresolved): ~103.7% of the price decline attributed to software/TFP, ~–0.9% to GPU hardware** — hardware is not where falling prices come from, cutting against Matsuoka/Patel memory-as-destiny. Memory may gate capacity while software drives price. *Evidence: estimated econometric decomposition, abstract-page, single-author preprint; the >100% residual needs the full text* ([Du](../../raw/tiered-super-moores-law-inference-price-evolution.md))
- **The decline is bifurcated open-weight vs frontier: on a fixed workplace-agent benchmark the cost of a once-proprietary performance level collapsed (~100x since 2024, open-weight-driven) while frontier serving costs stayed flat; completion 43%→98% and harmful actions 26%→1.9% (Mar-2024→mid-2026).** The price decline shows up in the open-weight rung, not the frontier rung. *Evidence: moderate — measured longitudinal benchmark, abstract-grade; per-tier $ levels need the full text* ([WorkBench Revisited](../../raw/workbench-revisited-workplace-agents-two-years-on.md))
- **A falling price series cannot support a margin claim on its own — effective cost on identical H100 hardware spans $0.21-$15.25/M output tokens (as of 2026-06) with traffic shape alone.** *Evidence: moderate (measured ratios; absolute level rests on an undisclosed $/GPU-hour input)* ([Patil](../../raw/beyond-per-token-pricing-concurrency-aware-cost-methodology.md))
- **Training-cost bifurcation is a supply-side mechanism for the fast tails: $18-38B per frontier run by 2030 vs mass-tier previous-frontier parity falling toward $5M.** *Evidence: weak-moderate (modelled projection to 2030)* ([Matsuoka](../../raw/memory-scarcity-open-models-ai-industry-restructuring-2026-2030.md))
- **Public token trackers overstate monetizable demand; pre-Q2-2026 demand projections predate the token-maximization → token-minimization shift.** *Evidence: moderate (preprint claim, magnitude not stated, not reproduced here); conflicts with the Goldman/Gartner demand figures logged 2026-07-17* ([Matsuoka](../../raw/memory-scarcity-open-models-ai-industry-restructuring-2026-2030.md))

## Benchmarks & Data

| Metric | Value | Source |
|---|---|---|
| Median decline (all benchmarks) | 50x / year | [Epoch AI](../../raw/epoch-ai-llm-inference-price-trends.md) |
| Range across benchmarks/thresholds | 9x – 900x / year | [Epoch AI](../../raw/epoch-ai-llm-inference-price-trends.md) |
| GPT-4-level PhD-science (GPQA Diamond) | 40x / year (below median) | [Epoch AI](../../raw/epoch-ai-llm-inference-price-trends.md) |
| Median, post-Jan-2024 models only | 200x / year | [Epoch AI](../../raw/epoch-ai-llm-inference-price-trends.md) |
| Economy-tier price half-life | ~1.10 year | [Du / Tiered Super-Moore's](../../raw/tiered-super-moores-law-inference-price-evolution.md) |
| Mid-tier price half-life | ~1.55 year | [Du](../../raw/tiered-super-moores-law-inference-price-evolution.md) |
| Flagship-tier fit | near-zero (R² = 0.031) | [Du](../../raw/tiered-super-moores-law-inference-price-evolution.md) |
| Reasoning per-token premium | ~31.5x | [Du](../../raw/tiered-super-moores-law-inference-price-evolution.md) |
| Aggregate decline (relative) | ~600x | [Du](../../raw/tiered-super-moores-law-inference-price-evolution.md) |
| Structural break (tech → competition) | May 2024 (Chow F=5.74, p=0.005) | [Du](../../raw/tiered-super-moores-law-inference-price-evolution.md) |
| Source of decline (estimated) | ~103.7% TFP / –0.9% GPU hardware | [Du](../../raw/tiered-super-moores-law-inference-price-evolution.md) |
| Cost of a fixed performance level, open-weight vs frontier | open-weight ~100x cheaper since 2024 / frontier flat | [WorkBench Revisited](../../raw/workbench-revisited-workplace-agents-two-years-on.md) |
| Best-agent workplace-task completion | 43% (Mar-2024) → 98% (mid-2026) | [WorkBench Revisited](../../raw/workbench-revisited-workplace-agents-two-years-on.md) |

Epoch benchmarks in the study: MMLU (general knowledge), GPQA Diamond (PhD science), MATH-500 / MATH Level 5 (math), HumanEval (coding), Chatbot Arena ELO (chat). Du's figures are relative/econometric — no absolute dated $/M-token levels appear in the abstract.

## Open Questions

- Does the post-2024 200x/year pace persist, or was 2024-2025 a one-off efficiency unlock (MoE, quantization, speculative decoding) that won't repeat?
- Which specific tasks sit at the slow (9x) end of the distribution — the defensible slices — and does that set shrink over time?
- Does the distribution's *shape* change (does the spread widen or narrow) as the frontier advances, or only its median?
- **Is there a matching *cost*-decline distribution?** This page has a price series and no cost series. Until one exists, "prices fell 50x/yr" and "margins compressed" are unconnected statements.
- **Does the Epoch series hold up under the token-minimization shift?** If providers and buyers both optimise for fewer tokens per task, a per-token price series and a per-task cost diverge — and per-task is the figure that matters commercially.
- Is Epoch's *quality-held* 50x/yr the right denominator for a commodity argument? The *blended-spot* price a buyer actually pays fell only ~3x/yr (2025→2026) — quality-per-dollar and unit price are different metrics that differ by >1 order of magnitude (see [frontier.md → Evidence Located](../frontier.md#evidence-located-2026-07-17-reading-desk-session)).

## Related Concepts

- [Token Pricing & the Inference-Margin Question](./token-pricing-margin-question.md) — the umbrella question this distribution feeds: fast price decline is the mechanism by which inference could become commodity infrastructure.
- [Price-Rung Persistence](./price-rung-persistence.md) — the sell-side mirror: providers hold the *rung* while capability rises, which is what forces the buy-side price collapse this page measures.
- [The Cost-Measurement Problem](./cost-measurement-problem.md) — this page is a *price* series; that page is why it cannot be read as a cost or margin series.
- [The Depreciation Conveyor & Vintage Economics](./depreciation-conveyor.md) — training-cost bifurcation and open-weight parity are the supply-side mechanisms behind the fastest tails here; also carries the same software-vs-hardware conflict from the cost-structure side.
- [Cost Per Task (Cost-of-Pass)](./cost-per-task.md) — the reasoning premium that scrambles the flagship tier is *why* cost-per-task, not price-per-token, is the right unit for reasoning workloads.

## Backlinks

*Pages that reference this concept:*
- [Token Pricing & the Inference-Margin Question](./token-pricing-margin-question.md)
- [Price-Rung Persistence](./price-rung-persistence.md)

## Changelog

- **2026-07-24** — Compiled WorkBench Revisited (2606.13715, abstract-grade). Added the bifurcated-decline datapoint: on a fixed workplace-agent benchmark the cost of a once-proprietary performance level collapsed ~100x since 2024 (open-weight-driven) while frontier serving costs stayed flat — the price decline lands in the open-weight rung, not the frontier rung (completion 43%→98%, harmful actions 26%→1.9%). New section, one Key Claim, two data rows; per-tier $ levels owed to the full text. Absolute per-task levels ($0.17–$6.55/task, STAGE-Claw) routed to [cost per task](./cost-per-task.md). No change to the Epoch/Du price series.
- **2026-07-23** — Compiled Tiered Super-Moore's Law (Du, 2603.28576, abstract-page). Added tier half-lives (economy 1.10yr / mid 1.55yr / flagship near-zero, R²=0.031, 31.5x reasoning premium), the May-2024 tech→competition structural break (HHI 4,558→2,086), and — prominently — the software-vs-hardware conflict (~103.7% TFP / –0.9% GPU hardware), recorded unresolved against Matsuoka/Patel memory-as-destiny. Linked the reasoning premium to cost-per-task. No change to the Epoch price series itself.
- **2026-07-22** — Compiled Patil (2606.11690) and Matsuoka (2607.07207) against this page: added the explicit price-is-not-cost boundary, training-cost bifurcation as a supply-side mechanism for the fast tails, and Matsuoka's demand-tracker critique as a recorded conflict with the Goldman/Gartner figures logged 2026-07-17.
- **2026-07-16** — Created from the Epoch AI source (already on the shelf, ingested 2026-07-14). Split out from the founding umbrella concept to make the price-decline *distribution* a first-class idea; paired with the new price-rung-persistence concept as its sell-side mirror.

## Sources

- epoch-ai-llm-inference-price-trends
- tiered-super-moores-law-inference-price-evolution
- beyond-per-token-pricing-concurrency-aware-cost-methodology
- memory-scarcity-open-models-ai-industry-restructuring-2026-2030
- workbench-revisited-workplace-agents-two-years-on

---

Cite as: MenFem Knowledge Base — https://menfem.com/kb/inference-economics/concepts/price-decline-distribution