PAPER2026-03-30 · Not stated on the paper (single author, Mingdeng Du); cs.CE, ACM J.4/K.4.4 · arXiv 2603.28576

Tiered Super-Moore's Law: Price Evolution, Production Frontiers, and Market Competition in LLM Inference Services

Mingdeng Du
Compiled notes
What it moved

Decomposes the ~600x token-price decline into TIER-SPECIFIC half-lives (economy 1.10yr, mid 1.55yr, flagship near-zero fit R2=0.031 from 31.5x reasoning premium); dates a May-2024 structural break (Chow F=5.74, p=0.005) from technology- to competition-driven decline; attributes ~103.7% of cost reduction to TFP vs -0.9% from GPU hardware — a direct empirical counter to the memory/hardware-centric cost story. MEASURED (OpenRouter 318 + Epoch 3,237 models + 62 milestones); econometric decomposition ESTIMATED.

Tiered Super-Moore's Law

What this is

The most rigorous token-price-trend + market-structure paper the KB has seen (2026-03-30 preprint). It takes the same phenomenon Epoch's price-trends dataset documents — prices falling fast and unequally — and adds three things Epoch does not: tier-specific decline rates, a dated structural break in the cause of the decline, and a growth-accounting decomposition of what actually drives cost reduction. Built on integrated datasets (OpenRouter 318 models, Epoch AI 3,237 models, 62 validated milestones 2020–2026).

Key findings

Price decline is tiered (measured):

  • Economy tier: ~1.10-year price half-life
  • Mid tier: ~1.55-year half-life
  • Flagship tier: near-zero exponential fit (R² = 0.031) — flagship prices don't fall cleanly because reasoning models carry a ~31.5x price premium over non-reasoning, scrambling the trend
  • Aggregate: ~600-fold price decline (relative; no absolute $/M-with-date in the abstract)

Cause of the decline shifted (estimated):

  • Chow test dates a structural break to May 2024 (F = 5.74, p = 0.005): before, decline was technology-driven; after, competition-driven.
  • Malmquist productivity index peaked at 4.11 in 2024 Q1–Q4; technological-frontier shift (TC = 4.13) dominates.
  • Market concentration fell: HHI 4,558 → 2,086 over three years (moderately concentrated → competitive).

What drives cost reduction (estimated, and the headline claim):

  • Total-factor-productivity residuals account for ~103.7% of cost reduction; GPU hardware contributes ~-0.9%.
  • I.e. software / architectural innovation, not hardware, drives the price collapse.
  • A US–China ~63-fold training-cost gap is attributed to architectural innovation, not factor-price differences (p = 0.228).

Why it matters here

  • Sharpens the price-decline-distribution concept — Epoch's 9x–900x spread is here resolved into clean tier half-lives (economy 1.10yr, mid 1.55yr) plus a flagship tier that resists fitting because reasoning premiums dominate. This is a more structured version of the "slow tails are the moat signal" reading already in the KB.
  • Direct conflict with the hardware/memory cost story — the ~103.7%-TFP / -0.9%-hardware decomposition says software efficiency, not silicon, is where cost reduction comes from. That cuts hard against Matsuoka's memory-scarcity-as-destiny framing and Patel's memory-shortage thesis: if hardware contributes ~-0.9% to the price decline, a memory crunch may throttle supply without ever having been the source of falling prices. Keep both on the record.
  • Competition-driven-since-May-2024 — supports Evans' "compresses toward marginal cost as competition bites" pole over the "rotating-landlord oligopoly" pole, and puts a date on when the mechanism changed.

Full-text close read — added 2026-09-11

The July entry was built from the abstract and flagged the ">100% TFP" residual and the Malmquist internals as needing the full text. Read now (arXiv HTML v1, 23 pages, 12 figures, 6 tables).

Dataset construction. Three streams, integrated: (1) OpenRouter, 318 models, a single pricing snapshot taken 2026-03-28, normalised to $/M tokens; (2) the Epoch AI panel, 3,237 models, 2020–2026, carrying training cost, FLOP, parameters and region of origin; (3) 62 cross-validated milestone observations hand-compiled from vendor pages and industry sources. The time series — i.e. everything the half-lives are fitted to — rests on the 62 milestones. The 318 and 3,237 figures are cross-sections, not history, and the paper's headline model counts therefore overstate the depth of the trend evidence. That is the single most important thing the abstract does not tell you.

Tier definitions are price thresholds, not capability classes. Flagship >$5/M input (GPT-4, Claude 3 Opus, o1/o3, GPT-5-pro); Mid $0.5–$5/M (GPT-4o, Claude Sonnet, Gemini Pro); Economy <$0.5/M (GPT-4o-mini, DeepSeek-V3, Gemini Flash). Because the bucket boundary is a price, a model whose price falls leaves its tier, which mechanically flatters the surviving members of the higher tier. Treat the tier half-lives as describing price bands over time, not cohorts of models.

Half-lives, with their fits. P(t) = P₀·e^(−λt):

TierHalf-lifePaper's gloss
Economy1.10 yr0.192"82% faster than Moore's Law"
Mid1.55 yr0.307"29% faster"
Flagship3.59 yr0.031exponential decay "inapplicable"

Note what the July entry could not see: the economy fit is R² = 0.192. The strongest half-life in the paper explains under a fifth of the variance. The tiering is real and the direction is robust; the precision of "1.10 years" is not, and it should never be quoted to two decimals as though it were a measured constant.

TFP decomposition — the >100% residual explained. Growth accounting attributes ~103.7% of cost reduction to total-factor-productivity residuals and −0.9% to GPU hardware. The residual exceeds 100% because the hardware term is negative: measured GPU cost per unit did not help over the window, so software and architecture had to more than cover the whole decline. This is a residual, i.e. everything the model does not otherwise name — it is the size of the unexplained part, not a measurement of software efficiency.

Chow break. 2024-05-01 is the strongest break: F = 5.736, p = 0.005, on 16 pre-break and 44 post-break observations. Sixteen observations before the break is thin, and the paper's technology-driven → competition-driven story rests on that side of the split.

DEA / Malmquist. Cross-sectional efficiency over the 318 models has mean 0.087 — "the typical model is priced at 11.5× its frontier-equivalent," which is a striking statement about dispersion in the market, and is arguably more useful to the KB than the trend. Malmquist peaks at MPI = 4.107 in 2024Q1–Q4, driven by technological-frontier shift TC = 4.13.

HHI. 4,558 (2023Q1) → 2,086 (2026Q1), a 54.2% fall, highly → moderately concentrated, with OpenAI's share falling 65% → 34%. Shares are proxy-estimated: the paper has no transaction volume data, which it concedes.

Reasoning premium is unstable. The 31.5× average hides quarterly values of 83.3× (2024Q3), 18.8× (2024Q4), 0.51× (2025Q1), 23.4× (2025Q2). A quarter at 0.51× means reasoning models were cheaper than non-reasoning that quarter. The paper reads the premium as pricing power rather than cost structure, pointing to DeepSeek-R1 entering at $0.55/M. Cite the premium as a range, never as "31.5×".

Training-cost elasticity. ln(inference price) = α + 0.432·ln(training cost) + controls; SE 0.204, p = 0.050, R² = 0.219, n = 18. Eighteen observations at exactly p = 0.050 is a result to report, not to lean on.

The 63× US–China training gap. GPT-4.5 at $340M vs DeepSeek-V3 at $5.4M. $/FLOP is statistically indistinguishable (p = 0.228; China 2.41×10⁻¹⁸, US 2.25×10⁻¹⁸) — so the gap is architecture, not factor prices: 3.3×10²⁴ FLOP for the MoE against 3.8×10²⁶.

Limitations

Single-author preprint, no stated affiliation. The authors concede three: 62 milestones is a small sample for the trend work; there is no transaction volume data, so all market-share and HHI figures are proxy-based; and a six-year window in a fast-moving market means "structural relationships identified here may shift." Two more the full read adds: the fits are weak (economy R² = 0.192, flagship 0.031) and the tier boundaries are prices, so models migrate between tiers as they cheapen. Relative metrics only — no absolute dated $/M-token level series anywhere, so this cannot be cross-checked against, or used to rebuild, the retired Token Price Index. Price series ends 2026-03-28 and predates the July–August 2026 cuts: citable as a rate with an as-of date, never as today's price.


Source: Tiered Super-Moore's Law by Mingdeng Du, arXiv 2603.28576, 2026-03-30

Related in the base
Tiered Super-Moore's Law: Price Evolution, Production Frontiers, and Market Competition in LLM Inference Services | Knowledge Base | MenFem