LLM Inference Prices Have Fallen Rapidly but Unequally Across Tasks
Median LLM inference price decline 50x/yr across 6 benchmarks (range 9x-900x/yr); GPT-4-level PhD-science (GPQA Diamond) milestone specifically 40x/yr, below the all-benchmark median; post-Jan-2024-only median rises to 200x/yr.
LLM Inference Prices Have Fallen Rapidly but Unequally Across Tasks
Primary-fetched from epoch.ai. Data insight originally published 2025-03-12; ingested to this KB 2026-07-14 as a foundational reference for this week's token-pricing coverage — an older publish date than the other two sources this week, but newly brought into the KB and directly load-bearing for the "40x is below average" claim in this week's content plan.
Abstract
Epoch AI measures how quickly the price to reach a fixed LLM performance milestone has fallen over time, combining LLM API pricing data with benchmark scores from Artificial Analysis and Epoch AI's own database. For each of six benchmarks, they identify the lowest-priced model that meets or exceeds a given score, then fit a log-linear regression to price-over-time to measure the annualized decline rate. The same method is applied across multiple performance thresholds per benchmark to show how much the decline rate varies by task.
Key Contributions
- Median decline: 50x per year, measured across six benchmarks and multiple performance thresholds.
- Range: 9x to 900x per year — the rate of price decline varies enormously depending on which benchmark/performance milestone is used.
- GPT-4-level PhD-science milestone (GPQA Diamond): 40x per year — notably below the all-benchmark median of 50x/year, despite being the headline example most often quoted.
- Recency effect: restricting the sample to models released after January 2024 raises the median decline rate to 200x per year — the pace of price decline has accelerated recently, and the fastest trends (approaching 900x/year) only emerge post-January-2024, so their persistence is uncertain.
Methodology
- Benchmarks used: MMLU (general knowledge), GPQA Diamond (PhD-level science), MATH-500 and MATH Level 5 (mathematics), HumanEval (coding), Chatbot Arena ELO (competitive chat performance).
- For each benchmark and performance threshold, identify the cheapest model (by API price) that clears the threshold at each point in time.
- Fit log-linear regression to the resulting price series to estimate the annualized rate of decline.
Results
| Metric | Value |
|---|---|
| Median price decline (all benchmarks) | 50x / year |
| Range across benchmarks/thresholds | 9x – 900x / year |
| GPT-4-level PhD-science (GPQA Diamond) specific rate | 40x / year (below median) |
| Median decline, post-Jan-2024 models only | 200x / year |
Limitations
- The fastest-declining trends (up to 900x/year) are drawn entirely from data after January 2024 — a short window, so whether this pace persists is unknown.
- Benchmark/threshold choice materially changes the headline number — there is no single "true" rate of inference price decline, only a distribution.
Full Content
See Results table above. This is a data-insight page (chart + methodology note), not a full paper; the key figures are captured in full above.
Source: LLM Inference Prices Have Fallen Rapidly but Unequally Across Tasks by Ben Cottier, Ben Snodin, David Owen, Tom Adamczewski, Epoch AI.