
In scope: token pricing trends, inference gross margins, the AI-lab P&L question, and how memory, optics and foundry cost structure feed through to the price of a token.
| Type | Source | Published |
|---|---|---|
| REPORT | LLM Inference Prices Have Fallen Rapidly but Unequally Across Tasks Ben Cottier, Ben Snodin, David Owen, Tom Adamczewski · Epoch AI Median LLM inference price decline 50x/yr across 6 benchmarks (range 9x-900x/yr); GPT-4-level PhD-science (GPQA Diamond) milestone specifically 40x/yr, below the all-benchmark median; post-Jan-2024-only median rises to 200x/yr. | 2025-03-12 |
| REPORT | Token Price Index — GPT-5.6 holds the prior-generation price rungs MenFem Token Price Index (own dataset) · MenFem — research/inference/token-prices.csv GPT-5.6 (2026-07-16) holds the exact prior-generation per-token price at two of three tiers: Sol (frontier) $5/$30 = GPT-5.5's rung, Terra (strong) $2.50/$15 = GPT-5.4's rung; Luna (mid) $1/$6 is a new point. Capability up, price rungs held. | 2026-07-16 |
| REPORT | LLM Inference Handbook Modular · Modular The measurement vocabulary for inference cost: TPOT = (E2EL - TTFT)/(output tokens - 1), ITL, RPS, and GOODPUT (requests/sec meeting SLOs — the metric that stops a cost-per-token figure flattering a system that is fast and unusable). Plus GPU-memory and KV-cache calculators. | — |