Rung 10 / What a task actually costs

Inference & Token-Pricing Economics

The economics of serving models — what a token costs to make, and what it sells for.

14
Sources
8
Concepts
4
Entities
A small precision balance with a copper weight on one pan, on paper.

In scope: token pricing trends, inference gross margins, the AI-lab P&L question, and how memory, optics and foundry cost structure feed through to the price of a token.

A brass-and-paper gauge with a copper bezel: token price in, task price out.
Token pricesTCOMarginLock-inRugs
Report only3Show all →
Sources compiled for this topic
TypeSourcePublished
REPORTLLM Inference Prices Have Fallen Rapidly but Unequally Across Tasks
Ben Cottier, Ben Snodin, David Owen, Tom Adamczewski · Epoch AI

Median LLM inference price decline 50x/yr across 6 benchmarks (range 9x-900x/yr); GPT-4-level PhD-science (GPQA Diamond) milestone specifically 40x/yr, below the all-benchmark median; post-Jan-2024-only median rises to 200x/yr.

2025-03-12
REPORTToken Price Index — GPT-5.6 holds the prior-generation price rungs
MenFem Token Price Index (own dataset) · MenFem — research/inference/token-prices.csv

GPT-5.6 (2026-07-16) holds the exact prior-generation per-token price at two of three tiers: Sol (frontier) $5/$30 = GPT-5.5's rung, Terra (strong) $2.50/$15 = GPT-5.4's rung; Luna (mid) $1/$6 is a new point. Capability up, price rungs held.

2026-07-16
REPORTLLM Inference Handbook
Modular · Modular

The measurement vocabulary for inference cost: TPOT = (E2EL - TTFT)/(output tokens - 1), ITL, RPS, and GOODPUT (requests/sec meeting SLOs — the metric that stops a cost-per-token figure flattering a system that is fast and unusable). Plus GPU-memory and KV-cache calculators.

Inference & Token-Pricing Economics | Knowledge Base | MenFem