Epoch AI

organization
serving-capacitymodellingprefill-decodecompute-crunch

Epoch AI

Type: Organization (research group)

Epoch supplies this rung's only supply-side instrument. Where every other source here optimises one box, Epoch models how many tokens the world's chips can serve, and whether that keeps pace with demand.

The distinguishing feature is that it is modelled but calibrated, not projected. A first-principles prefill/decode compute-and-bandwidth model is fitted against 111 measured SemiAnalysis InferenceX runs (Kimi K2.5), yielding hardware efficiency parameters that can be reused: compute efficiency 65%, bandwidth efficiency 30%, 5ms per step.

Key contributions

  • Per-chip serving throughput of ~400k tok/s on GB200 NVL72 (Epoch)
  • Global inference capacity growing ~3.4x/yr against demand growing ~10x/yr — the supply-demand gap behind sticky token pricing (Epoch)
  • Publishes reusable fitted efficiency parameters, so the model can be re-run against other hardware (Epoch)

What it deliberately does not say

No per-token dollar cost. The piece is denominated in tokens/sec of capacity, not dollars. Pairing it with the DigitalOcean measurement is what produces a cost view, and that pairing is an inference this rung has not yet done carefully — the two use different hardware, different models and different dates.

Mentioned in

Related concepts