The Markets Desk — Week 31 · Tue 28 Jul 2026

Markets

Token Price Index

PRIMARY-SOURCED FROM PROVIDER PRICING PAGES · OBSERVED 20 JUL 2026

If you’re shipping AI products on API inference, you’re paying token prices you can’t independently check — the invoice is usually the first place a team learns a model repriced. This is the hand-verified price of intelligence: every major model normalized to blended $/Mtok by capability tier, so you can see whether you’re overpaying before the invoice does. The hero metric is the open-vs-closed spread — how much cheaper you run open weights at the same capability bar.

6.21×

cheaper to run open weights at the Strong tier, blended. The frontier has no open peer — that’s the premium the whole economy hangs on.

Stack Watch

Free — the alert tier of Managed Inference

Get the alert when the rungs move

Every new model lands on the index the day its price verifies — GPT-5.6’s three rungs landed July 16. We re-sweep on every move and email you if the cheapest option at any capability tier changes. No account, no spam.

Open / Closed Spread by Tier

Blended $/Mtok · 3:1 input:output profile
FrontierBEST AVAILABLECLOSED $4.50NO OPEN PEERNO SPREAD
StrongNEAR-FRONTIERCLOSED $3.38OPEN $0.5446.21×OPEN CHEAPER
MidWORKHORSECLOSED $0.563OPEN $0.1753.21×OPEN CHEAPER
SmallCHEAP & FASTCLOSED $0.175OPEN $0.062.92×OPEN CHEAPER

All Tracked Models

$/Mtok · read straight from pricing pages
ModelProviderTierWeightsInputOutputBlended
Gemini 3.1 ProgoogleFRONTIERCLOSED2.0012.004.50
GPT-5.6 SolopenaiFRONTIERCLOSED5.0030.0011.25
Claude Opus 4.8anthropicFRONTIERCLOSED5.0025.0010.00
GPT-5.5openaiFRONTIERCLOSED5.0030.0011.25
Claude Fable 5anthropicFRONTIERCLOSED10.0050.0020.00
DeepSeek V4 ProdeepseekSTRONGOPEN0.430.870.54
Gemini 3.5 FlashgoogleSTRONGCLOSED1.509.003.38
Claude Sonnet 5anthropicSTRONGCLOSED2.0010.004.00
GPT-5.6 TerraopenaiSTRONGCLOSED2.5015.005.63
GPT-5.4openaiSTRONGCLOSED2.5015.005.63
Llama 3.1 405BmetaSTRONGOPEN3.503.503.50
DeepSeek V4 FlashdeepseekMIDOPEN0.140.280.18
Gemini 3.1 Flash-LitegoogleMIDCLOSED0.251.500.56
Llama 4 MaverickmetaMIDOPEN0.270.850.42
Gemini 2.5 FlashgoogleMIDCLOSED0.302.500.85
GPT-5.4-miniopenaiMIDCLOSED0.754.501.69
Llama 3.3 70BmetaMIDOPEN0.880.880.88
GPT-5.6 LunaopenaiMIDCLOSED1.006.002.25
Llama 3.2 3BmetaSMALLOPEN0.060.060.06
Llama 4 ScoutmetaSMALLOPEN0.080.300.14
Gemini 2.5 Flash-LitegoogleSMALLCLOSED0.100.400.18
GPT-5.4-nanoopenaiSMALLCLOSED0.201.250.46
Claude Haiku 4.5anthropicSMALLCLOSED1.005.002.00

Estimate Your Bill

10M input / 2M output per month, Strong tier
closed $33.00/moopen (hosted) $6.09/moyou save 82%
Cheapest closed$337.50/moGemini 3.5 Flash
Cheapest open$54.38/moDeepSeek V4 Pro

$283.13/mo

saved by running open — 84% cheaper

We do this for real — token-cost audit →

Or Have It Handled

Managed Inference — the standing version of the auditSee Managed Inference
£500/mo

Your bill won’t stay optimized — providers repriced or changed terms 32 times last quarter. Managed Inference re-benchmarks your actual model mix against this index every month and sends the switch list, with an alert any week your number moves. Built for teams spending £5k+/month on inference.

Price Is Half the Picture

The rest of the sovereignty desk
Methodology

Every price is read from the provider’s own public pricing page and dated — never a reseller. Blended $/Mtok weights input:output by workload profile; tiers are assigned by benchmark band and re-checked at each model release. One trap folded in: newer tokenizers emit more tokens for the same text, so compare price per task, not per raw token. Last observed 20 JUL 2026, re-swept monthly.