Inference & Token-Pricing Economics — Research Frontier
Research Frontier: Inference & Token-Pricing Economics
What's genuinely new and where the field is heading.
State of the topic as of the 2026-07-23 compile: 13 sources, 7 concepts, 4 entities — wiki now recompiled against all six newest sources. The 2026-07-23 discovery sweep ingested six genuinely-new, on-thesis sources and this compile filed them: it partially closed the topic's defining gap — a measured cost series now exists. DigitalOcean measured a 70B open model at $0.45–$20.32/M output tokens on one H200 purely across batch size (2026-07-08), independently reproducing Patil's utilization thesis in dollars; Epoch modelled serving capacity calibrated to 111 real SemiAnalysis InferenceX runs; and three arXiv economics papers (AI Tokenomics, Tiered Super-Moore's Law, Cost-of-Pass) plus an independent-provider price matrix filled the theory, price-trend-decomposition, cost-per-task, and independent-provider layers. The honest headline has moved from "a price series and no cost series" to "a single measured cost point for one open model, still no cost series for any frontier/proprietary model, no cost time-series, and still zero peer-reviewed papers." Two new pages were created: the cost-per-task concept and the independent inference providers entity. All six sources are now compiled; staleSources = 0.
Sweep also surfaced a new first-order conflict, now recorded on the concept pages: Tiered Super-Moore's Law attributes ~103.7% of the price decline to software/total-factor-productivity and –0.9% to GPU hardware, which cuts directly against the memory/hardware-scarcity-as-destiny framing of Matsuoka and Patel. Logged in Conflicts below (item 7) and prominently on the price-decline distribution and the depreciation conveyor; unadjudicated.
Update — 2026-07-24 compile (three sources added; now 16 sources, 8 concepts). Three surfaced-but-unread sources are now compiled at abstract grade, and they close or partly-close three of the gaps below. STAGE-Claw (arXiv 2606.10394) supplies the topic's first absolute dated cost-per-task levels — ~$0.17–$6.55/task across 11 frontier models (search-surfaced from its cost table, full-text confirmation owed); WorkBench Revisited (arXiv 2606.13715) adds the two-year trajectory and the open-weight-cheap / frontier-flat cost bifurcation (completion 43%→98%, harmful actions 26%→1.9%); and Photons = Tokens (arXiv 2603.06630) supplies the missing energy→token bridge (326 TWh → ≈6.5×10¹⁷ tokens/yr → ≈225,000 tokens/person/day), filed as a new concept — the energy cost of tokens. All three are single-preprint, abstract-grade, not peer-reviewed. Gap retags below carry [CLOSED 2026-07-24] / [PARTLY CLOSED 2026-07-24].
Active Frontiers
1. The cost denominator — measured for one open model, unmeasured for the frontier
Status: Partly closed (2026-07-23) — the topic's first measured cost floor exists, but only for one open model; still the most important frontier Key sources: DigitalOcean measured H200, Epoch serving-capacity model, Patil / arXiv 2606.11690, Matsuoka / arXiv 2607.07207 Key concept: The Cost-Measurement Problem
A measured cost-to-serve series now exists: DigitalOcean's single H200 (Llama-3.3-70B FP8, vLLM 0.24.0) runs $20.32/M at batch=1 → $0.45/M at batch=128 (~44x from traffic shape alone; 2026-07-08), independently reproducing Patil's utilization thesis in dollars on different silicon. Epoch's InferenceX-calibrated model supplies the physics under the floor (prefill compute-bound / decode bandwidth-bound; 65%/30% fitted efficiency). Patil's underlying finding stands: identical H100 yields $0.21-$15.25/M (up to 36.3x near idle) from offered request rate alone (× $6.98/GPU-hr, Azure H100 NVL list — now anchored). Matsuoka argues the unit is wrong too, proposing $/PB of bandwidth delivered as model-agnostic for bandwidth-bound decode. But the floor that exists is for an open 70B model; the frontier cost floor is still missing.
Open problems:
- [SHARPEST] No measured cost floor for any frontier/proprietary model — the measured floor is an open 70B; the Evans-vs-Patel debate is about GPT-5.x / Claude / Gemini serving, still unmeasured.
- The measured floor is a point, not a series: one SKU, one open model, one framework, one date. No cost time-series; no reproduction on Blackwell GB200/GB300, TPU, Trainium, custom ASIC; no long-context/reasoning-workload cost.
- Nothing reconciles per-token pricing with $/PB. The entire rest of this topic is denominated in tokens.
- $/PB is claimed model-agnostic for bandwidth-bound decode only; prefill, long-context and reasoning-heavy workloads are the growing share and are unaddressed.
2. Does inference become commodity infrastructure?
Status: Steady — now with a quantitative prior, but the prior is a coin flip Key sources: Ben Evans, Epoch AI, Dylan Patel, Matsuoka Key concept: Token Pricing & the Inference-Margin Question
Matsuoka is the first quantitative model to engage the question directly, and it prices the two poles at 25% each (Commoditization Crash 25% / Rotating Landlord Oligopoly 25%; Jevons Absorption 20%, System-Layer Re-differentiation 18%, Geopolitical Bifurcation 12%). Author-assigned priors from a single-author preprint with no elicitation method — weak evidence about the world, but a clear statement that the question is not currently decidable. The Evans-vs-Patel margin gap (40-50% vs 70%+) is meanwhile smaller than the cost range Patil measures on one SKU, so it is no longer a well-posed disagreement.
Open problems:
- No audited primary financial disclosure for any frontier lab. Anthropic's confidential draft S-1 (filed ~2026-06-01) is a date-certain resolution path — the only one in sight.
- "Does inference commoditize?" and "does the incumbent cost advantage persist?" may be different questions. Matsuoka's conveyor allows both: a commoditizing market in which incumbents remain permanently cheaper than entrants.
- Whether the post-2024 200x/yr price-decline pace (Epoch) is durable or a short-lived efficiency unlock.
3. The rung vs the curve — two mirror-image price mechanics
Status: Steady; the rung has been upgraded to a solvency instrument Key sources: Epoch AI, Token Price Index, Matsuoka Key concepts: The Price-Decline Distribution, Price-Rung Persistence
Buy-side: the price of fixed capability falls as a distribution, not a rate — Epoch's median 50x/yr masks a 9x-900x spread, and the slow tails are the moat signal. Sell-side: providers hold nominal per-token rungs while capability rises — GPT-5.6 (2026-07-16, own-dataset primary) held two of three prior-generation rungs to the cent. New this compile: Matsuoka names "sticky premium pricing" as one of two joint conditions for buildout solvency, which makes the Token Price Index an instrument on a modelled solvency variable — a frontier rung cut becomes a leading indicator, visible the day it happens, months before financials.
Open problems:
- Is rung-holding an OpenAI convention or industry-wide? Needs Anthropic / Google / DeepSeek generation-over-generation rows.
- How wide is the gap between list rungs and realised prices? Enterprise discounts, committed-use contracts and batch/cached tiers are invisible to the instrument, and could be cutting the effective rung while the published one holds.
- Does rung-holding survive frontier-capable open-weight models setting a free floor one rung down?
4. Vintage and depreciation economics
Status: Early stage — one modelled source, high consequence if true Key source: Matsuoka Key concept: The Depreciation Conveyor & Vintage Economics
The claim that a depreciation conveyor keeps the entrant-incumbent cost gap open indefinitely (3.2x 2026 → 1.9x 2027 → re-widening to 3-4x by 2029-30, all modelled), and that capacity is not fungible by vintage — 2026 and 2028-29 each fatally exposed to one pricing regime, only 2027 robust. If this holds, "who owns AI compute" is a weaker question than "which year's compute do they own, and at what memory price."
Open problems:
- Which pricing regime kills the 2026 vintage and which kills 2028-29 — the abstract states exposure without naming the regime. This is the single highest-value unread detail in the topic.
- The amortisation schedule and hardware-price path generating the 3.2x → 1.9x → 3-4x curve are undisclosed; the mechanism is plausible, the numbers are unfalsifiable without them.
- Hyperscaler useful-life assumptions are disclosed in 10-Ks and have been revised. That is a primary-source test of the conveyor's core premise and it has not been run.
5. Macro pass-through of inference cost
Status: Early stage / speculative — one unreplicated preprint Key source: ICPC preprint Key concept: Inference Cost as a Macro Input
Held at low confidence deliberately. If inference cost really passes through to the general price level at near-unit elasticity, a token price index becomes a macro instrument rather than an industry one — a large enough consequence to track at low confidence rather than discard. Three named reasons not to believe it yet: an implausibly clean R² = 0.998 over overlapping rolling windows, an undisclosed inference-cost series construction (now known: a GPU-list + IEA-electricity composite, utilization-naive), and lambda-bar calibrated at an implausible 0.18 with no stated source.
6. Cost per task — the right unit, first absolute levels now in
Status: Absolute dated levels now in (2026-07-24) — STAGE-Claw supplies ~$0.17–$6.55/task across 11 frontier models; a frontier/proprietary per-task time-series is still missing Key sources: Cost-of-Pass / arXiv 2504.13359, AI Tokenomics / arXiv 2606.24616, STAGE-Claw / arXiv 2606.10394, WorkBench Revisited / arXiv 2606.13715 Key concept: Cost Per Task (Cost-of-Pass)
The topic has repeatedly said the commercially meaningful unit is cost per task, not price per token, and now holds the framework for it. Cost-of-Pass defines it (per-token cost ÷ accuracy = the cost of a correct answer; frontier cost-of-pass = min across models or human experts) and measures it "roughly halving every few months" for complex quantitative tasks, driven by model-level not inference-time innovation. AI Tokenomics supplies the theory spine — token expenditure ≠ economic value, so per-token pricing prices the input not the output. Together they explain the reasoning-premium paradox (a 31.5x per-token premium is bought because accuracy lowers cost-per-pass) and echo Tiered Super-Moore's "software not hardware" decomposition.
Open problems:
No absolute dated cost-per-task ($/correct-answer) for a real 2026 agentic/reasoning workload.[CLOSED 2026-07-24] STAGE-Claw supplies it (~$0.17–$6.55/task across 11 frontier models, search-surfaced, full-text confirmation owed); WorkBench adds the two-year trajectory. Residual: a frontier/proprietary per-task cost time-series, and coverage beyond personal-computing / workplace agents (coding, research, long-context).- Cost-of-pass needs a per-attempt cost — the same missing, utilization-dependent denominator from Frontier 1. The unit is right; the input to it is still unmeasured for frontier models.
- Hidden reasoning tokens (unbilled) are named as a value-side cost with no measurement method.
Recent Breakthroughs
| Date | Breakthrough | By | Source |
|---|---|---|---|
| 2026-06-09 | First absolute dated cost-per-task levels in the KB — a real 2026 agentic workload costs ~$0.17–$6.55/task across 11 frontier models (~38x) on the 40-task STAGE-Claw state-based benchmark (search-surfaced cost table; full-text confirmation owed) | Liang et al. (arXiv 2606.10394) | preprint |
| 2026-06-10 | Workplace-agent benchmark two years on: completion 43%→98%, harmful actions 26%→1.9%; open-weight collapsed the cost of a once-proprietary performance level (~100x) while frontier serving costs stayed flat | Styles & Miller (arXiv 2606.13715) | preprint |
| 2026-02-23 | Token priced as a thermodynamic unit (Landauer/Shannon); 326 TWh 2028 US AI energy → ≈6.5×10¹⁷ tokens/yr ≈ 225,000 tokens/person/day, >3 OOM above mid-2024 — energy not the near-term binding constraint on token volume, direction is | Litowitz, Polson & Sokolov (arXiv 2603.06630) | preprint |
| 2026-07-08 | First MEASURED $/token cost series in the KB — one H200 = $0.45–$20.32/M output tokens (~44x) across batch size alone; own-GPU-vs-serverless crossover at ~73% utilization | DigitalOcean (Baranwal) | measured |
| 2026-05-25 | Prefill/decode serving-capacity model calibrated to 111 real SemiAnalysis InferenceX runs (bandwidth eff ~30%); global capacity growing ~3.4x/yr vs demand ~10x/yr → compute crunch | Epoch AI (Emberson & Sevilla) | report |
| 2026-03-30 | ~600x price decline decomposed into tier half-lives (economy 1.10yr / mid 1.55yr / flagship unfittable, 31.5x reasoning premium); May-2024 tech→competition break; GPU hardware only –0.9% of the cost fall | Mingdeng Du (arXiv 2603.28576) | preprint |
| 2026-06-10 | "AI tokenomics" formalized: token expenditure ≠ economic value; per-token pricing prices the input, not the output; hidden reasoning tokens named | Quanyan Zhu (arXiv 2606.24616) | preprint |
| 2026-02 (rev) | Cost-per-task framework "cost-of-pass" = per-token cost ÷ accuracy; frontier cost-of-pass on hard quant tasks halving every few months; model-level (not inference-time) innovation drives it | Erol et al. (arXiv 2504.13359) | preprint |
| 2026-04-24 | Independent providers observed: 6x same-model list-price spread (Together $0.65 batch → Cerebras $4.20); cheapest list sits at the measured cost floor → thin margin at batch tier | Digital Applied | analysis |
| 2026-07-08 | Inference economics reformulated in $/PB of bandwidth delivered; "depreciation conveyor" named; solvency priced as a two-condition corridor; commodity-vs-oligopoly tied at 25%/25% | Satoshi Matsuoka (arXiv 2607.07207) | preprint |
| 2026-06-10 | Identical H100 measured at $0.21-$15.25/M output tokens (up to 36.3x) from traffic shape alone; every surveyed cost calculator shown utilization-naive; vllm-cost-meter released | Chitral Patil (arXiv 2606.11690) | preprint |
| 2026-07-16 | GPT-5.6 holds two of three prior-generation price rungs to the cent (Sol $5/$30, Terra $2.50/$15); Luna a new mid point — capability up, rung flat | OpenAI (observed via MenFem Token Price Index) | TPI |
| 2026-07-10 | Attributed claim: Anthropic FCF-positive, $50B+ ARR, 70%+ gross margin | Dylan Patel (Podcast Alpha) | podcast |
| 2026-07-09 | Four-question framework for token-pricing outcomes (adoption breadth / frontier progress / competition / value capture) | Ben Evans | essay |
| 2026-05-19 | First model of inference cost as a first-order macro marginal-cost input (Inference-Cost Phillips Curve) | arXiv 2605.20281 | preprint |
Conflicts on the Record
Recorded rather than adjudicated. Where sources disagree, this topic keeps both.
- Unit conflict — per-token vs $/PB. Matsuoka argues $/PB of bandwidth delivered is the model-agnostic unit for bandwidth-bound decode; Epoch, the Token Price Index, Evans and Patel are all denominated per token. Nothing in this KB reconciles them. (preprint vs established practice; unresolved)
- Method conflict — margins without denominators. Patil's 36.3x utilization-driven cost spread does not numerically contradict Evans' 40-50% or Patel's 70%+, but it makes both unfalsifiable as stated. Both should be read as unstated-utilization estimates. (moderate vs moderate/weak)
- Demand conflict — trackers overstate vs bank forecasts. Matsuoka: public token trackers overstate monetizable demand, and all pre-Q2-2026 projections predate the shift from token maximization to token minimization. Against: Goldman's 24x-by-2030 token-consumption forecast and Gartner's 5-30x agentic per-task multipliers, both logged 2026-07-17 below and both largely predating that line. (preprint critique vs bank/analyst forecast; both moderate; unresolved)
- Rate conflict — which price decline is real. Epoch's quality-held 50x/yr vs the blended-spot ~3x/yr a buyer actually paid (Q1'25→Q1'26). They measure different things and differ by more than an order of magnitude; the gap between them is the rung-hold mechanic. Already logged 2026-07-17; unchanged.
- Margin conflict — Patel's 70% vs the reported present-period ~50%. Reported (unaudited) investor-deck lineage puts Anthropic's present gross margin near 50% with 77% projected only by 2028. Patel's "70%+ now" is not corroborated as a current figure. Logged 2026-07-17; unchanged.
- Internal tension in Matsuoka. He prices solvency on token-demand growth and argues the token trackers that measure demand are inflated. Whether the scenarios already internalise his own critique is not visible from the abstract.
- Source-of-cost-reduction conflict — software vs hardware (NEW 2026-07-23). Tiered Super-Moore's Law (Du, 2603.28576) attributes ~103.7% of the token-price decline to total-factor productivity / software-architectural innovation and –0.9% to GPU hardware, and dates a May-2024 shift to competition-driven decline — i.e. hardware is not where falling prices come from. This cuts directly against Matsuoka's and Patel's memory/HBM-scarcity-as-destiny framing, in which hardware cost structure is the load-bearing variable. Both can be partly true (memory gates supply/capacity while software drives price), but as stated they locate the cost lever in different places. Estimated econometric decomposition (abstract-page) vs modelled scenario + attributed podcast; unadjudicated, and the >100% TFP residual needs the full PDF to weigh.
Predictions & Trends
- If Evans is right, inference pricing keeps compressing toward marginal cost as the supply crunch eases — value migrates to the application layer.
- If the depreciation conveyor is real, both things happen: prices commoditize and incumbents keep a 2-4x structural cost advantage. Commoditization would then be a story about who dies, not about margins converging.
- The rung-hold pattern, if it persists across providers, implies labs capture capability gains as margin rather than passing them through — keeping headline list prices sticky even as the price of fixed capability falls fast underneath. Under Matsuoka's corridor, that stickiness is also load-bearing for buildout solvency.
- The token-maximization → token-minimization shift, if real, moves the commercially meaningful unit from price per token to cost per task — the exact unit this KB has zero data on.
Knowledge Gaps
Areas where the KB needs more sources. Specific, and ordered by what would most change the topic's conclusions. Status tags reflect the 2026-07-23 compile (all six swept sources now compiled).
- [NOW THE SHARPEST GAP] No cost floor for any FRONTIER / PROPRIETARY model — only for one open 70B. DigitalOcean's measured $0.45–$20.32/M is Llama-3.3-70B on an H200. The margin debate (Evans 40–50% vs Patel 70%+) is about GPT-5.x / Claude / Gemini serving, whose true cost-to-serve is still unmeasured anywhere in this KB. Until a frontier-model cost floor exists, the lab-margin claims stay unfalsifiable even though an open-model floor now exists. — suggested action: find/derive a measured serving cost for a frontier-class model; suggested search: "measured serving cost GPT-5 Claude Gemini per million tokens 2026", "InferenceX frontier model cost per token"
- [PARTLY CLOSED 2026-07-23] The measured cost series is a single POINT, not a series. DigitalOcean gives one SKU (H200), one model (open 70B FP8), one framework (vLLM 0.24.0), one date (2026-07-08). No cost time-series; no reproduction on Blackwell GB200/GB300, TPU, Trainium, or custom ASIC; no long-context or reasoning-workload cost. Patil's own
vllm-cost-meteragainst real self-hosted traffic remains the cleanest way to build the series. — suggested action: runvllm-cost-meteron a real vLLM server; suggested search: "measured LLM serving cost Blackwell GB200 per million tokens 2026" - [CLOSED 2026-07-24] Cost-per-task: framework in, absolute dated levels still missing. Cost-of-Pass (2504.13359) supplies the methodology (cost ÷ accuracy) and AI Tokenomics (2606.24616) the value≠cost theory, but neither gives a dated absolute cost-per-correct-answer for a real 2026 agentic or reasoning workload. STAGE-Claw / WorkBench-Revisited (surfaced, not ingested) carry measured per-task $ ($0.17–$6.55/task) and are the next ingest. → Now ingested (2026-07-24): STAGE-Claw (2606.10394) supplies the absolute levels — ~$0.17–$6.55/task across 11 frontier models (search-surfaced, full-text confirmation owed) — and WorkBench (2606.13715) the trajectory; both on cost per task. The residual (a frontier/proprietary per-task cost time-series, and coverage beyond agentic/workplace benchmarks) is tracked under Frontiers 1 and 6, not here. — suggested action: full-text-confirm STAGE-Claw's cost table; derive a frontier-proprietary per-task series
- [STILL OPEN — count updated] Still zero peer-reviewed papers. Now 6 preprints + 3 analysis + 3 technical-report + 1 opinion; 0 peer-reviewed. The Joule/Cell "Energy use of AI inference, efficiency pathways, and test-time scaling" (ScienceDirect S2542435126001145) is peer-reviewed and directly on-thesis but 403'd this sweep — retry via an institutional/arXiv mirror. MLSys / NeurIPS 2026 serving-cost proceedings untried. — suggested action: retry the Joule paper; suggested search: "MLSys 2026 inference serving cost paper", "inference economics peer-reviewed journal 2026"
- [PARTLY CLOSED 2026-07-24] No energy / power-cost pass-through leg. Not ingested (the peer-reviewed energy paper was blocked). The web literature gives ~5×10⁻⁴ Wh/token, reasoning-mode 5–12× (Gemini Deep Think ~6.2 Wh/query vs ~0.6 Wh), and ~$254/mo electricity per H100 at $0.20/kWh — but none is in the KB as a $/kWh→$/token conversion. → Photons = Tokens (arXiv 2603.06630) now ingested → new concept the energy cost of tokens: supplies the top-down energy→token volume bridge (326 TWh → ≈6.5×10¹⁷ tokens/yr → ≈225,000 tokens/person/day, >3 OOM above mid-2024) and the "energy is not the near-term binding constraint on token volume — direction is" framing. Still owed: a bottom-up $/kWh→$/token per-SKU cost, the reasoning-mode 5–12× multiplier folded in, and the peer-reviewed Joule paper (still 403'd). — suggested action: retry Joule S2542435126001145; do the ~5×10⁻⁴ Wh/token × $/kWh multiply for a real SKU
- [RECORDED, UNADJUDICATED 2026-07-23] The software-vs-hardware cost-lever conflict. Tiered Super-Moore's –0.9% hardware / ~103.7% TFP decomposition (Conflict 7) vs Matsuoka/Patel memory-as-destiny is now recorded on the price-decline distribution and the depreciation conveyor, kept unresolved. It still needs the full PDFs to weigh — the >100% TFP residual and Malmquist internals are not visible from the abstract. Directly bears on whether the memory crunch governs price or only capacity. — suggested action: full-text read 2603.28576 + 2607.07207 side by side
- [STILL OPEN] No audited primary financials for any frontier lab. Anthropic's confidential draft S-1 (~2026-06-01) is the only date-certain resolution path; everything held is reported/unaudited investor-deck lineage. — suggested action: standing watch for the public S-1 filing
- [STILL OPEN] Rung persistence is a one-provider, one-event observation. Only OpenAI, only GPT-5.6. Tiered Super-Moore gives tier curve half-lives (economy 1.10yr / mid 1.55yr) across many models, but not generation-over-generation rung rows for Anthropic / Google / DeepSeek. — suggested action: extend the Token Price Index to every provider release event
- [PARTLY CLOSED 2026-07-23] List vs realised price — an observed discount ladder now exists, enterprise realised prices still don't. Together batch $0.65 / reserved $0.95 vs Fireworks serverless $1.20 is an observed within-market latency-discount ladder. But enterprise committed-use / prompt-caching realised prices remain invisible to both the TPI and this matrix. — suggested search: "enterprise LLM committed use discount realized price per token 2026"
- [PARTLY CLOSED 2026-07-24] Open-weight self-host vs paid API / open-weight vs frontier — open-70B leg done, frontier legs open. DigitalOcean's open-70B cost floor can now be set against Together's open-70B API list ($0.65/M batch, right on the floor). Still missing: cost-to-serve a frontier open-weight (GLM-5.2 / Kimi) and any comparison against a paid frontier proprietary rung. → WorkBench Revisited (2606.13715) adds a two-year benchmark datapoint on the asymmetry: the cost of a once-proprietary performance level collapsed ~100x since 2024, open-weight-driven, while frontier serving costs stayed flat — the decline lands in the open-weight rung, not the frontier rung (direction stated; per-tier $ levels owed to the full text). Recorded on the price-decline distribution. — suggested search: "cost to self-host GLM Kimi frontier open weight vs API 2026"
- [PARTLY CLOSED 2026-07-23] Independent providers — a Q2-2026 list snapshot now exists; margins and Baseten still don't. The DigitalApplied matrix gives observed list prices (6x same-model spread) but not realised prices, not margins, not Baseten, and only one model (Llama 4 70B) — no reasoning/long-context tiers. Provider ARRs and the Cerebras ~$66B IPO are logged as secondary/unverified. — suggested search: "Together Fireworks Baseten Groq Cerebras gross margin unit economics 2026"
- [STILL OPEN] The depreciation conveyor has never been checked against actual accounting. Hyperscaler server useful-life disclosures (10-Ks, revised upward) are primary, free, and directly test the conveyor's premise. — suggested action: pull useful-life disclosures from the latest hyperscaler 10-Ks
- [STILL OPEN — now with a training-cost anchor] No non-US pricing. DeepSeek / Qwen / GLM inference rungs are still absent. Tiered Super-Moore does log a US–China 63x training-cost gap attributed to architecture, not factor prices — but inference pricing decoupled from the memory crisis remains the single biggest structural surprise available and is unmeasured here. — suggested search: "DeepSeek Qwen GLM API pricing 2026 per million tokens"
- [PARTLY OPEN] Vendor efficiency / specialty-silicon claims logged but not reproduced. Groq ~750 tps and Cerebras ~620 tps decode (vs H100 ~115 tps) are now in the KB as observed throughput, but no provider "Nx cheaper" cost claim has been independently reproduced. — suggested search: "Groq Cerebras cost per token independent reproduction benchmark 2026"
- [NEW 2026-07-23] The six newest sources are abstract/web reads, not full-text. DigitalOcean and Epoch were read in full; the four arXiv/analysis sources were read from abstract or article pages. Three load-bearing unread items: (a) Tiered Super-Moore's >100% TFP residual + Chow/Malmquist internals, (b) Cost-of-Pass's current model set and the exact "halved every few months" window, (c) AI Tokenomics' proposed hidden-token measurement method. — suggested action: full-text read all four on the next pass
Evidence Located (2026-07-17 reading-desk session)
Located during a reading-desk pass; addresses the Anthropic-financials and demand-side gaps. All items secondary/reported unless noted — none are audited filings; graded strong/moderate/weak. Located facts for the record, not an editorial or market conclusion. Retained unchanged through the 2026-07-22 compile.
Anthropic financials (vs Patel's attributed 70%+ / $50B+ ARR / FCF+)
- Confidential draft Form S-1 filed ~2026-06-01 (company-announced). Confidential ⇒ audited financials not public; they release with the public S-1 — a date-certain resolution path. Evidence: moderate. (Yahoo Finance)
- Run-rate revenue ~$47B (as of May 2026), Anthropic-stated via reporting → Patel's "$50B+ ARR" is a mild, consistent extrapolation; claim survives. Evidence: moderate. (Simon Willison)
- Gross margin ~50% present-period (reportedly revised down 50%→~40% for 2025 on third-party-cloud inference cost; 77% projected only by 2028) → Patel's "70%+ now" is not corroborated as a current figure; ~70% matches the 2028 projection, and present-day sits inside Evans' 40–50%. The weak link in Patel's account. Evidence: moderate (WSJ/Information investor-deck lineage; unaudited). (Forbes, Sacra)
- "First-ever quarterly operating profit $559M projected Q2 2026 on $10.9B rev" → operating profit ≠ FCF; "first-ever" ⇒ Patel's flat "FCF-positive" is a days-old inflection, plausible not established. Context: $65B Series H at $965B post-money. Evidence: moderate.
- Net: ARR survives; FCF knife-edge; 70% margin is a 2028 projection in present tense. Still no audited primary — gap stays open pending the public S-1.
Demand side (to pair against the price-decline distribution)
- Goldman Sachs: token consumption 24× → ~120 quadrillion tokens/month by 2030 (agentic ~12× consumer + enterprise). Evidence: moderate. (Goldman Sachs)
- DERIVED (desk, 2026-07-22, not a Goldman claim): compounding 24x implies ~2.21x/yr over four years or ~1.89x/yr over five — i.e. this forecast lands on Matsuoka's ~2x/yr solvency threshold, not comfortably above it. Base year is ambiguous in what this KB holds, so both figures are given.
- Blended token price fell ~67%/yr (~3×) Q1'25→Q1'26. ⚠ Metric caution — this blended-spot rate is not Epoch's 50×/yr quality-held rate (price-decline-distribution); they measure different things and differ by >1 order of magnitude. The gap between them is the rung-hold mechanic (price-rung-persistence): quality-per-dollar improves ~50× while the buyer's unit price falls only ~3×. Evidence: moderate. (NeuralWired/Gartner)
- Agentic workloads 5–30× more tokens/task (Gartner, Mar 2026); 10–20 LLM calls/task; RAG 3–5× context; inference ≈ 85% of enterprise AI budgets; 73% of enterprises overran cost forecasts. Evidence: weak-moderate. (Forbes Tech Council)
- Reading (located, not a view): unit price falling while per-task consumption explodes faster ⇒ total spend rises despite falling unit price (a Jevons dynamic). Sharpens the commodity question to which curve is steeper, not whether price falls. ⚠ Now contested — Matsuoka argues the trackers behind these figures overstate monetizable demand (Conflict 3 above).
Bonus — Patel's memory-shortage claim, verified
- Patel: capacity +20–30%/yr vs demand ~doubling. Actual: IDC 2026 DRAM supply +16% / NAND +17% (below Patel); HBM demand +130% (2025) / +70% (2026) (TrendForce); Micron "demand significantly exceeds supply," tightness beyond 2027 → structural-shortage thesis well-corroborated (supply growth if anything lower than Patel said). Asymmetry: Patel's checkable hardware claim holds; his conflicted Anthropic-margin claim dissolves. Evidence: moderate. (IDC, TrendForce)
Ingest Status
The three arXiv preprints registered as discovered in the 2026-07-16 sweep were ingested and compiled on 2026-07-22 — from their abstract pages only, with the full PDFs unread. The open question that sweep raised — "does the depreciation conveyor resolve, sharpen, or contradict Patel's attributed Anthropic-margin claim?" — is now answerable in part: it does neither, because Patil's work shows the margin claim has no denominator to compare against. The conveyor and the margin claim are about different quantities.
The 2026-07-23 discovery sweep added six new sources (now 13 total): DigitalOcean measured H200 cost series, Epoch serving-capacity model, and four economics sources — AI Tokenomics (2606.24616), Tiered Super-Moore's Law (2603.28576), Cost-of-Pass (2504.13359), and the DigitalApplied Q2-2026 independent-provider matrix. All six are now compiled (2026-07-23); staleSources = 0. The compile delivered what discovery flagged as owed: (1) the measured cost floor routed into the cost-measurement problem with its exact scope, plus the independent inference providers entity; (2) the software-vs-hardware cost-lever conflict (Conflict 7) recorded on the price-decline distribution and the depreciation conveyor, unresolved; (3) a new cost-per-task concept anchored on Cost-of-Pass + AI Tokenomics' value≠cost distinction; and (4) a beliefs check — no formal theses.md exists in this topic (confidence is tracked inline in concept prose + this frontier), so the software-vs-hardware conflict was recorded as a live, unadjudicated tension rather than a confidence move. The measured open-model floor (~$0.45–$1.17/M) and the –0.9%-hardware decomposition both bear on "does inference commoditize?" but neither resolves it — the frontier-model cost floor remains the gate.
The 2026-07-24 compile added three more sources (now 16 total): STAGE-Claw (2606.10394), WorkBench Revisited (2606.13715), and Photons = Tokens (2603.06630) — all previously surfaced/targeted, now compiled at abstract grade (STAGE-Claw's cost table search-surfaced, full-text confirmation owed). Net effect: the topic's first absolute dated cost-per-task levels are in (~$0.17–$6.55/task, STAGE-Claw); the open-weight-cheap / frontier-flat bifurcation gained a benchmark datapoint (WorkBench, on the price-decline distribution); and the energy→token leg opened as a new concept, the energy cost of tokens (Photons = Tokens). conceptCount = 8; staleSources = 0. Still zero peer-reviewed papers (9 preprints + 3 analysis + 3 technical-report + 1 opinion). The sharpest gaps are unchanged and unclosed: no measured cost floor for any frontier/proprietary model, and no frontier-proprietary cost-per-task time-series.