Memory Scarcity & Inference Economics ($/PB)
Active FrontierMemory Scarcity & Inference Economics ($/PB)
The 2026 DRAM/HBM price surge stopped being a component-cost story and became a market-structure story. The argument on this page comes almost entirely from one source — Satoshi Matsuoka, Memory Scarcity, Open Models, and the Restructuring of the AI Industry, 2026–2030 (arXiv:2607.07207, econ.GN, 8 July 2026) — plus the desk's own close read of the 22-page PDF. It is recorded here because the framework is useful, not because its numbers are established facts.
Read the provenance before the numbers. This is a single-author, non-peer-reviewed preprint by the director of RIKEN Center for Computational Science (Japan's national supercomputing centre — see RIKEN R-CCS). Every headline figure is a model output under stipulated parameters, not an observation. Its DRAM-price and HBM-wafer-share inputs are TrendForce numbers the paper cites — i.e. summary-derived twice over by the time they reach this page. Treat the whole page as an argument to test, capped at moderate evidence.
The unit: dollars per petabyte of bandwidth delivered
The paper's contribution is a model-agnostic cost unit for saturated, bandwidth-bound decode — the inference regime where the bottleneck is moving KV-cache and weights through memory, not raw FLOPs (see Prefill/Decode Disaggregation for why that scoping matters). Section 2, Eq. 1:
Cost_PB = [ (P_acc · k)/(8760·T) + (W/1000)·PUE·c_e + o ] / (B · η · 3.6) [$/PB]
Stipulated parameters: system multiplier k = 1.5×, depreciation life T = 4 years straight-line,
8760 hours/yr (implies 100% uptime), PUE 1.3, electricity $0.08/kWh, allocated opex
$0.25/GPU-hr, memory-bandwidth utilization η = 0.50 Hopper / 0.55 Blackwell / 0.60–0.62 Rubin.
The depreciation conveyor — the strongest idea here
New entrants always buy at spot prices; incumbents keep rolling fleets off 4-year amortization faster than hardware prices normalize. So the entrant-vs-incumbent cost gap never closes inside the horizon, and the cost advantage rotates among incumbents by vintage rather than ever transferring to entrants. This is a structural argument rather than a forecast, which is why it survives disagreement about the parameters.
| Year | Entrant $/PB ÷ incumbent floor $/PB | Note |
|---|---|---|
| 2026 | 3.2× (entrant ≈ $0.174/PB; floor ≈ $0.054/PB) | Peak of the shortage |
| 2027 | 1.9–2.0× | Narrowest point — the paper's "best entry year" |
| 2029–30 | 3–4× | Re-widens as B200-class incumbent fleets finish amortizing |
The vintage-breakeven grid — 16 cells, not 3 rows
How this is represented here matters, because the desk's first ingest of this paper got it wrong. The published result is §6, Figure 4 (p.8) — one paragraph of prose plus one bar chart, no table. It is a 4 vintages × 2 pricing regimes × 2 HBM branches = 16-cell grid. Values below were read off the bars (±0.5pp) in the desk's close read and cross-checked against the four numbers §6 states in prose.
The question each cell answers is narrow: what share of the tokens a purchase-year cohort serves must earn premium pricing for that fleet to break even? It is a break-even mix requirement — not an IRR, not a payback period, not an NPV.
| Vintage | Coupled, HBM normal | Coupled, shortage | Sticky, HBM normal | Sticky, shortage |
|---|---|---|---|---|
| 2026 | ~24.5% | ~24.5% | ~31.3% | ~31.3% |
| 2027 | ~7.5% | ~8.7% | ~9.6% | ~11.1% |
| 2028 | ~12.5% | ~21.1% | ~10.3% | ~17.3% |
| 2029 | ~25.2% | ~37.6% | ~11.6% | ~17.3% |
The mechanism — and the thing an ordinary summary loses. The two regimes are pegged differently, so they cross over:
- Coupled = routing arbitrage drags premium down with the mass floor:
P_premium = 7 × P_mass. A ratio peg. - Sticky = premium holds at $0.40/PB in absolute terms, defended by enterprise contracts and switching costs. An absolute peg.
- Mass price is defined as incumbent floor + 30% margin, and the floor falls every year as older fleets finish amortizing ($0.054/PB in 2026 → $0.022/PB by 2029). Every new vintage is priced against a floor set by someone else's sunk capital.
- In 2026 the coupled premium works out ≈ $0.49/PB — above the $0.40 sticky level. By 2029 the coupled premium ≈ $0.20/PB — half it.
Consequence: the 2026 vintage is hurt by premium stickiness, not helped by it (worst cell ~31.3% sticky, vs ~24.5% coupled). 2028–29 is the mirror image — comfortable under sticky (10.3–17.3%), broken by coupled-plus-shortage (21.1% and 37.6%). 2029 is nearly 2× as exposed as 2028; lumping them together hides the single most exposed cell in the paper. Only 2027 is robust in all four of its cells (7.5–11.1%).
The paper's own summary (§6, p.7): routing success breaks the Stargate-class 2028–29 commitments; routing failure breaks today's peak-price 2026 buyers. The pricing regime only selects who gets hurt.
Internal-consistency check that validates the chart reading: the 2026 row is identical across both HBM branches — as it must be, since that hardware was already bought and a future HBM branch cannot alter its capital cost.
The verdict hinges on an unsourced parameter
Every "unsustainable" call above is a comparison against a 10–20% plausible realized premium share band that is asserted in §6 with no citation. Move the band to 25–35% and the 2026 vintage clears, 2028 clears everywhere, and only the 2029 coupled-shortage cell fails — the headline conclusion substantially dissolves. The nearest empirical anchor appears three sections later (Anthropic taking ~42% of OpenRouter revenue on ~11% of token share): one platform, one month, and a revenue-share datum being used to bound a token-share parameter. Always report the band alongside the percentages, or report neither.
The demand-measurement critique
The paper assembles public token trackers (Google 9.7T/mo May 2024 → 3.2 quadrillion/mo May 2026; Microsoft Foundry ~2.9× annualized; OpenRouter ~4–6× YoY with agentic traffic overtaking human traffic ~1 Feb 2026 at ~15× tokens/request; China aggregate ~30T/day → ~180T/day; Goldman 24× by 2030 ≈ 2.2×/yr) and argues five upward biases inflate them:
| # | Bias | Evidence status |
|---|---|---|
| 1 | Supply-injected volume — operator pushes inference onto users (AI summaries, multimodal) | Asserted; no fraction estimated |
| 2 | Substitution — router growth includes migration onto the router from direct APIs | Asserted; no migration rate |
| 3 | Subsidy — volume demanded at below-cost prices (esp. the Chinese price war) | Asserted; weakest of the five |
| 4 | Quality composition — the ~15× agentic multiplier is largely redundant context re-reading | Best evidenced — three independent legs incl. ~60% of enterprises throttling AI spend |
| 5 | Base effects + disclosure selection (two mechanisms under one label) | Half-evidenced; disclosure selection is pure assertion |
The revised headline — 4–7×/yr → "plausibly 2–3×/yr" — is asserted, not decomposed. No bias is individually quantified, there are no error bars and no sensitivity. Four of the five contribute no quantity at all. This is the paper's most position-relevant number (it lands the central case barely above the paper's own 1.6–2.4× solvency threshold) and its least supported. Cite it as his estimate, never as a measurement.
Three further pieces of the critique, all operationally useful and none dependent on the 2–3× number:
- Denominate the corridor in dollars, not tokens. Neutral-router blended realization is ~$1 per million tokens against flagship closed pricing near $25–30. Bandwidth demand can grow while revenue per unit of capacity stalls. Stated crash signature: tokens growing while dollars stall.
- The reallocation channel (§9.2). Metered cloud token demand — the only demand that services the capex — can decelerate even while total inference grows, as inference migrates onto customer-owned hardware and as non-US enterprises enter directly at the on-premises open-weight stage, skipping the metered API entirely. Demand relocation reads as demand destruction on the balance sheets that matter. Effectively a sixth bias, and independent of any view on aggregate demand.
- The projection-vintage problem (§9.3). Every optimistic projection in the survey predates a Q2-2026 doctrinal inversion from token maximization to token minimization. All trailing-CAGR extrapolations — including the paper's own 2–3× estimate, which the author explicitly demotes to an upper bound — measure a regime that may no longer exist. Post-break data spanned only weeks at writing. Countervailing datum the author supplies: neutral-router volumes were still accelerating through late June 2026.
Custom-silicon entrant model
Removing NVIDIA's merchant margin (~20–35% of cost) does not remove the memory premium, because HBM is 45–50% of a custom accelerator's BOM and is "priced by the same three suppliers under the same shortage" regardless of who designs the logic die. A well-executed custom build reaches $0.082/PB — roughly merchant-GPU parity — against an incumbent depreciated floor of $0.022–0.037/PB (2.2–3.3× cheaper).
The frequently-quoted 25% success / 34% mediocre / 41% loss outcome split is elicited subjective probability, revised across five rounds, explicitly labelled in the paper as judgmental assessment and NOT model output. Never present it as a derived result. Stated mitigants (anchor demand, a secured HBM long-term agreement, or a non-HBM architecture) roughly double success probability to 45–50%; staged go/no-go capital gates cut deployed-capital loss probability to ~24%.
Scenario map (Section 8 — iterated expert judgment, not computation)
| Scenario | Prob. | Core mechanism |
|---|---|---|
| Rotating Landlord Oligopoly | 25% | Incumbents with depreciated fleets set mass-market prices from their amortized floor indefinitely |
| Commoditization Crash | 25% | Token growth <1.7×/yr and/or pricing decouples; routing + local inference hollow out centralized demand |
| Jevons Absorption | 20% | Token growth ≥2.5×/yr sustained by agentic workloads; all vintages solvent |
| System-Layer Re-differentiation | 18% | Sticky premium holds; value migrates to orchestration and proprietary RL environments |
| Geopolitical Bifurcation | 12% | Export controls extend bidirectionally; sovereign compute capacity appreciates |
Across five revision rounds the Crash scenario moved 15% → peak 30% → tempered to 25% "on external-review grounds."
Key Claims
- $/PB is a defensible model-agnostic unit for saturated, bandwidth-bound decode, given in closed form with explicit parameters. Evidence: moderate — framework, single non-peer-reviewed source (Matsuoka, close read)
- The depreciation conveyor keeps the entrant/incumbent gap open across 2026–2030 — 3.2× → 1.9× → 3–4×. Evidence: moderate — structural argument under stipulated parameters, not an observation (close read)
- 2027-vintage capacity is robust in all four regime/branch cells (7.5–11.1%); 2026 and 2028–29 are each fatally exposed to one regime — 2026 to sticky (~31.3%), 2029 to coupled-plus-shortage (~37.6%). Evidence: moderate — model output; verdict depends on an unsourced 10–20% band (close read)
- The two pricing regimes cross over because coupled is a ratio peg and sticky is an absolute peg — this is genuinely derived, not assumed, and it is the paper's real §6 contribution. Evidence: moderate (close read)
- Custom silicon removes merchant margin but not the memory premium — HBM 45–50% of BOM; best-case custom build $0.082/PB vs $0.022–0.037/PB incumbent floor. Evidence: moderate (Matsuoka)
- Metered-cloud demand can decelerate while total inference grows (on-prem migration + sovereignty channel). Evidence: moderate — mechanism is clean; magnitude unquantified (close read)
- DRAM contract prices rose ~90% Q1'26 vs Q4'25; memory is 40–50% of accelerator BOM; HBM took 23% of DRAM wafer output Q1–Q2'26. Evidence: weak — TrendForce figures cited by the paper, i.e. summary-derived twice over; not independently verified by this KB (Matsuoka)
- "Demand growth is 2–3×/yr." Evidence: weak — asserted after a qualitative bias list, no decomposition, self-labelled an upper bound. Cite as the author's estimate only.
- The 25/34/41 custom-silicon outcome split. Evidence: weak — elicited subjective probability revised over five rounds; explicitly not model output.
Where this model is weakest
- A 73× realized-cost spread swamps a 2–4× vintage effect. The paper's own cited reference shows effective cost on identical H100 hardware ranging $0.21 to $15.25 per million output tokens depending on utilization and offered load. The entire vintage grid discriminates between cohorts separated by 2–4×. The paper's §11 claim that "vintage timing dominates operating skill" is close to backwards on its own evidence — do not repeat it as fact.
- The 10–20% premium band carries the whole verdict and is unsourced.
- The 2–3×/yr demand revision is asserted, not computed.
- Undisclosed structural interest. The paper identifies mission-funded scientific compute (AI4SIS) — precisely the author's institution's demand class — as the corridor's "inelastic floor"; reads China's LineShine LX2 as vindicating an HBM-fed general-purpose CPU architecture in direct lineage from Fugaku's A64FX; and closes with policy addressed to Japan. Meticulous AI disclosure, no author-position disclosure. Weight those sections accordingly.
- AI-authorship reflexivity. Models/figures/drafts developed with Anthropic's Claude, plus review comments from OpenAI's ChatGPT adopted into the very sections that assess those companies' commercial prospects. The author flags this himself.
- Uptime and utilization are assumed away (
8760·Timplies 100% uptime; MBU is one scalar per platform). Every $/PB figure is a best-case floor — and entrants are likelier to be under-utilized than incumbents, so the true gap is probably wider than 3.2×. - No supplier-level memory-maker analysis whatsoever. No per-supplier capex, allocation, margin or strategy. Memory scarcity enters as a single exogenous input. This paper cannot corroborate any Samsung / SK Hynix / Micron-specific claim.
Open Questions
- Does the 10–20% realized-premium-share band survive contact with actual platform revenue disclosures, or is it the load-bearing guess it currently looks like?
- Does the "tokens growing while dollars stall" signature show up in Q3'26 hyperscaler disclosures?
- How large is the reallocation channel (on-prem + sovereign) in dollars, not anecdotes?
- Does the 2027-as-robust-vintage conclusion survive if new fab capacity lands earlier than 2027–28?
- Does anyone reproduce the vintage grid with a utilization distribution instead of 100% uptime?
Related Concepts
- HBM4 Memory Architecture — the supply-side object whose price this model treats as exogenous
- Custom Silicon vs GPU — the entrant model lives at this boundary
- Prefill/Decode Disaggregation — $/PB is scoped to bandwidth-bound decode only
- KV Cache Management — the software layer attacking the same bytes-per-token term
- Advanced Packaging & CoWoS Capacity — the other capacity gate on the same accelerators
Backlinks
Pages that reference this concept:
Changelog
- 2026-07-22 — Initial compilation from Matsuoka arXiv:2607.07207 + the desk's 22-page close read. Represented the vintage-breakeven result as the corrected 16-cell grid (4 vintages × 2 regimes × 2 HBM branches) with the 2026 sticky/coupled inversion fixed and 2028/2029 separated; custom-silicon build cost recorded as $0.082/PB; 25/34/41 split labelled elicited subjective probability. Demand-measurement critique recorded with per-bias evidence status, the undecomposed 2–3× revision flagged, and §9.2 reallocation + §9.3 projection-vintage added. Provenance capped at moderate: single-author non-peer-reviewed econ.GN preprint, RIKEN R-CCS structural interest noted, TrendForce inputs labelled summary-derived twice over.