Memory Scarcity & Inference Economics ($/PB)

Active Frontier
Sign in to track mastery·Sign in
memory-scarcityinference-economicsdramhbmmemory-pricingvintage-breakevendepreciation-conveyortoken-economicscapexscenario-analysis

Memory Scarcity & Inference Economics ($/PB)

The 2026 DRAM/HBM price surge stopped being a component-cost story and became a market-structure story. The argument on this page comes almost entirely from one source — Satoshi Matsuoka, Memory Scarcity, Open Models, and the Restructuring of the AI Industry, 2026–2030 (arXiv:2607.07207, econ.GN, 8 July 2026) — plus the desk's own close read of the 22-page PDF. It is recorded here because the framework is useful, not because its numbers are established facts.

Read the provenance before the numbers. This is a single-author, non-peer-reviewed preprint by the director of RIKEN Center for Computational Science (Japan's national supercomputing centre — see RIKEN R-CCS). Every headline figure is a model output under stipulated parameters, not an observation. Its DRAM-price and HBM-wafer-share inputs are TrendForce numbers the paper cites — i.e. summary-derived twice over by the time they reach this page. Treat the whole page as an argument to test, capped at moderate evidence.

The unit: dollars per petabyte of bandwidth delivered

The paper's contribution is a model-agnostic cost unit for saturated, bandwidth-bound decode — the inference regime where the bottleneck is moving KV-cache and weights through memory, not raw FLOPs (see Prefill/Decode Disaggregation for why that scoping matters). Section 2, Eq. 1:

Cost_PB = [ (P_acc · k)/(8760·T) + (W/1000)·PUE·c_e + o ] / (B · η · 3.6)   [$/PB]

Stipulated parameters: system multiplier k = 1.5×, depreciation life T = 4 years straight-line, 8760 hours/yr (implies 100% uptime), PUE 1.3, electricity $0.08/kWh, allocated opex $0.25/GPU-hr, memory-bandwidth utilization η = 0.50 Hopper / 0.55 Blackwell / 0.60–0.62 Rubin.

The depreciation conveyor — the strongest idea here

New entrants always buy at spot prices; incumbents keep rolling fleets off 4-year amortization faster than hardware prices normalize. So the entrant-vs-incumbent cost gap never closes inside the horizon, and the cost advantage rotates among incumbents by vintage rather than ever transferring to entrants. This is a structural argument rather than a forecast, which is why it survives disagreement about the parameters.

YearEntrant $/PB ÷ incumbent floor $/PBNote
20263.2× (entrant ≈ $0.174/PB; floor ≈ $0.054/PB)Peak of the shortage
20271.9–2.0×Narrowest point — the paper's "best entry year"
2029–303–4×Re-widens as B200-class incumbent fleets finish amortizing

The vintage-breakeven grid — 16 cells, not 3 rows

How this is represented here matters, because the desk's first ingest of this paper got it wrong. The published result is §6, Figure 4 (p.8) — one paragraph of prose plus one bar chart, no table. It is a 4 vintages × 2 pricing regimes × 2 HBM branches = 16-cell grid. Values below were read off the bars (±0.5pp) in the desk's close read and cross-checked against the four numbers §6 states in prose.

The question each cell answers is narrow: what share of the tokens a purchase-year cohort serves must earn premium pricing for that fleet to break even? It is a break-even mix requirement — not an IRR, not a payback period, not an NPV.

VintageCoupled, HBM normalCoupled, shortageSticky, HBM normalSticky, shortage
2026~24.5%~24.5%~31.3%~31.3%
2027~7.5%~8.7%~9.6%~11.1%
2028~12.5%~21.1%~10.3%~17.3%
2029~25.2%~37.6%~11.6%~17.3%

The mechanism — and the thing an ordinary summary loses. The two regimes are pegged differently, so they cross over:

  • Coupled = routing arbitrage drags premium down with the mass floor: P_premium = 7 × P_mass. A ratio peg.
  • Sticky = premium holds at $0.40/PB in absolute terms, defended by enterprise contracts and switching costs. An absolute peg.
  • Mass price is defined as incumbent floor + 30% margin, and the floor falls every year as older fleets finish amortizing ($0.054/PB in 2026 → $0.022/PB by 2029). Every new vintage is priced against a floor set by someone else's sunk capital.
  • In 2026 the coupled premium works out ≈ $0.49/PB — above the $0.40 sticky level. By 2029 the coupled premium ≈ $0.20/PB — half it.

Consequence: the 2026 vintage is hurt by premium stickiness, not helped by it (worst cell ~31.3% sticky, vs ~24.5% coupled). 2028–29 is the mirror image — comfortable under sticky (10.3–17.3%), broken by coupled-plus-shortage (21.1% and 37.6%). 2029 is nearly 2× as exposed as 2028; lumping them together hides the single most exposed cell in the paper. Only 2027 is robust in all four of its cells (7.5–11.1%).

The paper's own summary (§6, p.7): routing success breaks the Stargate-class 2028–29 commitments; routing failure breaks today's peak-price 2026 buyers. The pricing regime only selects who gets hurt.

Internal-consistency check that validates the chart reading: the 2026 row is identical across both HBM branches — as it must be, since that hardware was already bought and a future HBM branch cannot alter its capital cost.

The verdict hinges on an unsourced parameter

Every "unsustainable" call above is a comparison against a 10–20% plausible realized premium share band that is asserted in §6 with no citation. Move the band to 25–35% and the 2026 vintage clears, 2028 clears everywhere, and only the 2029 coupled-shortage cell fails — the headline conclusion substantially dissolves. The nearest empirical anchor appears three sections later (Anthropic taking ~42% of OpenRouter revenue on ~11% of token share): one platform, one month, and a revenue-share datum being used to bound a token-share parameter. Always report the band alongside the percentages, or report neither.

The demand-measurement critique

The paper assembles public token trackers (Google 9.7T/mo May 2024 → 3.2 quadrillion/mo May 2026; Microsoft Foundry ~2.9× annualized; OpenRouter ~4–6× YoY with agentic traffic overtaking human traffic ~1 Feb 2026 at ~15× tokens/request; China aggregate ~30T/day → ~180T/day; Goldman 24× by 2030 ≈ 2.2×/yr) and argues five upward biases inflate them:

#BiasEvidence status
1Supply-injected volume — operator pushes inference onto users (AI summaries, multimodal)Asserted; no fraction estimated
2Substitution — router growth includes migration onto the router from direct APIsAsserted; no migration rate
3Subsidy — volume demanded at below-cost prices (esp. the Chinese price war)Asserted; weakest of the five
4Quality composition — the ~15× agentic multiplier is largely redundant context re-readingBest evidenced — three independent legs incl. ~60% of enterprises throttling AI spend
5Base effects + disclosure selection (two mechanisms under one label)Half-evidenced; disclosure selection is pure assertion

The revised headline — 4–7×/yr → "plausibly 2–3×/yr" — is asserted, not decomposed. No bias is individually quantified, there are no error bars and no sensitivity. Four of the five contribute no quantity at all. This is the paper's most position-relevant number (it lands the central case barely above the paper's own 1.6–2.4× solvency threshold) and its least supported. Cite it as his estimate, never as a measurement.

Three further pieces of the critique, all operationally useful and none dependent on the 2–3× number:

  • Denominate the corridor in dollars, not tokens. Neutral-router blended realization is ~$1 per million tokens against flagship closed pricing near $25–30. Bandwidth demand can grow while revenue per unit of capacity stalls. Stated crash signature: tokens growing while dollars stall.
  • The reallocation channel (§9.2). Metered cloud token demand — the only demand that services the capex — can decelerate even while total inference grows, as inference migrates onto customer-owned hardware and as non-US enterprises enter directly at the on-premises open-weight stage, skipping the metered API entirely. Demand relocation reads as demand destruction on the balance sheets that matter. Effectively a sixth bias, and independent of any view on aggregate demand.
  • The projection-vintage problem (§9.3). Every optimistic projection in the survey predates a Q2-2026 doctrinal inversion from token maximization to token minimization. All trailing-CAGR extrapolations — including the paper's own 2–3× estimate, which the author explicitly demotes to an upper bound — measure a regime that may no longer exist. Post-break data spanned only weeks at writing. Countervailing datum the author supplies: neutral-router volumes were still accelerating through late June 2026.

Custom-silicon entrant model

Removing NVIDIA's merchant margin (~20–35% of cost) does not remove the memory premium, because HBM is 45–50% of a custom accelerator's BOM and is "priced by the same three suppliers under the same shortage" regardless of who designs the logic die. A well-executed custom build reaches $0.082/PB — roughly merchant-GPU parity — against an incumbent depreciated floor of $0.022–0.037/PB (2.2–3.3× cheaper).

The frequently-quoted 25% success / 34% mediocre / 41% loss outcome split is elicited subjective probability, revised across five rounds, explicitly labelled in the paper as judgmental assessment and NOT model output. Never present it as a derived result. Stated mitigants (anchor demand, a secured HBM long-term agreement, or a non-HBM architecture) roughly double success probability to 45–50%; staged go/no-go capital gates cut deployed-capital loss probability to ~24%.

Scenario map (Section 8 — iterated expert judgment, not computation)

ScenarioProb.Core mechanism
Rotating Landlord Oligopoly25%Incumbents with depreciated fleets set mass-market prices from their amortized floor indefinitely
Commoditization Crash25%Token growth <1.7×/yr and/or pricing decouples; routing + local inference hollow out centralized demand
Jevons Absorption20%Token growth ≥2.5×/yr sustained by agentic workloads; all vintages solvent
System-Layer Re-differentiation18%Sticky premium holds; value migrates to orchestration and proprietary RL environments
Geopolitical Bifurcation12%Export controls extend bidirectionally; sovereign compute capacity appreciates

Across five revision rounds the Crash scenario moved 15% → peak 30% → tempered to 25% "on external-review grounds."

Key Claims

  • $/PB is a defensible model-agnostic unit for saturated, bandwidth-bound decode, given in closed form with explicit parameters. Evidence: moderate — framework, single non-peer-reviewed source (Matsuoka, close read)
  • The depreciation conveyor keeps the entrant/incumbent gap open across 2026–2030 — 3.2× → 1.9× → 3–4×. Evidence: moderate — structural argument under stipulated parameters, not an observation (close read)
  • 2027-vintage capacity is robust in all four regime/branch cells (7.5–11.1%); 2026 and 2028–29 are each fatally exposed to one regime — 2026 to sticky (~31.3%), 2029 to coupled-plus-shortage (~37.6%). Evidence: moderate — model output; verdict depends on an unsourced 10–20% band (close read)
  • The two pricing regimes cross over because coupled is a ratio peg and sticky is an absolute peg — this is genuinely derived, not assumed, and it is the paper's real §6 contribution. Evidence: moderate (close read)
  • Custom silicon removes merchant margin but not the memory premium — HBM 45–50% of BOM; best-case custom build $0.082/PB vs $0.022–0.037/PB incumbent floor. Evidence: moderate (Matsuoka)
  • Metered-cloud demand can decelerate while total inference grows (on-prem migration + sovereignty channel). Evidence: moderate — mechanism is clean; magnitude unquantified (close read)
  • DRAM contract prices rose ~90% Q1'26 vs Q4'25; memory is 40–50% of accelerator BOM; HBM took 23% of DRAM wafer output Q1–Q2'26. Evidence: weak — TrendForce figures cited by the paper, i.e. summary-derived twice over; not independently verified by this KB (Matsuoka)
  • "Demand growth is 2–3×/yr." Evidence: weak — asserted after a qualitative bias list, no decomposition, self-labelled an upper bound. Cite as the author's estimate only.
  • The 25/34/41 custom-silicon outcome split. Evidence: weak — elicited subjective probability revised over five rounds; explicitly not model output.

Where this model is weakest

  1. A 73× realized-cost spread swamps a 2–4× vintage effect. The paper's own cited reference shows effective cost on identical H100 hardware ranging $0.21 to $15.25 per million output tokens depending on utilization and offered load. The entire vintage grid discriminates between cohorts separated by 2–4×. The paper's §11 claim that "vintage timing dominates operating skill" is close to backwards on its own evidence — do not repeat it as fact.
  2. The 10–20% premium band carries the whole verdict and is unsourced.
  3. The 2–3×/yr demand revision is asserted, not computed.
  4. Undisclosed structural interest. The paper identifies mission-funded scientific compute (AI4SIS) — precisely the author's institution's demand class — as the corridor's "inelastic floor"; reads China's LineShine LX2 as vindicating an HBM-fed general-purpose CPU architecture in direct lineage from Fugaku's A64FX; and closes with policy addressed to Japan. Meticulous AI disclosure, no author-position disclosure. Weight those sections accordingly.
  5. AI-authorship reflexivity. Models/figures/drafts developed with Anthropic's Claude, plus review comments from OpenAI's ChatGPT adopted into the very sections that assess those companies' commercial prospects. The author flags this himself.
  6. Uptime and utilization are assumed away (8760·T implies 100% uptime; MBU is one scalar per platform). Every $/PB figure is a best-case floor — and entrants are likelier to be under-utilized than incumbents, so the true gap is probably wider than 3.2×.
  7. No supplier-level memory-maker analysis whatsoever. No per-supplier capex, allocation, margin or strategy. Memory scarcity enters as a single exogenous input. This paper cannot corroborate any Samsung / SK Hynix / Micron-specific claim.

Open Questions

  • Does the 10–20% realized-premium-share band survive contact with actual platform revenue disclosures, or is it the load-bearing guess it currently looks like?
  • Does the "tokens growing while dollars stall" signature show up in Q3'26 hyperscaler disclosures?
  • How large is the reallocation channel (on-prem + sovereign) in dollars, not anecdotes?
  • Does the 2027-as-robust-vintage conclusion survive if new fab capacity lands earlier than 2027–28?
  • Does anyone reproduce the vintage grid with a utilization distribution instead of 100% uptime?

Related Concepts

Backlinks

Pages that reference this concept:

Changelog

  • 2026-07-22 — Initial compilation from Matsuoka arXiv:2607.07207 + the desk's 22-page close read. Represented the vintage-breakeven result as the corrected 16-cell grid (4 vintages × 2 regimes × 2 HBM branches) with the 2026 sticky/coupled inversion fixed and 2028/2029 separated; custom-silicon build cost recorded as $0.082/PB; 25/34/41 split labelled elicited subjective probability. Demand-measurement critique recorded with per-bias evidence status, the undecomposed 2–3× revision flagged, and §9.2 reallocation + §9.3 projection-vintage added. Provenance capped at moderate: single-author non-peer-reviewed econ.GN preprint, RIKEN R-CCS structural interest noted, TrendForce inputs labelled summary-derived twice over.
Memory Scarcity & Inference Economics ($/PB) | KB | MenFem