Memory Scarcity, Open Models, and the Restructuring of the AI Industry, 2026-2030
Quantitative scenario model showing the 2026 DRAM/HBM price surge creates a persistent entrant-incumbent inference-cost gap that never closes over 2026-2030 (3.2x in 2026, narrowing to ~1.9x in 2027, re-widening to 3-4x by 2029-30) via a depreciation-conveyor mechanism; five probability-weighted scenarios (Rotating Landlord Oligopoly 25%, Commoditization Crash 25%, Jevons Absorption 20%, System-Layer Re-differentiation 18%, Geopolitical Bifurcation 12%) map the resolution.
Memory Scarcity, Open Models, and the Restructuring of the AI Industry, 2026-2030
Subtitle: A quantitative scenario analysis of inference economics, training-cost divergence, and infrastructure solvency. 22 pages. Primary category econ.GN; cross-listed cs.AI, cs.AR, cs.CE, cs.PF.
Abstract (paraphrased)
The paper argues four forces are jointly reshaping the AI industry between 2026 and 2030: the DRAM/HBM price surge, the arrival of frontier-capable open-weight models (the paper's reference point is GLM-5.2), fast inference-efficiency gains (near-Shannon-limit KV-cache compression, lightweight local runtimes), and Meta and xAI entering the compute-resale market on fleets bought before memory repricing. The author's central metric is inference cost expressed in dollars per petabyte of bandwidth delivered ($/PB) — a model-agnostic unit for bandwidth-bound decode workloads. Using that metric, the paper finds the cost gap between new market entrants and depreciated incumbents never closes across the horizon: incumbents keep getting newly amortized fleets faster than hardware prices normalize. Training splits into a luxury tier (frontier runs costing $18-38B by 2030) and a mass tier (previous-frontier-parity models falling toward roughly $5M via RL/distillation). Solvency of the announced compute buildout is confined to a narrow corridor requiring sustained ~2x/year token-demand growth for four years plus sticky premium pricing — and the paper argues public token-usage trackers systematically overstate real monetizable demand. A vintage-by-vintage breakeven analysis finds 2026- and 2028/29-built capacity each fatally exposed under one pricing regime or the other, with only 2027-vintage capacity robust across regimes. A custom-silicon entrant removes merchant margin but not the memory premium itself (headline outcome distribution: 25% success / 34% mediocre / 41% loss, improvable via staged go/no-go capital gates). China's LineShine LX2 system, built on domestic HBM, decouples its cost curve from the global memory crisis entirely. Direct from source (arXiv:2607.07207 abstract): "Solvency now depends on monetized bandwidth demand, premium stickiness, and vintage ownership."
Key Contributions
- A model-agnostic cost unit ($/PB, dollars per petabyte of bandwidth delivered) for comparing inference economics across accelerator generations and buyers, formalized in Section 2 (Eq. 1): cost combines accelerator amortization, power (watts × PUE × electricity cost), and other opex, divided by effective bandwidth throughput.
- The "depreciation conveyor" mechanism — the structural reason the entrant-incumbent cost gap doesn't close even as hardware prices eventually normalize: incumbents' fleets keep rolling off 4-year amortization schedules faster than new entrants can buy in at parity.
- A vintage-breakeven framework classifying compute purchased in different years (2026 / 2027 / 2028 / 2029 — four separate vintages, not three; the PDF Fig. 4 grid is 4 vintages x 2 regimes x 2 HBM branches = 16 cells) by how much premium-pricing share they need to break even under both "coupled" (routing arbitrage drags premiums to the mass floor) and "sticky" (premiums hold in absolute dollars) pricing regimes.
- A demand-measurement critique — five systematic upward biases the author argues inflate public token-usage trackers (supply-injected volume, substitution, subsidy, quality composition, and "base effects and disclosure selection" — the fifth label covers two distinct mechanisms; only the quality-composition bias is genuinely evidenced with cited data, per the close-read), plus the argument that all pre-Q2-2026 projections (including the author's own prior work) predate a "token-maximization → token-minimization" doctrinal shift and should be treated as upper bounds only.
- Five-scenario probability map (Section 8) spanning the demand-growth × pricing-stickiness × HBM-supply-branch space, each with a named narrative and assigned probability.
- A custom-silicon entrant model (Section 10.2) quantifying that removing NVIDIA's merchant margin (~20-35% of cost) does not remove the HBM/memory premium, because HBM is priced by "the same three suppliers under the same shortage" regardless of who designs the logic die.
- The LineShine LX2 case study (Section 10.1) — a Shenzhen-built, Huawei Armv9-based, domestic-HBM system (#1 on the June 2026 TOP500) as an existence proof that a sovereign compute stack can structurally decouple from the global memory shortage, at the cost of being generations behind on HBM (paper says it trails HBM4) and 3.6x mixed-precision uplift vs. 9-11x for GPU systems.
Methodology
Core model (Section 2, Eq. 1) computes $/PB as: accelerator price amortized over an assumed 4-year depreciation life and annualized hours, plus power draw × PUE × electricity cost, plus other opex — all divided by delivered bandwidth (TB/s) times a memory-bandwidth-utilization factor (η), converted to a per-petabyte basis. The author frames $/PB as the correct lower-level unit specifically for saturated, bandwidth-bound decode (i.e., the inference regime where the bottleneck is moving KV-cache/weights through memory, not raw FLOPs).
Vintage tracking uses straight-line 4-year amortization; once a fleet is fully depreciated it's modeled as operating near marginal (non-capital) cost — this is the mechanism underlying the "depreciation conveyor."
Scenarios (Section 8) are defined as regions of a three-dimensional space: token-demand growth rate, premium-pricing stickiness, and which of two HBM-supply branches obtains (a "base" branch normalizing by 2028, or a "shortage" branch with elevated pricing persisting to 2030).
Demand measurement (Section 9.1) cross-checks and discounts several public token-growth data points (Google's reported ~7x YoY to 3.2 quadrillion tokens/month as of May 2026; Microsoft Foundry's ~2.9x annualized rate in Q3 FY2026; OpenRouter's ~4-6x YoY spring-2025-to-May-2026; a Goldman Sachs 24x-by-2030 projection) against the five bias sources above to argue realized monetizable demand growth is materially lower than headline token-volume growth.
The author discloses (Acknowledgments) that the quantitative models, figures, and drafts were developed working with Anthropic's Claude (identified in-paper as "Fable 5"), with a caution that AI-assisted output may contain errors and that load-bearing figures were checked against cited primary sources but should be independently verified before use in investment, procurement, or policy decisions. The paper explicitly states it does not constitute investment advice.
Results — Quantitative Claims
Memory pricing & supply (Section 1, Fig. 1)
| Data point | Value |
|---|---|
| DRAM contract price rise, Q1 2026 vs Q4 2025 | ~90% |
| Memory share of accelerator bill-of-materials | 40-50% |
| HBM share of DRAM wafer output (Q1-Q2 2026, cited to TrendForce) | 23% |
| SK Hynix leadership warning (cited) | Shortage could persist to 2030 |
| Meaningful new fab capacity arrival | 2027-2028 |
DRAM → HBM wafer reallocation (Section 1)
The paper states the three major DRAM suppliers are "reallocating the large majority of wafer capacity toward HBM and server products" — treated as the structural root cause of the conventional-DRAM price surge, since HBM allocation directly competes with commodity-DRAM output on the same wafer starts. The paper does not name individual supplier capex commitments or break the reallocation down by company (Samsung/SK Hynix/Micron) — this is a gap for anyone trying to map the claim onto a specific print (see "Samsung Q2 relevance" below).
Entrant-incumbent cost gap ($/PB, Section 3 / Fig. 1)
| Year | Gap (entrant $/PB ÷ incumbent floor $/PB) | Notes |
|---|---|---|
| 2026 | 3.2x (entrant ≈ $0.174/PB; incumbent floor ≈ $0.054/PB) | Peak of current shortage |
| 2027 | 1.9-2.0x | Narrowest point — "best entry point" per the paper |
| 2029-30 | 3-4x | Re-widens as newer (B200-class) incumbent fleets amortize |
The paper's headline finding: this gap "never closes within the horizon" — new entrants are structurally always buying at spot prices while incumbents are running down already-amortized capital.
Capacity buildout (Section 7, Fig. 5)
- Installed capacity under announced plans grows ~2.6x by 2029.
- Solvency of that buildout requires ~2x/year token-demand growth sustained for four years, with sticky premium pricing.
- Baseline assumed efficiency gain: ~30%/year in bytes-per-token.
Training-cost bifurcation
- Luxury/frontier tier: $18-38B per frontier training run by 2030.
- Mass tier: previous-frontier-parity achievable via RL/distillation, falling toward ~$5M.
Vintage-breakeven analysis (Section 6, Fig. 4, Table-style breakdown)
Premium-pricing share needed to break even, by vintage and regime:
CORRECTED 2026-07-20 from the primary PDF — the figures below originally came from a structured web synthesis and were wrong in two ways that flip a position read. See the close-read companion
memory-scarcity-2607.07207-closeread.mdfor the full 16-cell grid. (1) The 2026 vintage's regime exposure was inverted: it is worse under STICKY pricing, not coupled. (2) 2028 and 2029 are separate vintages and materially different — lumping them hides that 2029 is nearly 2x as exposed as 2028.
| Vintage | Coupled pricing (routing arbitrage) | Sticky pricing (premiums hold) |
|---|---|---|
| 2026 | ~24.5% | ~31.3% → worst cell for this vintage (bought at peak hardware price) |
| 2027 | 7.5-11.1% → robust in all four cells | 7.5-11.1% → robust |
| 2028 (Stargate-class) | 21.1% in the shortage branch → unsustainable | survives |
| 2029 (Stargate-class) | 37.6% in the shortage branch → most exposed cell in the grid | survives |
Plausible-share band the paper judges these against: 10-20% (asserted, not derived).
Conclusion in-paper: routing success (efficiency/local-inference eating into centralized demand) breaks the 2028-29 Stargate-class commitments; routing failure breaks today's peak-price 2026 buyers. 2027 is the only vintage robust across both regimes. Note the structure driving this: P_mass is pinned to the incumbent depreciated floor plus 30% margin, so every new vintage is priced against a floor set by someone else's sunk capital — and the two regimes cross over (2026 coupled premium ~$0.49/PB sits above sticky's flat $0.40/PB; by 2029 coupled ~$0.20/PB is half of it), which is what produces the U-shape across vintages.
Custom-silicon entrant (Section 10.2, Table 3-4)
- Unconditional outcome distribution: 25% success / 34% mediocre / 41% loss. These are elicited SUBJECTIVE probabilities revised across five rounds, explicitly not model output (verified in the PDF, Table 3) — do not cite them as a derived result.
- Conditional on a downside stress case (bandwidth demand peaking ~2028): ~10% success / ~30% mediocre / ~60% loss.
- HBM is 45-50% of a custom accelerator's BOM, "priced by the same three suppliers under the same shortage" regardless of who owns the logic design — so a well-executed custom build reaches ~$0.082/PB (roughly merchant-GPU parity) [CORRECTED 2026-07-20 from the PDF; this file previously said $0.072/PB, which is actually the 2029 normalized-HBM new-build figure from Fig. 1, not the custom-build number] while the incumbent depreciated floor sits at $0.022-0.037/PB (2.2-3.3x cheaper).
- Mitigants (anchor demand, a secured HBM long-term agreement/LTA, or a non-HBM architecture) can roughly double success probability to 45-50%, and staged go/no-go capital gates cut deployed-capital loss probability to ~24%.
Scenario probabilities (Section 8)
| Scenario | Probability | Core mechanism |
|---|---|---|
| Rotating Landlord Oligopoly | 25% | Demand near threshold; incumbents with sunk/depreciated fleets (Meta/xAI-class) set mass-market prices from their amortized cost floor indefinitely. |
| Commoditization Crash | 25% | Token growth <1.7x/yr and/or pricing decouples; routing + local inference hollow out centralized demand as 2027-28 fab capacity lands — hurts whoever bought at peak vintages. |
| Jevons Absorption | 20% | Token growth ≥2.5x/yr sustained via agentic workloads; all vintages solvent, entrant-incumbent gap persists but matters less because raw access beats cost optimization. |
| System-Layer Re-differentiation | 18% | Sticky premium pricing holds; value migrates to orchestration/proprietary RL environments rather than raw inference capacity. |
| Geopolitical Bifurcation | 12% | Export controls extend bidirectionally; Western labs regain pricing power, China standardizes on open weights, sovereign compute capacity appreciates. |
LineShine LX2 (Section 10.1)
Shenzhen-built system, 40,960 Huawei Armv9 LX2 processors, ranked #1 on the June 2026 TOP500. Domestic HBM means its cost curve is "priced domestically and directed by the state, immune to the shortage inflating everyone else's capex" per the paper. Technical caveats acknowledged: mixed-precision uplift 3.6x vs. 9-11x for GPU systems; 52 GFLOPS/W (trails El Capitan); HBM generations behind HBM4. The paper frames it as evidence for the "sovereign, non-CUDA, HBM-integrated CPU path" being real at national scale, connecting to the Dongarra-Hoefler-Matsuoka "Do We Still Need GPUs?" debate (refs [17],[18]).
HBM vs. Conventional DRAM Allocation — What the Paper Does and Doesn't Say
The paper treats the DRAM→HBM wafer reallocation (Section 1) as the root cause of the broader DRAM price surge — i.e., conventional/commodity DRAM tightness is presented as a direct consequence of suppliers redirecting wafer starts toward HBM, not an independent phenomenon. It cites TrendForce for the "HBM at 23% of DRAM wafer output" figure but does not break this allocation shift down by named supplier, does not give capex figures for Samsung/SK Hynix/Micron individually, and does not analyze supplier-side profitability or pricing strategy — the shortage is modeled as an exogenous constraint feeding into AI-industry economics, not an endogenous strategic choice by the memory makers. This is the paper's most important limitation for anyone trying to read it directly against Samsung's numbers (see below).
Limitations (author-acknowledged)
- Pre/post regime-break risk (Section 9.3): all optimistic demand projections — including the author's own prior work — predate what the paper calls the "Q2 2026 doctrinal inversion from token maximization to token minimization," and should be treated as upper bounds until at least two quarters of post-break data are available. Only weeks of post-break data existed at time of writing.
- Model is directional, not precise: the $/PB framework is described as "intended for iteration rather than false precision," with conclusions stated to be robust to roughly ±20% parameter perturbation except where flagged (Appendix A).
- Realized cost variability: the paper cites a reference ([25]) showing effective cost on identical H100 hardware ranging from $0.21 to $15.25 per million output tokens depending on utilization and request-rate — i.e., the model's per-vintage cost floors are averages, not guarantees, across deployments.
- Policy fragility: the June 2026 suspension of a newly released frontier model is cited as evidence that premium supply can be policy-fragile (feeding the Geopolitical Bifurcation scenario) — a risk factor outside the model's core economic parameters.
- No supplier-level memory-maker analysis (see above) — the paper does not model Samsung/SK Hynix/Micron capacity-allocation decisions, margins, or capex individually; it treats the memory shortage as a single aggregate exogenous input.
- AI-assisted authorship disclosed — figures/models developed working with an AI system per the Acknowledgments; author states load-bearing figures were checked against primary sources but urges independent verification before use in investment/procurement/policy decisions. Explicitly not investment advice.
Why This Matters for the Desk
This is a demand-side/systems paper, not a supply-side memory-maker earnings model — its numbers describe how AI buyers' economics respond to memory scarcity, not how memory makers' margins or capex evolve. The load-bearing facts that translate most directly to reading Samsung's Q2 print: the ~90% Q1 2026 DRAM contract-price spike, the 40-50% memory share of accelerator BOM, the 23% HBM share of DRAM wafer output (TrendForce-sourced, aggregate not per-supplier), the "shortage to 2030" framing (SK Hynix-sourced), and the general thesis that new fab capacity doesn't land meaningfully until 2027-2028 — meaning the current pricing regime, from the buy side, is expected to persist through at least the next several quarters. The paper offers no independent view on Samsung's specific mix shift, margin trajectory, or capex guidance; it should be read as corroborating context for the shortage persisting, not as a source for Samsung-specific numbers.
Source: Memory Scarcity, Open Models, and the Restructuring of the AI Industry, 2026-2030, Satoshi Matsuoka, arXiv:2607.07207 [econ.GN] (cross-listed cs.AI, cs.AR, cs.CE, cs.PF), submitted July 8, 2026. 22 pages.