LONGActiveMed3Y

Inference Costs Will Fall 90% by 2028 — The Memory Stack Is the Mechanism

By MenFem Editorial·AI Infrastructure·13 April 2026·Methodology·
ai-infrastructurememory-over-computeinferencecost-collapsecreator-class
Share
Inference Costs Will Fall 90% by 2028 — The Memory Stack Is the Mechanism

Key Points

  • Four independent memory stack improvements compound: HBM4 (2x) × KV compression (4-6x) × optical (70% energy) × architecture (2-10x)
  • Even three of four layers delivering = 5-8x cost reduction by 2028
  • Inference cost collapse is the mechanism of creator-class wealth transfer
  • Sub-$1 per million tokens for 70B+ models is the milestone to watch
  • Hyperscaler capex shift from training to inference infrastructure validates the thesis

The combination of four independent improvements in the AI memory stack will compound to deliver approximately 10x reduction in per-token inference cost by 2028. Each layer delivers 2-6x improvement independently — together, they multiply. Layer 1: HBM4 hardware delivers 2x bandwidth improvement, meaning fewer memory chips per inference server and lower hardware cost per query. Layer 2: KV cache compression (TurboQuant and successors) delivers 4-6x reduction in memory consumption per inference pass, allowing more concurrent requests per GPU. Layer 3: Optical interconnect (CPO) reduces data movement energy by 70%, cutting operating costs for every inference request. Layer 4: Architecture evolution (hybrid SSM-Transformer models) could deliver an additional 2-10x memory efficiency improvement if the quality gap closes. This is not a forecast about any single technology — it is a thesis about compounding improvements across a stack. Even if only three of the four layers deliver, the result is a 5-8x cost reduction. The implications connect directly to the creator-class transfer thesis: cheap inference means universal access to AI capabilities. When running GPT-4-class inference costs less than a Google search, the creator class gains tools that only enterprises could afford in 2024. This is the mechanism of wealth transfer from institutions to individuals.

Research Log

Sources rebuilt primary-first (catalogue II item 69): 4 primary, 0 secondary kept.

source: docs/plans/markets-work-catalogue-2026-09-07.md

Re-underwritten 2026-09-06. The mechanism (memory) is doing the opposite of what the call needs in 2026: DRAM contract prices rose two quarters running and the flagship list price went up, not down. The claim is about cost per task by 2028, which no instrument on the site measures since the Token Price Index retired; the price leg is now hand-kept in research datasets/model-list-prices.csv. KEEP, conviction to MEDIUM. Grading date: the claim says "by 2028" — Connor ruled the title wins over the THREE_YEARS enum; grade on 2028-12-31 (gradeBy override follows once the column ships).

source: docs/plans/call-reviews-2026-09-06-thematic.md

Sources filled from the shelf and the 2026-09-06 review docs (catalogue item 3). 4 entries.

source: docs/plans/markets-work-catalogue-2026-09-06.md

Bull Case

By 2028, running GPT-4-class inference costs less than a Google search. AI becomes a utility like electricity. The creator class gains access to capabilities that only enterprises could afford in 2024. New business models emerge that were impossible at 2025 inference costs.

Bear Case

Cost reductions are real but Jevons paradox means prices stay elevated as demand absorbs capacity. Only hyperscalers benefit from lower unit costs because they can afford the hardware capex. Small creators and startups remain priced out of frontier model access.

What would prove this wrong

The reason for holding this stops being true if the four layers stop being independent — if one constraint gates two or more of them, they add rather than multiply and the compounding argument, which is the whole thesis, fails. It also stops being true if the savings are captured rather than passed through: providers holding price while their costs fall would mean the technology thesis was right and the cost-collapse thesis — the one actually published — was wrong.

Catalysts

Sub-$1/M token inference for 70B+ modelsMacro

When inference providers offer sub-$1 per million tokens for 70B+ parameter models, the cost collapse is confirmed and the market re-prices accordingly.

Hyperscaler capex shift from training to inferenceMacro

Watch Meta, Google, Microsoft capex breakdowns. The shift from training infrastructure to inference infrastructure validates that inference cost is the binding business constraint.

Risk factors

Jevons paradox absorbs cost savingsMedium

If inference demand grows faster than cost reductions, per-unit prices may not fall as fast as the technology enables. Providers may capture margin rather than pass savings to users.

Geopolitical disruption to HBM supplyMedium

HBM manufacturing is concentrated in South Korea (SK Hynix, Samsung). Geopolitical risk in the Korean peninsula could constrain supply at a critical moment.

Conviction

Conviction History

HighMed

6 Sept 2026

Review 2026-09-06: PARTLY FIRED on both legs — layers 1–2 share a memory-supply gate (DRAM contract prices +58–63% QoQ in Q2 after +90–95% in Q1) and GPT-6 Astra listed at 2.5× its predecessor on 3 Sep. Conviction HIGH → MEDIUM.

Key Metrics

HBM4 Bandwidth Improvement
2x
KV Cache Compression
4-6x
CPO Energy Reduction
70%
Combined Target Reduction
~10x by 2028
See all AI Infrastructure calls →