The combination of four independent improvements in the AI memory stack will compound to deliver approximately 10x reduction in per-token inference cost by 2028. Each layer delivers 2-6x improvement independently — together, they multiply. Layer 1: HBM4 hardware delivers 2x bandwidth improvement, meaning fewer memory chips per inference server and lower hardware cost per query. Layer 2: KV cache compression (TurboQuant and successors) delivers 4-6x reduction in memory consumption per inference pass, allowing more concurrent requests per GPU. Layer 3: Optical interconnect (CPO) reduces data movement energy by 70%, cutting operating costs for every inference request. Layer 4: Architecture evolution (hybrid SSM-Transformer models) could deliver an additional 2-10x memory efficiency improvement if the quality gap closes. This is not a forecast about any single technology — it is a thesis about compounding improvements across a stack. Even if only three of the four layers deliver, the result is a 5-8x cost reduction. The implications connect directly to the creator-class transfer thesis: cheap inference means universal access to AI capabilities. When running GPT-4-class inference costs less than a Google search, the creator class gains tools that only enterprises could afford in 2024. This is the mechanism of wealth transfer from institutions to individuals.
Sources rebuilt primary-first (catalogue II item 69): 4 primary, 0 secondary kept.
source: docs/plans/markets-work-catalogue-2026-09-07.md
