Test-time compute models like o3 and Claude's extended thinking generate massive KV caches that consume 40GB or more of HBM per long-context prompt. This is the single biggest bottleneck in AI inference today — not compute, not power, but memory. Google's TurboQuant demonstrates 6x KV cache compression with minimal quality loss, applied directly at inference time without retraining. NVIDIA's NVFP4 pushes further with hardware-accelerated 4-bit KV cache quantization. Software-layer memory optimization is shipping faster than HBM4 hardware — and the inference cost collapse it enables will be the most important driver of AI democratization in 2026-2027. The implications compound: 6x compression means either 6x more concurrent users per GPU or dramatically longer contexts without adding hardware. Companies building inference-as-a-service on compressed KV cache architectures will capture margins that GPU-heavy providers cannot match. Long-context models become viable at scale, and classical RAG — the retrieval overhead that adds latency and complexity — becomes unnecessary for most use cases. This is not a theoretical improvement. It is already shipping in production inference systems and will define the cost structure of AI for the next two years.
Sources rebuilt primary-first (catalogue II item 69): 5 primary, 1 secondary kept.
source: docs/plans/markets-work-catalogue-2026-09-07.md
