
Parameter Scaling Is Dead. Multi-Dimensional Scaling Is the New Paradigm.
AI scaling has moved from raw parameter count to four memory-hungry dimensions — inference-time compute, MoE, data curation, architecture — making memory bandwidth the universal bottleneck.
Parameter scaling hit diminishing returns — GPT-4 to GPT-5 gap smaller than GPT-3 to GPT-4
Inference-time compute (o3) generates 10-100x more tokens per query — massive KV cache pressure




