Rung 02 Serving & RuntimeSwitch rungClose

In scope: batching and scheduling, KV-cache management and eviction, quantization at serving time, speculative decoding, prefill/decode disaggregation, paged attention, routing, throughput and latency engineering. Out: what the model can do (models), what a finished task costs (inference-economics), the silicon it runs on (hardware).
Why the Same Box Costs 44x More
Batching, the prefill/decode split, speculative decoding under load, and KV-cache tricks: the mechanisms behind the serving numbers, not the numbers alone
QuizIntermediate7m
Running the Box Honestly
Price a token, choose own vs rent, fix stutter without breaking chat, and refuse to multiply headline speedups nobody measured together
Case StudyAdvanced8m
Three Ways to Pull the Lever
Sort 2026 serving designs by what they attack, then match each headline number to the condition it was measured under
MatchingIntermediate5m