03HOW YOU RUN IT CHEAPLY, PER TOKEN· SURGING
Serving & Runtime
How a model is actually run, per token — the mechanics that turn capability into throughput. In scope: batching and scheduling, KV-cache management and eviction, quantization at serving time, speculative decoding, prefill/decode disaggregation, paged attention, routing, throughput and latency engineering. Out: what the model can do (models), what a finished task costs (inference-economics), the silicon it runs on (hardware).
15SOURCES
6CONCEPTS
0ENTITIES
SOURCE MIX
12 P2 R1 A0 N
ACTIVITY · 20W
No practice challenges available for this topic yet.
Challenges are generated from compiled research.