Rung 3 · Gated
Inference Engineering
Philip Kiely · Baseten · 2026freeSource ↗
Inference as a SYSTEM — serving, batching, KV cache, the production cost surface. It sits above the memory rung because KV-cache and batching arithmetic IS memory-hierarchy arithmetic: taken first, cost-per-token is a number you look up; taken here, it is a number you can derive and therefore contest.
0 / 8 chapters
Progress
Allocated
Quota
Read it
Seat
Gated
State
Second technical slot, once the memory rung closes.
Gate
Opens when the memory rung (L2a, L9, L12–L15) closes — the hardware has to be under it.
Segments
The named subset — the scope above is the count| Ref | Title | Done | Artifacts | Note |
|---|---|---|---|---|
| ch1 | Inference | · | NBLABGMEBTL0/0 | |
| ch2 | Prerequisites | · | NBLABGMEBTL0/0 | |
| ch3 | Models | · | NBLABGMEBTL0/0 | Arithmetic intensity — pairs with Comp Arch L9. |
| ch4 | Hardware | · | NBLABGMEBTL0/0 | |
| ch5 | Software | · | NBLABGMEBTL0/0 | |
| ch6 | Techniques | · | NBLABGMEBTL0/0 | |
| ch7 | Modalities | · | NBLABGMEBTL0/0 | |
| ch8 | Production | · | NBLABGMEBTL0/2 |
Artifacts
Consume → do → output- The token-price → task-price bridgeNot yet — nothing has landed for this slot.
- The memory roofline, verifiedEarmarked for inference-bench — building on the workshop; not counted until it ships there.
- ops:byte and KV-cache formulas, spacedNot yet — nothing has landed for this slot.
Feeds
What closing this sharpensThe reading
Sections of the catalogue this unit draws on- §3Inference Efficiency and KV Cache8 papersServing · Inference economics
- §4Sparse Attention and Long Context9 papersServing · Models
The papers
The rows behind the sections above- Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery ↗
- Efficient Remote KV Cache Reuse with GPU-Native Video Codec ↗
- SpargeAttention2: Trainable Sparse Attention via Hybrid Top-K+Top-P Masking and Distillation Fine-Tuning ↗
- Decoding as Optimisation on the Probability Simplex: From Top-K to Top-P (Nucleus) to Best-of-K Samplers ↗
- FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling ↗★
- LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation ↗
- Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning ↗
- A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints ↗
- Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection ↗
- IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse ↗
- Mixture-of-Depths Attention ↗
- TriAttention: Efficient Long Reasoning with Trigonometric KV Compression ↗
- Sessa: Selective State Space Attention ↗
- Contexts Are Never Long Enough: Structured Reasoning for Scalable Question Answering over Long Document Sets ↗
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence ↗★
- Long Context Pre-Training with Lighthouse Attention ↗
- Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps ↗