Rung 3 · Gated
Efficiency in LLMs
Alex Smola · companion to rung 3LINK PENDING
Interleaved with the chapter above rather than sequenced after it. The same economics from the modelling side — the second angle is what turns a chapter into a model you can argue with.
0 / 6 sections
Progress
Allocated
Quota
Read it
Seat
Gated
State
Interleaved with the rung-3 slot.
Gate
Shares rung 3 with Inference Engineering and opens with it.
Artifacts
Consume → do → output- Efficiency levers and what each one costs, spacedNot yet — nothing has landed for this slot.
Feeds
What closing this sharpensThe reading
Sections of the catalogue this unit draws on- §3Inference Efficiency and KV Cache8 papersServing · Inference economics
- §4Sparse Attention and Long Context9 papersServing · Models
The papers
The rows behind the sections above- Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery ↗
- Efficient Remote KV Cache Reuse with GPU-Native Video Codec ↗
- SpargeAttention2: Trainable Sparse Attention via Hybrid Top-K+Top-P Masking and Distillation Fine-Tuning ↗
- Decoding as Optimisation on the Probability Simplex: From Top-K to Top-P (Nucleus) to Best-of-K Samplers ↗
- FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling ↗★
- LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation ↗
- Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning ↗
- A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints ↗
- Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection ↗
- IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse ↗
- Mixture-of-Depths Attention ↗
- TriAttention: Efficient Long Reasoning with Trigonometric KV Compression ↗
- Sessa: Selective State Space Attention ↗
- Contexts Are Never Long Enough: Structured Reasoning for Scalable Question Answering over Long Document Sets ↗
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence ↗★
- Long Context Pre-Training with Lighthouse Attention ↗
- Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps ↗