← The reading
The bibliography
§3 · Inference Efficiency and KV Cache
Inference Engineering · inference-bench
CAPTURED
8 On the list1 Starred0 In the Atlas8 To read
Rungs
Where this section moves a number- Inference EngineeringRung 3
- Efficiency in LLMsRung 3
The pick
The source author's must-read for this sectionThe papers
8 papers- Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery ↗
- Efficient Remote KV Cache Reuse with GPU-Native Video Codec ↗
- SpargeAttention2: Trainable Sparse Attention via Hybrid Top-K+Top-P Masking and Distillation Fine-Tuning ↗
- Decoding as Optimisation on the Probability Simplex: From Top-K to Top-P (Nucleus) to Best-of-K Samplers ↗
- FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling ↗★
- LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation ↗
- Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning ↗
- A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints ↗
What counts as read
A paper is on this list because someone worth reading put it there. That is a pointer, not a claim: it counts as read only once it has a close-read file in kb/<topic>/raw/, which is what a close-read link on a row means. There is deliberately nothing to tick off here — the study desk is the only writer of study state, and a second way to mark something done is a second version of the truth.