← The reading
The bibliography
§4 · Sparse Attention and Long Context
Inference Engineering · inference-bench
CAPTURED
9 On the list1 Starred0 In the Atlas9 To read
Rungs
Where this section moves a number- Inference EngineeringRung 3
- Efficiency in LLMsRung 3
The pick
The source author's must-read for this sectionThe papers
9 papers- Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection ↗
- IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse ↗
- Mixture-of-Depths Attention ↗
- TriAttention: Efficient Long Reasoning with Trigonometric KV Compression ↗
- Sessa: Selective State Space Attention ↗
- Contexts Are Never Long Enough: Structured Reasoning for Scalable Question Answering over Long Document Sets ↗
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence ↗★
- Long Context Pre-Training with Lighthouse Attention ↗
- Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps ↗
What counts as read
A paper is on this list because someone worth reading put it there. That is a pointer, not a claim: it counts as read only once it has a close-read file in kb/<topic>/raw/, which is what a close-read link on a row means. There is deliberately nothing to tick off here — the study desk is the only writer of study state, and a second way to mark something done is a second version of the truth.