← The reading
The bibliography
§2 · Efficient Training and Scaling
CAPTURED
15 On the list1 Starred0 In the Atlas15 To read
Rungs
Where this section moves a numberThe pick
The source author's must-read for this sectionThe papers
15 papers- Self-Distillation Enables Continual Learning ↗
- Shaping Capabilities with Token-Level Data Filtering ↗
- Self-Improving Pretraining: Using Post-Trained Models to Pretrain Better Models ↗
- BLOCK-EM: Preventing Emergent Misalignment via Latent Blocking ↗
- OPUS: Towards Efficient and Principled Data Selection in LLM Pre-training in Every Iteration ↗
- Test-Time Training with KV Binding Is Secretly Linear Attention ↗
- Progressive Residual Warmup for Language Model Pretraining ↗
- Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights ↗
- MiCA Learns More Knowledge Than LoRA and Full Fine-Tuning ↗
- In-Place Test-Time Training ↗
- Hybrid Policy Distillation for LLMs ↗
- Large Language Models Explore by Latent Distilling ↗
- Efficient Training on Multiple Consumer GPUs with RoundPipe ↗
- Efficient Pre-Training with Token Superposition ↗★
- MinT: Managed Infrastructure for Training and Serving Millions of LLMs ↗
What counts as read
A paper is on this list because someone worth reading put it there. That is a pointer, not a claim: it counts as read only once it has a close-read file in kb/<topic>/raw/, which is what a close-read link on a row means. There is deliberately nothing to tick off here — the study desk is the only writer of study state, and a second way to mark something done is a second version of the truth.