Rung 4 · Gated
CS336 — Language Models from Scratch
StanfordfreeSource ↗
Building the thing comes after understanding the machine it runs on. The assignments are the point — the credibility here is a working model, not a completed lecture list.
0 / 19 lectures
Progress
Allocated
Quota
Read it
Seat
Gated
State
Takes the technical slot when rung 3 closes.
Gate
Follows rung 3. A PyTorch readiness check precedes Assignment 1; if fluency is thin, a nanoGPT warm-up comes first.
Artifacts
Consume → do → output- I built a language model from scratch — the build, in publicNot yet — nothing has landed for this slot.
- The from-scratch model, runningNot yet — nothing has landed for this slot.
Feeds
What closing this sharpensThe reading
Sections of the catalogue this unit draws on- §1Architecture and Model Design22 papersModels
- §2Efficient Training and Scaling15 papersModels · Data
The papers
The rows behind the sections above- Deep Delta Learning ↗
- MiMo-V2-Flash Technical Report ↗
- Ministral 3 ↗
- Scaling Embeddings Outperforms Scaling Experts in Language Models ↗
- LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs ↗
- ERNIE 5.0 Technical Report ↗
- ViT-5: Vision Transformers for the Mid-2020s ↗
- Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters ↗
- Nanbeige4.1-3B: A Small General Model That Reasons, Aligns, and Acts ↗
- Symmetry in Language Statistics Shapes the Geometry of Model Representations ↗
- GLM-5: From Vibe Coding to Agentic Engineering ↗
- Arcee Trinity Large Technical Report ↗
- The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks ↗
- Tiny Aya: Bridging Scale and Multilingual Depth ↗
- Attention Residuals ↗
- Mamba-3: Improved Sequence Modeling Using State Space Principles ↗
- Attention to Mamba: A Recipe for Cross-Architecture Distillation ↗
- Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning ↗★
- ZAYA1-8B Technical Report ↗
- Delta Attention Residuals ↗
- Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention ↗
- The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence ↗
- Self-Distillation Enables Continual Learning ↗
- Shaping Capabilities with Token-Level Data Filtering ↗
- Self-Improving Pretraining: Using Post-Trained Models to Pretrain Better Models ↗
- BLOCK-EM: Preventing Emergent Misalignment via Latent Blocking ↗
- OPUS: Towards Efficient and Principled Data Selection in LLM Pre-training in Every Iteration ↗
- Test-Time Training with KV Binding Is Secretly Linear Attention ↗
- Progressive Residual Warmup for Language Model Pretraining ↗
- Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights ↗
- MiCA Learns More Knowledge Than LoRA and Full Fine-Tuning ↗
- In-Place Test-Time Training ↗
- Hybrid Policy Distillation for LLMs ↗
- Large Language Models Explore by Latent Distilling ↗
- Efficient Training on Multiple Consumer GPUs with RoundPipe ↗
- Efficient Pre-Training with Token Superposition ↗★
- MinT: Managed Infrastructure for Training and Serving Millions of LLMs ↗