Off the ladder · Ready
Build a Reasoning Model From Scratch
Sebastian RaschkapaidSource ↗
Runs alongside the ladder rather than on it — eight chapters that each end at a keyboard. It is the unit most likely to earn a genuine “ran it” seat first, because the first five chapters need nothing but a laptop.
0 / 8 chapters
Progress
Allocated
Quota
Read it
Seat
Ready
State
Runs alongside the technical slot.
Environment
Setup is a real gatepartial~/dev/study/verified
Local setup complete — Qwen3-0.6B generating at 20.9 tok/s on MPS. Chapters 6–8 need a rented GPU (Modal).
Evidence
Measured — it does not move the seat- On the record — Local environment generates at 20.9 tokens/second on Apple MPS (Qwen3-0.6B). · Generation benchmark on the working setup at ~/dev/study/. ·
Segments
The named subset — the scope above is the count| Ref | Title | Done | Artifacts | Note |
|---|---|---|---|---|
| ch1 | Understanding reasoning models | · | NBLABGMEBTL0/0 | |
| ch2 | Text generation with pretrained LLMs | · | NBLABGMEBTL0/0 | |
| ch3 | Model evaluation | · | NBLABGMEBTL0/0 | |
| ch4 | Inference-time scaling | · | NBLABGMEBTL0/1 | |
| ch5 | Self-refinement | · | NBLABGMEBTL0/0 | |
| ch6 | RL training with GRPO | · | NBLABGMEBTL0/0 | NEEDS — Needs a rented GPU (Modal). |
| ch7 | Advanced GRPO | · | NBLABGMEBTL0/0 | NEEDS — Needs a rented GPU (Modal). |
| ch8 | Distillation | · | NBLABGMEBTL0/0 | NEEDS — Needs a rented GPU (Modal). |
Artifacts
Consume → do → output- Reasoning models, measured on one small enough to watchNot yet — nothing has landed for this slot.
- A self-refinement loop, playableNot yet — nothing has landed for this slot.
- GRPO mechanics and where it breaks, spacedNot yet — nothing has landed for this slot.
Feeds
What closing this sharpensThe reading
Sections of the catalogue this unit draws on- §5Reasoning and Test-Time Compute15 papersModels · Inference economics
- §6Reinforcement Learning and RLVR25 papersModels
The papers
The rows behind the sections above- PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning ↗
- Learning to Reason in 13 Parameters ↗
- InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning ↗
- Learning to Self-Verify Makes Language Models Better Reasoners ↗
- Does Your Reasoning Model Implicitly Know When to Stop Thinking? ↗
- Reasoning Models Struggle to Control Their Chains of Thought ↗
- Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs ↗
- Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation ↗
- Test-Time Scaling Makes Overtraining Compute-Optimal ↗★
- Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets ↗
- When Can LLMs Learn to Reason with Weak Supervision? ↗
- AI Co-Mathematician: Accelerating Mathematicians with Agentic AI ↗
- LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling ↗
- Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling ↗
- Share More, Search Less: Collaborative Parallel Thinking for Efficient Test-Time Scaling ↗
- GDPO: Group Reward-Decoupled Normalization Policy Optimization for Multi-Reward RL Optimization ↗
- Your Group-Relative Advantage Is Biased ↗
- Reasoning Models Generate Societies of Thought ↗
- Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability ↗
- Harder Is Better: Boosting Mathematical Reasoning via Difficulty-Aware GRPO and Multi-Aspect Question Reformulation ↗
- Reinforcement Learning via Self-Distillation ↗
- Expanding the Capabilities of Reinforcement Learning via Text Feedback ↗
- iGRPO: Self-Feedback-Driven LLM Reasoning ↗
- Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning ↗
- Experiential Reinforcement Learning ↗
- Discovering Implicit Large Language Model Alignment Objectives ↗
- OpenClaw-RL: Train Any Agent Simply by Talking ↗
- Rethinking Token-Level Policy Optimization for Multimodal Chain-of-Thought ↗
- Self-Distilled RLVR ↗
- Target Policy Optimization ↗
- FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling ↗
- SPPO: Sequence-Level PPO for Long-Horizon Reasoning Tasks ↗
- Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges ↗
- SEIF: Self-Evolving Reinforcement Learning for Instruction Following ↗
- Process Rewards with Learned Reliability ↗
- Self-Distilled Agentic Reinforcement Learning ↗
- You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories ↗★
- DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning ↗
- Agent Explorative Policy Optimization for Multimodal Agentic Reasoning ↗
- Trust Region On-Policy Distillation ↗