← Study
Off the ladder · Ready

Build a Reasoning Model From Scratch

Sebastian RaschkapaidSource ↗

Runs alongside the ladder rather than on it — eight chapters that each end at a keyboard. It is the unit most likely to earn a genuine “ran it” seat first, because the first five chapters need nothing but a laptop.

0 / 8 chapters
Progress
Allocated
Quota
Read it
Seat
Ready
State

Runs alongside the technical slot.

Environment

Setup is a real gate

partial~/dev/study/verified

Local setup complete — Qwen3-0.6B generating at 20.9 tok/s on MPS. Chapters 6–8 need a rented GPU (Modal).

Evidence

Measured — it does not move the seat
  • On the record — Local environment generates at 20.9 tokens/second on Apple MPS (Qwen3-0.6B). · Generation benchmark on the working setup at ~/dev/study/. ·

Segments

The named subset — the scope above is the count
RefTitleDoneArtifactsNote
ch1Understanding reasoning models·NBLABGMEBTL0/0
ch2Text generation with pretrained LLMs·NBLABGMEBTL0/0
ch3Model evaluation·NBLABGMEBTL0/0
ch4Inference-time scaling·NBLABGMEBTL0/1
ch5Self-refinement·NBLABGMEBTL0/0
ch6RL training with GRPO·NBLABGMEBTL0/0NEEDS — Needs a rented GPU (Modal).
ch7Advanced GRPO·NBLABGMEBTL0/0NEEDS — Needs a rented GPU (Modal).
ch8Distillation·NBLABGMEBTL0/0NEEDS — Needs a rented GPU (Modal).

Artifacts

Consume → do → output
NBNotebookLABLabGMEGameBTLBottle0/3
  • Reasoning models, measured on one small enough to watchNot yet — nothing has landed for this slot.
  • A self-refinement loop, playableNot yet — nothing has landed for this slot.
  • GRPO mechanics and where it breaks, spacedNot yet — nothing has landed for this slot.

Feeds

What closing this sharpens

The reading

Sections of the catalogue this unit draws on
  • §5Reasoning and Test-Time Compute15 papersModels · Inference economics
  • §6Reinforcement Learning and RLVR25 papersModels

The papers

The rows behind the sections above

§6 · Reinforcement Learning and RLVR

25 papers · ★1 · 0 in the AtlasSection page