← Study
Off the ladder · Awaiting Connor

RLHF & Post-Training

Nathan LambertfreeSource ↗

Added 2026-08-06 at Connor’s direction and deliberately NOT gated on the memory rung: the author states no prior RL background is needed, so a dependency would be invented. It has no weekly slot either, and pretending otherwise would be the same error in the other direction.

0 / 13 video lectures
Progress
Pending Connor
Quota
Read it
Seat
Awaiting Connor
State

No weekly slot allocated. Draft Day decision.

Artifacts

Consume → do → output
NBNotebookLABLabGMEGameBTLBottle0/1
  • Post-training objectives and their failure modes, spacedNot yet — nothing has landed for this slot.

Feeds

What closing this sharpens

The reading

Sections of the catalogue this unit draws on
  • §6Reinforcement Learning and RLVR25 papersModels

The papers

The rows behind the sections above

§6 · Reinforcement Learning and RLVR

25 papers · ★1 · 0 in the AtlasSection page