Off the ladder · Awaiting Connor
RLHF & Post-Training
Nathan LambertfreeSource ↗
Added 2026-08-06 at Connor’s direction and deliberately NOT gated on the memory rung: the author states no prior RL background is needed, so a dependency would be invented. It has no weekly slot either, and pretending otherwise would be the same error in the other direction.
0 / 13 video lectures
Progress
Pending Connor
Quota
Read it
Seat
Awaiting Connor
State
No weekly slot allocated. Draft Day decision.
Artifacts
Consume → do → output- Post-training objectives and their failure modes, spacedNot yet — nothing has landed for this slot.
Feeds
What closing this sharpensThe reading
Sections of the catalogue this unit draws on- §6Reinforcement Learning and RLVR25 papersModels
The papers
The rows behind the sections above- GDPO: Group Reward-Decoupled Normalization Policy Optimization for Multi-Reward RL Optimization ↗
- Your Group-Relative Advantage Is Biased ↗
- Reasoning Models Generate Societies of Thought ↗
- Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability ↗
- Harder Is Better: Boosting Mathematical Reasoning via Difficulty-Aware GRPO and Multi-Aspect Question Reformulation ↗
- Reinforcement Learning via Self-Distillation ↗
- Expanding the Capabilities of Reinforcement Learning via Text Feedback ↗
- iGRPO: Self-Feedback-Driven LLM Reasoning ↗
- Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning ↗
- Experiential Reinforcement Learning ↗
- Discovering Implicit Large Language Model Alignment Objectives ↗
- OpenClaw-RL: Train Any Agent Simply by Talking ↗
- Rethinking Token-Level Policy Optimization for Multimodal Chain-of-Thought ↗
- Self-Distilled RLVR ↗
- Target Policy Optimization ↗
- FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling ↗
- SPPO: Sequence-Level PPO for Long-Horizon Reasoning Tasks ↗
- Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges ↗
- SEIF: Self-Evolving Reinforcement Learning for Instruction Following ↗
- Process Rewards with Learned Reliability ↗
- Self-Distilled Agentic Reinforcement Learning ↗
- You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories ↗★
- DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning ↗
- Agent Explorative Policy Optimization for Multimodal Agentic Reasoning ↗
- Trust Region On-Policy Distillation ↗