← The reading
The bibliography
§6 · Reinforcement Learning and RLVR
Reasoning book ch6–8 · RLHF book
CAPTURED
25 On the list1 Starred0 In the Atlas25 To read
Rungs
Where this section moves a number- RLHF & Post-TrainingOff the ladder
- Build a Reasoning Model From ScratchOff the ladder
The pick
The source author's must-read for this sectionThe papers
25 papers- GDPO: Group Reward-Decoupled Normalization Policy Optimization for Multi-Reward RL Optimization ↗
- Your Group-Relative Advantage Is Biased ↗
- Reasoning Models Generate Societies of Thought ↗
- Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability ↗
- Harder Is Better: Boosting Mathematical Reasoning via Difficulty-Aware GRPO and Multi-Aspect Question Reformulation ↗
- Reinforcement Learning via Self-Distillation ↗
- Expanding the Capabilities of Reinforcement Learning via Text Feedback ↗
- iGRPO: Self-Feedback-Driven LLM Reasoning ↗
- Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning ↗
- Experiential Reinforcement Learning ↗
- Discovering Implicit Large Language Model Alignment Objectives ↗
- OpenClaw-RL: Train Any Agent Simply by Talking ↗
- Rethinking Token-Level Policy Optimization for Multimodal Chain-of-Thought ↗
- Self-Distilled RLVR ↗
- Target Policy Optimization ↗
- FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling ↗
- SPPO: Sequence-Level PPO for Long-Horizon Reasoning Tasks ↗
- Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges ↗
- SEIF: Self-Evolving Reinforcement Learning for Instruction Following ↗
- Process Rewards with Learned Reliability ↗
- Self-Distilled Agentic Reinforcement Learning ↗
- You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories ↗★
- DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning ↗
- Agent Explorative Policy Optimization for Multimodal Agentic Reasoning ↗
- Trust Region On-Policy Distillation ↗
What counts as read
A paper is on this list because someone worth reading put it there. That is a pointer, not a claim: it counts as read only once it has a close-read file in kb/<topic>/raw/, which is what a close-read link on a row means. There is deliberately nothing to tick off here — the study desk is the only writer of study state, and a second way to mark something done is a second version of the truth.