RLHF and RLVR
RLHF (reinforcement learning from human feedback) aligns a pretrained model to human preferences: humans rank model outputs, a reward model is fit to the rankings, and the policy is optimised against it (InstructGPT, 2022). Variants replace parts of the pipeline: Constitutional AI / RLAIF uses AI feedback guided by written principles (2022); DPO (2023) optimises directly on preference pairs without an explicit reward model.
RLVR (reinforcement learning with verifiable rewards) replaces the learned reward with a programmatic check, such as a unit test passing or a maths answer matching. The term was introduced in the Tulu 3 post-training recipe (2024), which names the method “Reinforcement Learning with Verifiable Rewards”. DeepSeek-R1 (2025) popularised reward-by-verification with GRPO (introduced in DeepSeekMath) to elicit long reasoning, which is why RLVR underpins reasoning-models.
Contrast
| RLHF | RLVR | |
|---|---|---|
| Reward source | learned from human preferences | automatic verifier |
| Good for | tone, helpfulness, safety | maths, code, checkable tasks |
| Risk | reward hacking, sycophancy | only works where answers can be checked |
Related: low-rank-adaptation, ai-safety-and-governance, deepseek-family, large-language-model.
Sources
- https://arxiv.org/abs/2203.02155 (InstructGPT, 2022)
- https://arxiv.org/abs/2212.08073 (Constitutional AI, 2022)
- https://arxiv.org/abs/2305.18290 (DPO, 2023)
- https://arxiv.org/abs/2411.15124 (Tulu 3, 2024)
- https://arxiv.org/abs/2402.03300 (DeepSeekMath/GRPO, 2024)
- https://arxiv.org/abs/2501.12948 (DeepSeek-R1, 2025)
- https://en.wikipedia.org/wiki/Reinforcement_learning_from_human_feedback (secondary)
- Titles and dates confirmed via arXiv API 2026-10-07; Tulu 3 (RLVR coined), DeepSeekMath (GRPO introduced), DeepSeek-R1 (adopts GRPO) and InstructGPT (reward model + PPO) confirmed in the PDF text 2026-10-07.