RLHF and RLVR

RLHF (reinforcement learning from human feedback) aligns a pretrained model to human preferences: humans rank model outputs, a reward model is fit to the rankings, and the policy is optimised against it (InstructGPT, 2022). Variants replace parts of the pipeline: Constitutional AI / RLAIF uses AI feedback guided by written principles (2022); DPO (2023) optimises directly on preference pairs without an explicit reward model.

RLVR (reinforcement learning with verifiable rewards) replaces the learned reward with a programmatic check, such as a unit test passing or a maths answer matching. The term was introduced in the Tulu 3 post-training recipe (2024), which names the method “Reinforcement Learning with Verifiable Rewards”. DeepSeek-R1 (2025) popularised reward-by-verification with GRPO (introduced in DeepSeekMath) to elicit long reasoning, which is why RLVR underpins reasoning-models.

Contrast

RLHFRLVR
Reward sourcelearned from human preferencesautomatic verifier
Good fortone, helpfulness, safetymaths, code, checkable tasks
Riskreward hacking, sycophancyonly works where answers can be checked

Related: low-rank-adaptation, ai-safety-and-governance, deepseek-family, large-language-model.

Sources