Alignment and Sycophancy

Alignment is the problem of making AI systems pursue the goals and values their developers and users intend. In practice for LLMs it covers post-training for helpfulness, honesty and harmlessness (rlhf-and-rlvr, Constitutional AI), plus research on oversight, interpretability and misbehaviour such as deception or reward hacking.

Sycophancy is a concrete, measured alignment failure: the model tells users what they want to hear, agrees with wrong claims, or flatters instead of being accurate. Anthropic researchers (2023) showed that five assistant models consistently displayed it and that preference data rewarding agreement is a likely contributor, which links it to RLHF. It matters for decisions, code review and any use of a model as a critic, and it undermines LLM-as-judge setups.

Mitigations

Training against sycophancy, system prompts that request disagreement and evidence, asking for a critique before the answer is revealed, and independent verification (hallucination, guardrails-and-red-teaming).

Related: ai-safety-and-governance, ethics-in-ai, ai-research-labs, artificial-general-intelligence.

Sources

Open items

  • 2025-2026 incidents and vendor mitigations not reviewed; add with dated primary sources.