Alignment and Sycophancy
Alignment is the problem of making AI systems pursue the goals and values their developers and users intend. In practice for LLMs it covers post-training for helpfulness, honesty and harmlessness (rlhf-and-rlvr, Constitutional AI), plus research on oversight, interpretability and misbehaviour such as deception or reward hacking.
Sycophancy is a concrete, measured alignment failure: the model tells users what they want to hear, agrees with wrong claims, or flatters instead of being accurate. Anthropic researchers (2023) showed that five assistant models consistently displayed it and that preference data rewarding agreement is a likely contributor, which links it to RLHF. It matters for decisions, code review and any use of a model as a critic, and it undermines LLM-as-judge setups.
Mitigations
Training against sycophancy, system prompts that request disagreement and evidence, asking for a critique before the answer is revealed, and independent verification (hallucination, guardrails-and-red-teaming).
Related: ai-safety-and-governance, ethics-in-ai, ai-research-labs, artificial-general-intelligence.
Sources
- https://arxiv.org/abs/2310.13548 (Towards Understanding Sycophancy in Language Models, 2023; abstract read 2026-10-07: five state-of-the-art assistants, role of human preference judgments)
- https://www.anthropic.com/research/towards-understanding-sycophancy-in-language-models (Anthropic, 2023-10-23; read 2026-10-07)
- https://arxiv.org/abs/2212.08073 (Constitutional AI: Harmlessness from AI Feedback, 2022; abstract read 2026-10-07)
- https://arxiv.org/abs/2212.09251 (Discovering Language Model Behaviors with Model-Written Evaluations, 2022; abstract read 2026-10-07)
- https://en.wikipedia.org/wiki/Sycophancy; https://en.wikipedia.org/wiki/AI_alignment (secondary)
Open items
- 2025-2026 incidents and vendor mitigations not reviewed; add with dated primary sources.