Hallucination
A hallucination is fluent model output that is false or unsupported by the provided source: invented facts, citations, quotes, APIs or file paths. In the author’s view it is tied to how next-token prediction works rather than being a bug with one fix.
Kalai, Nachum, Vempala and Zhang (2025-09-04, “Why Language Models Hallucinate”) argue that training and evaluation reward guessing over admitting uncertainty, so models learn to bluff; they propose scoring that credits calibrated abstention. Note this is one proposed explanation, not the only one.
Mitigations
- Grounding: RAG with citations, graphrag, search tools; check with rag-evaluation.
- Constraining output: structured-outputs, tool calls for facts and calculations.
- Verification: second-model checks (llm-as-judge-and-evals), tests for generated code, human review.
- Prompting: permit “I don’t know”, require quotes from the source (context-engineering).
- Model choice: see reasoning-models; whether reasoning models hallucinate less was not checked here.
Related: alignment-and-sycophancy, large-language-model, how-to-cure-llm-weaknesses-with-vector-databases, ai-slop.
Sources
- https://arxiv.org/abs/2509.04664 (Why Language Models Hallucinate, 2025)
- https://en.wikipedia.org/wiki/Hallucination_(artificial_intelligence) (secondary)
- Title, authors and date re-checked on the arXiv abstract page 2026-10-07.
Open items
- Author affiliations (the paper is commonly attributed to OpenAI researchers) not confirmed from a primary page; OpenAI blog returned 403.