Hallucination

A hallucination is fluent model output that is false or unsupported by the provided source: invented facts, citations, quotes, APIs or file paths. In the author’s view it is tied to how next-token prediction works rather than being a bug with one fix.

Kalai, Nachum, Vempala and Zhang (2025-09-04, “Why Language Models Hallucinate”) argue that training and evaluation reward guessing over admitting uncertainty, so models learn to bluff; they propose scoring that credits calibrated abstention. Note this is one proposed explanation, not the only one.

Mitigations

Related: alignment-and-sycophancy, large-language-model, how-to-cure-llm-weaknesses-with-vector-databases, ai-slop.

Sources

Open items

  • Author affiliations (the paper is commonly attributed to OpenAI researchers) not confirmed from a primary page; OpenAI blog returned 403.