Published on September 19, 2026
LLM Hallucinations: Why AI Makes Things Up and How Not to Get Fooled
- artificial-intelligence
- llm
In the first two articles of this series I explained what sets an LLM apart from artificial intelligence, machine learning, and deep learning and how it actually works, from tokenization to token-by-token generation. I'm closing the series with the topic I think matters most for anyone using these tools every day: why an LLM sometimes states false things with total confidence, and what we can actually do about it.
The phenomenon has a name that's entered everyday language by now: hallucination. It isn't an occasional glitch or a "bug" waiting for a software update to fix it — it's a direct, almost unavoidable consequence of how these models work at their core. Understanding why it happens isn't just a matter of technical curiosity: it's what tells you when to trust the output more and when to double-check it.
What we actually mean by "hallucination"
A hallucination, in this context, is a statement generated by the model that sounds plausible, written in exactly the same confident tone as any other response, but is false or unsupported — an invented fact, a citation that doesn't exist, a correct number attributed to the wrong source, a law or a scientific paper that simply doesn't exist. What makes hallucinations particularly insidious isn't the falsehood itself (any source, human included, can be wrong) — it's the total absence of any signal of uncertainty. The model doesn't "hesitate," doesn't shift tone, doesn't flash a warning. It writes the invented thing with exactly the same fluency and confidence it would use to write something true.
Try it yourself
In what year did OpenAI publish the paper 'Language Models are Few-Shot Learners', the one that introduced GPT-3?
Both answers below are written with the same confidence. Pick the one that seems right to you.
The root cause: a statistical model, not a database of facts
To understand why this happens, we need to go back to what we covered in the previous article: an LLM doesn't consult an archive of verified facts when it generates a response. It generates text by estimating, token after token, the statistically most plausible continuation given the text sequence that precedes it, based on regularities observed during training on enormous amounts of text. There's no mechanism inside the model that explicitly distinguishes "this sentence is a verified fact" from "this sentence merely sounds plausible." From the point of view of the generation mechanism, the two are treated the same way: sequences of tokens with a certain probability.
A research paper published by OpenAI in September 2025, titled precisely "Why Language Models Hallucinate," formalizes this intuition and goes beyond a simple description: the authors — Adam Tauman Kalai, Ofir Nachum, and Edwin Zhang of OpenAI, together with Santosh Vempala of Georgia Tech — argue that hallucinations "need not be mysterious": they originate, fundamentally, as ordinary errors of the kind that show up in any binary (true/false) classification task, and the paper analyzes the statistical causes of the phenomenon within the modern training pipeline (Kalai, Nachum, Vempala, Zhang, Why Language Models Hallucinate, OpenAI, 2025). A central point of the paper is that generating answers is structurally harder than classifying them as true or false: a model might correctly recognize whether a statement is true or false when it's shown to it already written out, but when it has to generate that statement from scratch, the chance of error increases, because the generation task inherits, and amplifies, the errors of the underlying classification task.
But the paper goes further than the technical explanation and points to a perverse incentive in how models are trained and evaluated: standard training and evaluation procedures reward guessing over admitting uncertainty. The authors use an effective comparison: a language model facing a hard question behaves a bit like a student facing an exam question they don't know the answer to — attempting a plausible answer, even a made-up one, is statistically more likely to be rewarded (in standard benchmark scores) than honestly answering "I don't know." As long as evaluation systems keep penalizing declared uncertainty more harshly than a confident wrong answer, models will keep being mathematically pushed toward "always attempt an answer."
A second mechanism: what happens when the model recognizes a name but not the facts
Interpretability research from Anthropic, which examined the internal mechanisms of its own Claude model directly, offers a second, particularly concrete piece of the puzzle. The somewhat counterintuitive finding is that the model's "default" behavior isn't to answer — it's to refuse to answer: researchers identified "a circuit that is 'on' by default and that causes the model to state that it has insufficient information to answer any given question" (Anthropic, Tracing the thoughts of a large language model). In other words, cautious refusal is the model's starting position, not the exception.
What happens when the model does answer confidently is that this default refusal circuit gets suppressed by another internal feature, tied to recognizing "known entities": when the model recognizes a name, topic, or concept it's familiar with, this feature inhibits the default refusal and lets the model respond. The problem — and this is the interesting part — is that this mechanism can misfire: it can activate even when the model recognizes only the name of something or someone, without having reliable information about anything else. In those cases, the refusal still gets suppressed, and the model "proceed[s] to confabulate: to generate a plausible — but unfortunately untrue — response" (Anthropic, Tracing the thoughts of a large language model). It's a technical explanation for a pattern anyone who's used an LLM has probably noticed: models tend to hallucinate most confidently about exactly the people, companies, or references they recognize by name but actually know very little about.
Real, documented, verifiable cases
Hallucinations aren't a theoretical, lab-only risk: they've already had concrete consequences, in court and on financial markets. Here are three real cases, all documented by verifiable sources.
Mata v. Avianca: invented legal citations
In 2023, in a personal injury lawsuit filed against the airline Avianca in the U.S. District Court for the Southern District of New York, the plaintiff's lawyers used ChatGPT to prepare a legal brief supporting their position. The brief cited six prior court cases as precedent. The problem: none of those six cases actually existed; they had been entirely invented by the model, complete with plausible-sounding but fictitious internal citations and judge names. When opposing counsel couldn't locate the cited rulings and flagged this to the court, rather than immediately admitting the mistake, one of the attorneys went so far as to submit a sworn affidavit containing supposed excerpts from the rulings — themselves also generated by ChatGPT and equally fabricated. The judge held a hearing on June 8, 2023, and issued a sanctions order on June 22, 2023, imposing a $5,000 fine, jointly and severally, on the two attorneys involved (U.S. District Court, S.D.N.Y., Mata v. Avianca, Inc., Opinion and Order on Sanctions, June 22, 2023).
Air Canada: the chatbot that invented a policy
In November 2022, a passenger named Jake Moffatt, after losing his grandmother, used the chatbot on Air Canada's website to ask about the airline's bereavement fare policy. The chatbot told him he could apply for the discounted bereavement fare retroactively, within 90 days of the ticket's issuance. That wasn't actually the airline's policy, which didn't allow retroactive requests at all. When Moffatt later asked for a refund of the difference, Air Canada refused, arguing before the tribunal that the chatbot was "a separate legal entity" responsible for its own statements. British Columbia's Civil Resolution Tribunal rejected that argument on February 14, 2024, ruling that a company remains responsible for all the information on its website, whether it comes from a static page or a chatbot, and ordered Air Canada to compensate the passenger (Civil Resolution Tribunal of British Columbia, Moffatt v. Air Canada, 2024 BCCRT 149).
Google Bard: an error in a promotional video
You don't even need a user asking tricky questions — sometimes a hallucination shows up in a product's official launch demo. In February 2023, in a promotional video Google published to introduce Bard (its then-new chatbot competing with ChatGPT), the model was asked what new discoveries from the James Webb Space Telescope it could tell a nine-year-old about. Bard answered, among other things, that the telescope had taken "the very first pictures of a planet outside of our own solar system." That statement was false: the first image of an exoplanet had actually been taken almost two decades earlier, in 2004, by the European Southern Observatory's Very Large Telescope. The error was first reported by Reuters, and within a day Alphabet's (Google's parent company) shares fell 7.7%, wiping out roughly $100 billion in market value (CNN Business, Google shares lose $100 billion after company's AI chatbot makes an error during demo, February 8, 2023).
Three very different cases — a legal proceeding, a contract dispute, a promotional video — but with the exact same thread running through them: a statement made with complete confidence, with no signal of doubt, that turned out to be false only once someone actually checked.
How widespread is the problem today
Here it's important to be honest about the limits of what can be stated with certainty. The 2025 AI Index Report from Stanford's Institute for Human-Centered Artificial Intelligence (HAI) — arguably the most authoritative independent report on the state of AI — devotes a section to measuring hallucinations, noting that older benchmarks like HaluEval and TruthfulQA failed to gain widespread adoption within the research community, and that newer evaluation tools have emerged in their place, such as the updated Hughes Hallucination Evaluation Model leaderboard, FACTS, and SimpleQA (Stanford HAI, 2025 AI Index Report — Responsible AI).
I'm deliberately not quoting specific hallucination-rate percentages here, because the figures circulating online for 2026 aren't, as far as I can verify at the time of writing, backed by sources as solid as the ones cited above, and they vary enormously depending on the task being measured (summarizing a given text is a completely different challenge from answering an open-ended factual question). Following the same principle that guided this whole article — better to admit a limit than to invent a number — I'll simply say that the field now clearly treats this as a priority, that the benchmarks used to measure it have multiplied and improved in recent years, and that measured error rates vary widely depending on the type of task asked of the model.
Hallucinations aren't the only limit
It's worth spending a few lines on other LLM limitations too — less discussed, but just as real, because "LLM limits" isn't a synonym for "hallucinations." It's a broader category.
The knowledge cutoff. An LLM is trained on a body of text collected up to a certain date, after which training stops. Everything that happened in the world after that date is simply absent from its internal "weights" — not because the model ignores it by mistake, but because that information was never part of the material it learned from. Some systems compensate for this by connecting the model to real-time external search tools, but the base model, on its own, stays anchored to the moment it was trained.
Sensitivity to how a question is phrased. The exact same question, asked with slightly different wording or in a different order, can produce different answers — sometimes substantially different ones. This is a direct consequence of the fact that the model generates tokens based on statistical patterns learned from text, and small variations in a prompt can shift which patterns get "activated" during generation.
No genuine internal fact-checking. Unlike a search engine or a database, an LLM has no native way to check its own statements against an external source before writing them down: verification, when it happens, happens through additional tools (like the ability to search the web mid-conversation), not as an intrinsic property of the base language model.
Biases inherited from training data. A model learns statistical patterns from the text it's trained on; if that text reflects biases, imbalances, or the over-representation of certain viewpoints over others, the model tends to absorb and reproduce those too, not just grammar and facts. This is an active area of research — the same "Responsible AI" chapter of Stanford HAI's 2025 AI Index Report covers it alongside hallucination measurement, as part of the broader picture of model safety and reliability (Stanford HAI, 2025 AI Index Report — Responsible AI) — one more reason not to treat an LLM's output as a neutral source by default.
Concrete practices for not getting fooled
All of this points to a few practical habits that, in my experience, genuinely make a difference in everyday LLM use.
- Always verify specific citations and references. If a model cites a court ruling, a scientific paper, a law, or a statistic with a source attached, the bare minimum is checking that the source actually exists and says what the model claims it says — this is exactly the check that, had it been done before filing the brief, would have prevented the entire Mata v. Avianca case.
- Be more suspicious, not less, when an answer is very specific about something obscure. As Anthropic's research on the "known entities" mechanism suggests, the risk of confabulation is highest precisely when you ask for details about people, companies, or events the model "recognizes by name" but has seen very little text about during training.
- Cross-check with multiple independent sources, especially for numbers, dates, proper names, and claims that would have real consequences if wrong (a company policy, a legal deadline, a dosage).
- Don't treat a confident tone as a sign of reliability. This is perhaps the hardest habit to build, because it runs against a natural reflex: normally, when a person speaks confidently about something, we tend to trust them more. That heuristic doesn't work with an LLM, because — as shown above — the model generates correct answers and invented ones with exactly the same fluency.
- Use LLMs as a starting point for research, not a substitute for final verification, especially in domains where a mistake has a real cost: law, medicine, finance, safety.
- Give the model the actual source documents or context when you can, instead of relying only on its internal memory: a model that answers based on a text supplied directly in the prompt tends to be more reliable than one that has to "recall" a fact from its training.
A limit worth knowing, not a reason to give up
Understanding why LLMs hallucinate doesn't mean concluding they're unreliable across the board, or that they should be avoided. It means knowing what to expect from a tool that, given how it's built — a statistical model predicting the most probable word, not a database of verified facts — works great for summarizing, rephrasing, brainstorming, or explaining general concepts, and should always be double-checked whenever the answer depends on a specific, verifiable fact.
If you've read the whole series, you now have the tools to do exactly that, with real understanding: you know where an LLM sits relative to other forms of artificial intelligence, you know how it actually works under the hood, and now you also know why it sometimes gets things wrong with complete confidence. That awareness — not blanket distrust, and not blind faith — is the most useful way to use these tools today.