Why do models hallucinate?
TL;DR#
- Model hallucination is what researchers call it when AI models make stuff up
- This can lead to widespread misinformation, poor decision-making, and even psychological harm for users
- Hallucinations happen when models are fed bad training data…but they’re also part and parcel of how AI works
- Researchers are working on some promising methods to reduce hallucination
Remember the early days of ChatGPT, when screenshots like this one were making the rounds on Twitter?
This is a classic example of model hallucination, a term for when AI generates content that is factually inaccurate, misleading, or illogical. Hallucinations in language models can manifest as anything from arithmetic errors to false claims about history to declaring love for a human user (I’m looking at you, Sydney Bing). And because today’s models are so good at stringing words together, hallucinated content can often seem plausible at first glance.
For reasons I’ll explain in this post, hallucinations aren’t just “a side effect to be fixed”—they are pretty integral to how AI works. Some tech executives say that hallucination actually adds value to AI systems, by representing existing information in new and creative ways. But the phenomenon can also have serious repercussions. Like that one time a lawyer used ChatGPT and ended up citing fake cases in court. Yikes.
Does this mean you should avoid using AI tools altogether? Of course not—or this post would be a whole lot shorter. There are still tons of useful applications of generative AI tools, like to help you summarize a meeting transcript or brainstorm project titles. The key is figuring out how to use these tools responsibly, before they land you on the front page of Forbes (for the wrong reasons).