Why a model can be confidently wrong

One failure mode from a curriculum I'm writing on how and why AI models fail.


Picture someone who has read an enormous amount and has never once been taught to say "I don't know." Every time they were asked something, they were rewarded for an answer that sounded right. So when you ask them something obscure, they don't pause. They build the most plausible answer they can, calmly and fluently. Sometimes it's correct, sometimes it isn't, and you can't tell which from how they sound.

That's a language model. It's trained to predict what plausibly comes next, based on the text it learned from. Nothing in that training asks whether a sentence is true. Hallucination isn't the model breaking. It's the model doing what it was built to do.

Two ways it goes wrong

Ask for a company's exact revenue in one quarter, a rare fact. Sometimes the model is spread across many possible numbers, and a careful system can see that it's unsure. The dangerous case is when it has latched onto one number, from a similar company or a different quarter, and states it with full confidence. That answer reads exactly like a correct one.

Unsure: probability spread across many answers
Confidently wrong: one tall bar on a single wrong answer

What helps

Make every claim traceable. Give the model verified sources to work from and require each claim to point to one. Then check each claim. No source: flag it. Source doesn't actually say it: flag it. Otherwise it's grounded. This doesn't stop the model from making things up. It makes the made-up parts visible.

Why it matters

A bigger model or longer training doesn't guarantee a fix, because the problem isn't how well the model copies its data. It's that nothing in training tells it what's true. The checking has to be built around the model.