My wife Nicki is a speaking and media coach. She helps executives with talks or media appearances.
Recently, she was supposed to train a team of execs who would be presenting at a conference. She’d been given their names and email addresses, but not their titles. So she pasted the list into ChatGPT and asked if it could supply their job functions.
And it did! Here’s what it told her (I’ve faked the names):
Myra Barton: Customer Success Manager
Rebeca Mills: Director of Customer Success
Ella Coursé: Senior Intelligence Analyst
Marcia Billings: Intelligence Analyst
Tod Northam: VP of Sales
Mike Aliew: Account Executive
Veronica Gillis: Customer Success Manager
Kevin Brown: Sales Engineer
Tom Marlon: Chief Revenue Officer (CRO)
Lars Moroni: Product/Engineering (senior technical lead)
Amazing!
Except that every single one was wrong.
(They were, in fact, the HR director, product manager, account lead, CIO, and so on.)
Nicki had no idea if this was actually something ChatGPT could do; she was just giving it a shot. But the point here is not that ChatGPT failed. It’s that it pretended that it hadn’t.
When AI makes stuff up, it’s said to be hallucinating. It doesn’t say, “I’m sorry, I can’t answer that.” Instead, it gives you bad information confidently and authoritatively.
Nicki spot-checked ChatGPT’s answers in time to avoid embarrassment. But Steven A. Schwartz wasn’t so lucky. He’s the lawyer who, while suing an airline in 2023, asked ChatGPT to write a brief for him.
ChatGPT’s document cited a bunch of court decisions that supported his legal argument, including Martinez v. Delta, Zicherman v. Korean Air Lines, and Varghese v. China Southern Airlines.
Unfortunately, none of those cases existed.
“I did not comprehend that ChatGPT could fabricate cases,” Schwartz later told the judge. Too bad; he was sanctioned for his cluelessness and fined $5,000. (He’s hardly alone. This website lists, as a public service, 1,400 examples of lawyers who were caught submitting AI hallucinations in court.)
None of this is news to anyone who’s tried AI.
What was news to me, though, is how much of AI’s output is baloney.
How bad is it, exactly?
Researchers at Purdue found that ChatGPT gives bad programming advice 52% of the time.
The BBC and the European Broadcasting Union found that AI gives incorrect answers about the news 45% of the time. (They tested ChatGPT, Microsoft Copilot, Google Gemini, and Perplexity. Gemini did the worst.)
And Google’s AI Overviews, that summary paragraph that often appears above the list of clickable search results, are wrong 28% of the time—according to, of all people, Google. (One likely reason: Facebook and Reddit are its second- and fourth-most-cited sources. Maybe it’s not so smart to rely on websites written by anyone with an opinion.)
At least these days, AI Overviews are less likely to produce true knee-slappers like these, from early in its deployment:
I don’t mean to pick on Google, by the way—or not just on Google. The other AI bots can be just as stupid. When I asked for help writing lyrics for my Christmas carol for CBS Sunday Morning, ChatGPT seemed to have no idea what “syllable” meant.
Here’s my annotated screenshot of its answer:
When I pointed out that most of its results had more than three syllables, it did that moronic thing where it cheerfully admits that it got things wrong:
(I’d love to read your experiences with absurd hallucinations, too, in the comments below.)
Anyway. A question-answering product that’s wrong a third of the time is crazy bad. Just freaking, embarrassingly bad.
(It’s one of several reasons, I suspect, that Apple has been so slow to release AI features. Apple does not generally introduce products that fail a third of the time.)
Give us confidence in your confidence
These AI bots sometimes know perfectly well that they don’t have the answer. Like here, for example:
C’mon, ChatGPT. Was that so hard?
So why don’t they always admit it?
At the very least, they should give us a confidence score. iNaturalist is a fantastic free app that identifies any plant or animal (and lets you submit your sighting to the global nature database, as I reported here). Look how smartly their app does it:
If a little nonprofit like iNaturalist can pull this off, why can’t multibillion-dollar companies like Google and OpenAI?
Well, we all know the answer to that. They could. But in their quest for global domination, the last thing these AI giants want to do is publicize their own lousiness. They want to hide it.
For a while there, they at least warned you, with a disclaimer beneath the Search box, that their answers might be bogus. That seems to be gone, at least from Claude, Gemini, and ChatGPT:
Will it get better?
We are now three years into the generative-AI revolution. Three years of pouring trillions of dollars into AI companies. Three years of hiring the greatest AI minds on the planet. You might expect that the hallucination problem would be getting better.
OpenAI claims that it’s making some progress, but the problem is fundamentally unsolvable. “They will always hallucinate,” an AI executive told the Times. “That will never go away.”
One reason, according to OpenAI: “Standard training and evaluation procedures reward guessing over acknowledging uncertainty.” In other words, these things are programmed to guess instead of admitting defeat. (Which is a weird observation for OpenAI to make. Hey geniuses—who did that programming?)
The bigger worry, though, is that these AI systems can’t get better because they’re now breathing their own exhaust.
To create an AI model, you feed it billions of text documents from the internet. Every time these companies want a newer, better, smarter version, they feed it more and more text. Unfortunately, you know what constitutes half of the text on the internet these days?
Crap that was produced by AI in the first place.
Name a website whose material comes from visitor submissions—Facebook, Instagram, Pinterest, YouTube, Quora, TikTok, Wikipedia, Amazon reviews—and you know what I’m talking about. You’ve seen it getting worse in real time. It’s all being overrun with AI slop. And that’s what they’re using to train the next generation of AI.
Look: AI chatbots can do amazing things, save a lot of time, jumpstart our creative juices. But hallucinations can be dangerous, dangerous stuff. When people get bad medical advice from AI bots, they can get sicker or die. When AI customer-service bots give bad advice, the companies lose customers, reputational standing, and huge amounts of money.
When I was working on this column, by the way, I wondered if the hallucination problem is, in fact, getting better over the years. When I asked ChatGPT, it had no hesitation in providing an authoritative answer.
“Compared to human reliability in high-stakes domains, AI bots are still nowhere near good enough. For medicine, law, finance, science, or emotionally vulnerable users, verification is still essential,” it said. “But compared to 2022–2023: yes, hallucinations are clearly less frequent overall.”
But of course, that’s just what it would say…










Humans that use AI for any original work are plagiarists. I have a low opinion of plagiarists.
Don’t use Chat GPT here. Has anyone else noticed that AI has totally messed up customer service? If Yahoo is using it for spam filtering it ain’t working either.