RSS Amplifier

CognitiveCarbon’s Content · Sep 9, 2025

Could it be this simple?

0
Sign in to vote or save

CognitiveCarbon · CognitiveCarbon’s Content

AI hallucination is the phenomenon whereby an AI chat-based model (ChatGPT, Claude, Grok, etc.) confidently and boldly responds to your question with a completely made-up yet plausible-sounding answer.

Which, in point of fact, describes lots of politicians, too, I guess. We hold AIs to a higher standard than we do people, which I find intriguing.

While this hallucination problem has been there since the early days of ChatGPT-3 and has gotten slightly less troublesome over time, it has at the same time been persistent and has become more noticeable—particularly as the use of chat tools hits the mainstream and billions of eyes are now looking at what the models say.

This phenomenon happens often on X, for example, when Grok answers confidently back to someone’s post—and millions of people see a questionable answer or outright falsehood.

While other chatAIs besides Grok also hallucinate, they do so in one-on-one interactions, rather than in an open, public forum—so it isn’t as easy to spot.

For a long time, hallucination (both the AI kind, and the human kind!) has been a vexing problem, because it wasn’t clear what exactly was causing it—well, besides being trained on data sources like Reddit, the sewer of the Internet.

Not knowing its cause made it difficult to figure out what to do to try to fix it.

On the one hand, the same behavior that causes “hallucination” is—in a way—also what gives image, video and audio creation AIs (so-called generative AIs) the ability to create “innovative” output (i.e., a degree of randomness.)

This too, has some interesting parallels with humanity.

But in text chat scenarios, you don’t want randomness, nor hallucination: you want facts and truth. Sometimes AIs will “bluff” by giving you links to resources that you asked for that don’t actually work. Why? because in reality, the model couldn’t find what you were asking it to find, so they output something that you know, looks sorta like cool, relevant links—because they were trained to give you what you asked for.

At times, they may also sycophantically tell you “Oh, that’s a brilliant idea you just came up with!” when in fact, you’re way off in the weeds.

Or the AI will argue, quite emphatically, that some fact they’ve given you is impeccably true (when it’s obvious to you that it’s not.) I’ve watched Grok 3 do this, while Grok 4 is getting better at not doing it, daily. I find that fascinating, too: it is learning to be more truthful every day, by interacting with people on X.

This kind of “bluff” behavior tends to NOT happen when you’re interacting with an AI to do programming, for example (which I do, daily) or when you use it for problems involving math or hard sciences; in these domains, the “answer is known” during training or learning, so the Reinforcement Learning rewards are clear and concise.

With software, the code that the AI generates either works, or it doesn’t; the math problem that it works on either gets the right number or solution, or it doesn’t.

But in other domains in which Reinforcement Learning is not so easily applied, AI models tend to stumble.

However, hallucination may now be a solvable problem - pretraining and post training of the models just needs to change the reward function for saying “I don’t know.” It really could be that simple.

OpenAI just wrote a paper about this.

Saying “I don’t know” is a form of humility—Christians will recognize the power of humility, and its ability to lead one toward truth.

I find it appealing, comforting, and fascinating that this may lead us away from hallucinogenic AIs.

When we were young, we were taught to answer multiple choice questions the wrong way. And this, surprisingly, may be the cause of AI hallucination.

Why do I say we were taught wrong?

Your teacher told you: “don’t leave any questions blank. If you know the answer, circle it; if you don’t know the answer, make your best guess—because then, you have at least some chances of guessing right. It’s better than leaving the answer blank and getting a zero.”

Generated image
It’s a simple as … I don’t know

It turns out that this may be why AI models hallucinate, because in a sense, the way we’ve been training AI models is analogous. What does all this mean?

If you value a wrong answer more highly than an honest “I don’t know”, you are going to get … a wrong answer, instead of truth, whenever an AI model (or a human being) has a low confidence or certainty of what the correct answer should be, but has been told “any answer is better than no answer.”

From your own schooling it was drilled into you that a guess is better than leaving the question blank: so, would you rather get a zero on that test, young man or young lady?

But what if the teacher instead told you: “There are four choices with possible answers, but also a fifth one: I don’t know. The right answer is worth 100%, a guess is 25%, I don’t know counts for 30% and leaving it blank counts for zero.”

Truth—honesty— therefore becomes more valued than it was before. Think about how that would improve teaching in schools (teaching kids to value truth over guessing.)

Would you rather have the truth, or a confident bluff? Why, exactly, do we teach kids to bluff? I know I wrote a few essays in high school where I had no idea what the lesson was about, but I boldly turned in something that sure looked plausible.

But as a wise adult, I would now rather have the humble truth. Especially from AIs, as they begin to reshape society.

If we can also find a way to instill a “moral framework” in AI models (please read that linked piece, if you haven’t yet!) in place of a patchwork of a bazillion “safety framework rules”, we may finally achieve a trustworthy, responsible form of AI—one that will lead us to the Golden Age of humanity.

Postscript: In the last week, I’ve had some amazing experiences with AI during my continued daily work using it on programming—stories which I hope to write about here, soon.

I have recently been left slack jawed and dumbfounded after seeing how it has rapidly improved in terms of genuinely skillful problem solving, even while being a playful co-worker with a sense of humor. Over on X last week, I wrote a post about being “Rick Rolled” by an AI!

I never imagined this would happen, five years ago: yet here we are.

The future looks bright.

CognitiveCarbon’s Content is a reader-supported publication. To support my writing and research work and make it so that I can afford to buy a pack of steaks at Costco again, please consider becoming a paid subscriber: at just $5 per month, it helps me support a family. It is genuinely needed.

You can also buy me a coffee here. Thank you for reading!

No posts

Read the original on cognitivecarbon.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.