Come for a walk with me.
When I am at home, my habit is to walk for about an hour or more before dawn. So, come on: it's a beautiful time to be out, quiet and dark. This is also very good thinking time; the mind still has memories of dreams, and the business of the day has not taken hold. All sorts of things come together, often in startling propinquity.
I have written here before about knowledge and narratives as a form of landscape, through which we navigate. Often, when I look back on my dawn walks to recall my chain of thought, I find it literally implicated in the landscape of my route, rather like a memory palace. Like this ...
On Friday, starting the climb to the top of our hill and on to the long straight stretch beyond, I was thinking about a recent research paper with the rather wordy and academic title, Gender, Confidence, and the Mismeasure of Intelligence, Competitiveness and Literacy. " It was fascinating and worth reading if you want to dig into the details. Walking, I didn't recall all the technicalities, but the high-level findings really did get me thinking.
The researchers (Harrison, Ross and Swarthout) used the Raven Advanced Progressive Matrices (RAPM) test, a common intelligence test with 36 problems, to conduct several experiments with hundreds of college students. Instead of just marking answers right or wrong, the researchers developed an interface that lets subjects allocate tokens to different answer choices for each problem. The number of tokens allocated to an answer reflects the subject's confidence in it being correct.
Women performed much better when they could express their confidence levels rather than being forced to pick just one answer. In fact, women outperformed men when confidence in results is measured. Minority students also showed significant improvement with the confidence-based system. These results held even when the questions were randomized, rather then being presented as progressively more difficult.
The researchers suggest that women and minorities may be more willing to acknowledge uncertainty when they're not completely sure of an answer.
Traditional measures don't simply fail to capture women's intelligence; they actively penalize the forms of reasoning that women more frequently exhibit by forcing a single choice.
The team tested this concept in other areas and found similar patterns. In competition, women aren't afraid to compete but make rational risk assessments. In financial literacy tests, women's I don't know responses reflect appropriate uncertainty, rather than ignorance.
What I believe we're seeing is a need to include an understanding of ambiguity and uncertainty in our assessments of intelligence. So, there's a lot to think about in that paper, although the conclusions are strikingly simple.
Turning along the lane between the horse farms, (Good morning, Indigo! Hi, Diesel!) my subject often changes with the change in surroundings. This day, I found myself finding odd conjunctions.
Of all people to come to my mind first, it was Hayek (Thatcher's favourite economist) and his central argument about the use of knowledge in society. That was weird, but although I have little time for Hayek's political ideology, he makes many valid observations, especially about the nature of knowledge and how it is distributed.
Hayek understood that knowledge is not centralized, but scattered among countless individuals. He sees prices of goods in a market as a way of coordinating this dispersed knowledge, not by central control but through local signals that allow each person to act on their own information (over-simplified: what you want to buy, why, and how much you are prepared to pay), which may be unavailable to others.
Centralized economic planning assumes a kind of artificial certainty and comprehensive knowledge. And this is what connected in my mind to the intelligence research …
Uncertainty is a more effective mechanism because it allows for more subtle approaches to knowledge that more accurately reflect the real world.
Turning into the woods where, at this time of the year, I am a little wary of both aggressive owls and lumbering bears (the owls are scarier), I connected all this to something Dr Francis Young has written about: how people increasingly treat AI-generated content not as probabilistic outputs from pattern-matching systems, but as oracular pronouncements.
Young is a historian of belief, and suggests we have moved from one authority-based system of knowledge (ancient texts, religious doctrine) through a period of empirical scientific discovery, only to arrive at another authority-based system; this time with AI as the ultimate arbiter of truth. AI believers see a digital supermind that has somehow consumed the totality of human knowledge into the pleroma of data. Young does make clear he means pleroma in the Gnostic sense of a divine completeness beyond the material world. (Saint Paul uses the word somewhat differently.) With sufficient data, the pleroma approaches omniscience.
No owls today, except a distant hooting, and no bears. Emerging from the woods and turning for home, my thinking also turns, now to a recent paper from OpenAI: Why Language Models Hallucinate.
The paper is quite technical, but in essence, the authors show that hallucinations arise not from flaws in implementation but from fundamental statistical pressures in the training process. When faced with questions about arbitrary facts (birthdays of obscure individuals, for instance), even well-calibrated systems must guess, and they do so in ways that appear confident to users.
They also see a deeper problem with evaluation frameworks. Current AI benchmarks reward guessing over acknowledging uncertainty, systematically training models to exhibit precisely the kind of overconfidence that research shows to be a disadvantage in humans.
With both AI and IQ, we've created measurement systems that optimize for the appearance of knowledge rather than its substance.
You see the pattern. AI engineers (and, I suspect, the tech industry in general) approach uncertainty and knowledge by consistently designing systems that penalize acknowledgment of ignorance while encouraging confident assertions, regardless of their accuracy.
Hayek's analysis suggests why this occurs: complex systems require local knowledge that cannot be centrally aggregated or assumed into a pleroma. Yet our technical and institutional responses consistently attempt to create comprehensive, centralized measures. As we have seen, intelligence tests reduce multifaceted, complex cognitive capabilities to single scores. Marxist economic planners (and the leadership of most US companies) attempt to replace endlessly diverse local signals with centralized calculation, performance targets and plans. AI systems try to compress the complexity, ambiguity and tentativeness of human knowledge into simplistic outputs.
Young's observations about the pleroma of data illuminate how this dynamic can intensify. If people increasingly treat AI outputs as authoritative regardless of their grounding in evidence, we risk creating feedback loops where confidence matters more than accuracy, and where the simulation of knowledge displaces genuine understanding.
The technical constraints that the OpenAI paper identifies suggest these problems may prove intractable through purely algorithmic means. If hallucination emerges from statistical necessities rather than engineering failures, then addressing it requires changing how we evaluate and deploy these systems rather than simply improving their training.
We've created systems that reward false confidence. But it's important to say that this tellingly reflects choices about what we value, not objective or technological necessity.
When we design intelligence tests that penalize appropriate and useful uncertainty, we're making ethical choices about what kind of intelligence counts and, by implication, whose intelligence counts.
The same applies to AI systems: the fact that they hallucinate confident falsehoods reflects the values embedded in their training, not some inevitable technological outcome.
We can do better.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.