RSS Amplifier

The AutSide · Aug 9, 2026

When I Couldn't Quite Let the Pangram Thing Go

0
Sign in to vote or save

Jaime Hoerricks, PhD · The AutSide

I disabled Pangram—not because it was wrong, but because I couldn't independently validate its claims. Years in forensic science taught me that evidence must be reproducible. Black-box AI detection cannot yet meet that standard.

I’ve disabled Pangram on my Substacks, not because I had concluded it was wrong, but because I had concluded I could not evaluate whether it was right.

Those are different conclusions.

For several days I kept returning to the question—not obsessively, exactly, though I know how easily that word gets attached to autistic attention when what is really happening is that something has failed to resolve. The score sat there with the peculiar authority numbers acquire once they are placed inside an interface, and my autistic spideysense would not leave it alone. Something about the whole arrangement felt unstable. Not merely inaccurate. Unavailable to inspection.

Long before I became a teacher, I spent a career working in forensic science. Part of that work involved acting as an independent validation expert. I would be given the same evidence examined by another expert—sometimes for the prosecution, sometimes for the defence—and asked to conduct my own analysis. My report went to the Trier of Fact.

Sometimes the question was broader: what does the evidence support?

Sometimes it was narrower and, in many ways, more demanding. I would receive another expert’s notes alongside the underlying evidence and be asked whether I could reliably reproduce the result they had reported. Not whether I agreed with the conclusion. Not whether the report sounded technically sophisticated. Whether another competent examiner, working from the same evidentiary set and following the documented method, could arrive at the same place.

That question became part of how I understand evidence.

It is still there.

So when Pangram began assigning AI scores to my writing, my first instinct was not to accept the result or reject it. My instinct was to validate it.

Except I could not.

The model is proprietary. The training data are unavailable. The weighting of linguistic features is undisclosed. The decision process is hidden. The precise version of the system may change without the user knowing, and the result arrives without the kind of technical record that would allow another examiner to reconstruct how it was produced.

There is no shared evidentiary set in the forensic sense because the text itself is only one part of the process. The detector’s model is part of the process too, as are its training corpus, thresholds, calibration assumptions, preprocessing rules and internal representation of whatever it has decided “AI-like” writing is. Without access to those things, an independent examiner cannot repeat the original analysis.

She can only submit the same text to a different black box.

That is not reproduction.

It is another opinion.

This distinction matters because every legitimate measurement system contains uncertainty. DNA interpretation contains uncertainty. Fingerprint comparison contains a lot of uncertainty. Digital forensics contains complex uncertainty. The existence of uncertainty does not make a method useless, but the uncertainty must be available for examination. Another expert must be able to inspect the evidence, repeat the process, challenge the assumptions, identify interpretive choices, and show where the result does or does not hold.

The disagreement then becomes part of the evidence.

With AI detectors, disagreement simply accumulates. One detector reports a high probability of AI authorship. Another reports almost none. A third highlights several paragraphs whilst leaving the rest untouched. The systems do not explain why they disagree because none of them exposes enough of its method to permit comparison at the level where the disagreement was actually produced.

This does not necessarily prove that every detector is wrong.

It suggests something more fundamental: they may not be measuring the same thing.

One may be responding to lexical predictability. Another may privilege sentence variation. Another may have learned a cluster of stylistic features associated with a particular model, genre or training set. Each returns a score under the broad category of AI detection, but the common label may conceal entirely different operational definitions beneath it.

That possibility is what my forensic mind cannot quite release.

If institutions are going to accuse students, researchers, journalists or authors of misrepresenting their work, the evidential standard ought to become more rigorous, not less. An accusation of fabrication or academic dishonesty is not a gentle intervention. It carries consequences for credibility, employment, publication, education and reputation. The stronger the allegation, the more available its method should be to independent challenge.

Instead, the method is protected as proprietary whilst the output is treated as evidence.

The company is permitted opacity.

The author is required to explain herself.

There is another reason the Pangram question keeps catching.

Autistic people have spent generations being told that our communication is somehow unnatural—too formal, too repetitive, too precise, too explicit, too patterned, too dependent upon recurring structures, too insistent upon distinctions other people seem able to leave unstated. These descriptions are rarely offered neutrally. They form part of a longer history in which autistic language is treated as evidence of distance from the fully human rather than as language produced by a differently organised mind.

Now machines have been trained upon enormous archives of human writing, including writing produced by people whose language was already considered unusual, and new systems attempt to identify machine authorship by locating statistical resemblance to those machines.

The direction of imitation has disappeared.

The machine learns from human language, then the human is accused of resembling the machine.

I do not yet know whether neurodivergent writers are disproportionately flagged by AI detectors. There is evidence that some linguistic populations are, and there are plausible reasons autistic writers might be especially vulnerable: formal register, patterned syntax, explicit transitions, recursive clarification, repetition used structurally rather than accidentally. But plausible is not proven, and I have become increasingly wary of explanations that arrive before the evidence has earned them.

So I disabled Pangram.

Not because I had solved the mystery, and not because I had demonstrated that its result was false. I disabled it because I could not investigate the result in the way my former professional life taught me an evidentiary claim should be investigated. I could not inspect the method. I could not reproduce the analysis. I could not determine whether disagreement between systems reflected error, differing definitions or a deeper instability in the object being measured.

All I could see was the score.

And a score without an examinable path from evidence to conclusion is not validation.

It is an assertion wearing the shape of measurement.

Perhaps future systems will become transparent enough to permit genuine independent review. Perhaps common standards will emerge. Perhaps researchers will identify stable features that can be tested across models, genres and human populations without disproportionately misclassifying the people whose language already falls outside the expected centre.

Perhaps they will not.

For now, my autistic spideysense keeps circling the same unresolved point: the problem may not be that the detector sometimes produces the wrong answer. The problem may be that no one outside the black box can determine what question it has actually answered.

No posts

Read the original on autside.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.