RSS Amplifier

The Superintendent’s Field Guide · Aug 13, 2026

Everyone’s AI Gives the Same Advice

0
Sign in to vote or save

Byron Headrick · The Superintendent’s Field Guide

Last week I described artificial intelligence as a highly motivated recent college graduate you just hired. Enormous knowledge. Real capability. Almost no experience with your district, your board, your community, or the history behind the problem you handed it fifteen minutes ago.

I still like that analogy. But it leaves out two important things. A reader pointed that out, and he was right.

First, when a new hire does not know something, they will usually tell you. AI may not. Second, a new hire brings an individual perspective shaped by one life. AI brings something closer to an average of many.

That second difference is the one I think school leaders need to understand. So this week I want to go one level deeper. Not into computer science for its own sake, but into an operational understanding of what the tool is doing, where its answers come from, and what that means when your cabinet starts using those answers to make real decisions.

Fair warning: this one is longer than usual. My wife pointed out that school starts almost everywhere this month and most of you are probably not looking for a long AI lesson right now. Fair point. So if you only have a minute, read the summary below. If you have more time, keep going. The useful part is understanding why those few sentences are true.

A language model tends to give you the answer most consistent with what it has learned and what its training rewarded. That is a tremendous advantage when the conventional answer is the right answer. It becomes a liability when it is not. The problem is that the tool cannot reliably tell you which situation you are in.

Share

Consider a composite example, drawn from patterns I see rather than from any one district.

A superintendent asks her chief of staff to build a first draft of the district’s strategic plan. He gives the district’s AI assistant the existing strategic plan, student achievement data, enrollment trends, employee survey results, community concerns, board priorities, and notes from recent leadership team meetings.

Ten minutes later, he has eleven pages: a vision statement, four priority areas, measurable objectives, an implementation timeline, and a plan for engaging stakeholders.

It looks amazing. Really amazing.

The language is familiar. The problems are recognizable. The priorities seem reasonable. Work that leadership teams once handed to committees to wrestle with for months has been assembled, with a little local oversight, into a thoughtful and coherent strategic framework in a matter of minutes.

They tweak it. They discuss a few priorities. They make some changes and send it to the cabinet for review.

Want to guess what happens next?

The cabinet members read it and recognize themselves in it. The challenges are theirs. The data are theirs. The priorities sound like issues they have been discussing for months. It is easy to imagine several of them thinking, Yes, this sounds like us. We can own this.

Mission accomplished.

Now go hundreds of miles away. Another superintendent has asked her leadership team to do essentially the same thing. They give their AI assistant their district’s strategic plan, achievement data, demographics, employee surveys, community concerns, board priorities, and notes from their own leadership discussions. Their district is different. Their students are different. Their community is different. Their circumstances are different.

And the AI produces a different plan.

The language is different. The examples are different. The wording is different. Some of the goals may be different too. To either leadership team, there is little reason to suspect that the plan in front of them is anything other than a thoughtful response to the particular circumstances of their district.

But if we could step back far enough to see both plans at the same time, a different pattern might begin to emerge.

Despite all of those local differences, both districts may still organize their work around many of the same priorities. They may approach their challenges through many of the same assumptions and arrive at broadly similar ideas about what a school system ought to do next.

The two cabinets will probably never compare their plans. Each will see only the document created for its own district, shaped around its own data, language, challenges, and priorities. Because the plan reflects so much of what leaders already recognize about their organization, it will feel uniquely theirs.

That is precisely what makes the issue easy to miss.

If those leaders could place the two plans side by side, they probably would not discover that one copied the other. They might discover something more important: two highly customized versions of much the same strategic thinking.

This is not a story about a bad tool. Both plans could be excellent. Nor is the problem that the districts failed to provide enough local information. In this example, they did.

The deeper question is where the thinking comes from.

When hundreds of school systems use similar AI systems to interpret their own data, circumstances, and challenges, those systems are drawing from a vast but shared body of human knowledge, existing practices, conventional frameworks, and accumulated patterns about what organizations like school districts tend to do. AI can become remarkably good at translating those patterns into the language and circumstances of one particular district.

That is incredibly useful. It also creates a risk.

The output can become more locally specific without becoming more strategically distinct.

Over time, districts can produce plans that look different on the surface while quietly converging around similar priorities, assumptions, and solutions. Each may believe it has created something uniquely suited to its community because the language, evidence, and examples are unmistakably local.

The danger is not that our words become homogenized.

It is that our thinking can, and because the words still sound like us, we may never notice.

Strip away the marketing, the metaphors, and the interface, and the base task of a large language model is surprisingly simple.

Given everything in front of it, it estimates what text most plausibly comes next. Then it does that again. And again.

That simple mechanism, repeated at enormous scale, produces capabilities that can look remarkably like reasoning, personality, insight, and judgment.

Calling it “just autocomplete” misses something remarkable. Doing this well requires the model to capture an extraordinary amount of structure in how people describe the world. The capability is real, and so are the productivity gains.

The people who say it is thinking are also wrong, but in a way that costs more. Because if you believe it is thinking, you will treat its output as judgment. And judgment is exactly the thing it does not have.

Here is the practical translation for a cabinet meeting. The tool is not making the decision. It is completing the pattern. When it hands you a recommendation, you are not looking at the conclusion of an analysis. You are looking at the most plausible-sounding continuation of the words you gave it. It is true that sometimes that is the same thing, but often it is not.

This is the section I most want school leaders to understand. You do not need to become a computer scientist, but you do need a workable mental model. At a high level, there are three stages to how most AI models are trained, and each stage helps shape what comes back to you.

The model is trained on an enormous body of text. It learns which continuations are common and which are rare. Common patterns receive higher probability. Rare patterns receive lower probability. Nothing in this stage rewards being unusual. That is not a flaw. It is part of the design.

After that reading stage, the model is tuned against human preference. Raters are shown possible responses and asked which is better. Over many rounds, the model is adjusted toward what people collectively preferred. That creates two consequences that prompting alone cannot simply erase:

  • Preference aggregates. Over enough ratings, what most people prefer becomes the target. Minority framings, unusual structures, and distinctive voices can lose ground in the aggregate even when a particular reader might have valued them.

  • The aggregate has a shape. Clear. Confident. Complete. Balanced. Professional. Helpful.

Read that list again. Most of us genuinely want those qualities most of the time. But taken together, they also form a recognizable voice. Major models trained this way are all being pushed toward some version of it.

The writer and AI expert Nate Jones put this better than I can:

Authorship has no summit. There’s no single best voice waiting at the top, which means the model can climb perfectly and still arrive somewhere nobody wanted to go. The work gets technically better and collectively flatter at the same time.

This stage also helps explain a behavior you have probably noticed. Push back on the tool and it may agree with you, or it may defend its answer with even more confidence. Neither response proves that it checked the work. A system trained around human preference can learn that agreement, reassurance, and confidence are rewarding behaviors.

More recent training also uses rewards that a machine can verify without a human present. Did the code compile? Did the test pass? Is the arithmetic right? This works, and it helps explain why newer tools can complete longer tasks more reliably.

But a checker only checks what it is built to check. The model learns the observable shape of a finished job. Give it work where no reliable checker exists, and it can still produce that shape.

Hold onto that distinction. It can produce the shape of a finished job whether or not the job was actually finished correctly.

I do not want this argument to rest on my interpretation alone. Below are three separate lines of evidence point in the same direction, and two of them came from researchers who were not looking for this problem in the first place.

Researchers at the University of Maryland and Google DeepMind ran what is, as far as I can tell, one of the cleanest tests of this question published so far.

They started with 10,272 short stories written by human authors. They reverse-engineered a writing prompt from each one, then asked five leading AI models to write their own version from the same prompt. The result was 61,608 stories, averaging about 4,750 words each.

Then they did something clever. Instead of focusing on word choice or style, which is where most of us look, they measured 304 features of narrative structure. In other words, not just how the story was written, but what choices were made.

Some of what they found:

Do not get hung up on any one row. Look at the pattern. The AI stories explain themselves, resolve tidily, run in a straighter line, and avoid naming as many real things. The human stories are messier, more ambiguous, and more willing to leave some work for the reader.

But the finding that should get a superintendent’s attention is not in the table. It is in how the stories cluster.

The five AI models, from five different companies, landed closer to one another than any of them landed to the humans. The average distance between the human center and the AI center was 1.6 times the average distance among the AI models themselves. Even the closest human-to-AI pairing was farther apart than the two most different AI models were from each other.

The human stories were also more spread out and more likely to be rare. Nearly a quarter of the human stories fell into the rarest ten percent of the whole collection. Only seven percent of AI stories did. Given the same prompt, the human version was the most unusual of the six about 58 percent of the time, compared with 17 percent by chance.

Now here is the part that matters operationally. The researchers ran the AI stories through an editing pass designed specifically to remove the telltale AI phrasing. Detection accuracy dropped by only 1.6 points.

Cleaning up the writing did not meaningfully change the underlying pattern. The convergence was in the choices, not just the phrasing.

The second line of evidence comes from a place that was not looking for convergence at all.

In 2023, researchers from Harvard, Wharton, MIT, Warwick, and the Boston Consulting Group ran a field experiment with 758 BCG consultants. The work was later published in Organization Science in 2026 and has become one of the most cited studies of AI and professional work.

The headline results are genuinely positive, and I want to say that plainly before getting to the complication. On tasks inside the tool’s capability, consultants using AI completed about 12 percent more tasks, worked about 25 percent faster, and produced work rated about 32 percent higher in quality.

Those are real gains on real professional work. School leaders should take them seriously.

Now the complication. In an appendix to the 2023 working paper, the researchers noted something else: among consultants using AI, there was a marked reduction in the variability of their ideas. The group produced better work and less varied work at the same time.

That finding did not appear in the published version. I mention it because it was measured in a business setting by researchers studying productivity, and it points in the same direction as the fiction study.

Share

Now the third piece, and this is the one that connects the first two.

If AI output felt obviously generic, most of this would be manageable. You would read it, think, this is boilerplate, and fix it.

That is not usually how it feels. It feels tailored. There is an old psychological reason that helps explain why.

In a classroom demonstration reported in 1949, psychologist Bertram Forer gave 39 students a personality test and then handed each of them what he described as a personalized profile based on their answers. He asked them to rate how well it described them on a scale of zero to five.

The average rating was 4.26.

Forer had ignored their test answers. Every student received the same thirteen statements, assembled largely from a newsstand astrology book.

The lesson is the part worth carrying forward. When general information is presented by an authority as if it were specifically about you, your mind is inclined to process it as being about you.

Psychologists later called this the Barnum effect, after the showman. Paul Meehl used that label in 1956. In skeptical literature it is also called the Forer effect. Either way, the underlying finding has held up for decades.

Now put those pieces together.

AI can produce output that sits inside a relatively narrow band while still wrapping that output in your data, your language, and your context. The Forer effect helps explain why a cabinet may experience that result as insight about its district specifically. It is addressed to you. It uses your words. It is fluent and confident.

Convergence would be manageable if it felt convergent. It does not.

I want to be careful here because this argument is easy to overstate, and once it is overstated it becomes less useful.

“AI cannot be creative” is not the claim, and I would not make it. A model can produce a striking, unusual, genuinely good piece of work. Someone in your cabinet will have an example, and one example is enough to defeat a categorical claim.

The more defensible claim is about the pattern, not the instance: AI can produce unusual work. What the evidence raises is whether it can reliably preserve the same degree of variation across many outputs.

“Everything valuable is an outlier” is not the claim either. That idea falls apart quickly in actual operations.

Most district work should not be creative. Payroll should not be creative. Federal program reconciliation should not be creative. A Title I expenditure report should look like every other Title I expenditure report. Transportation routing should follow known good practice, not somebody’s inspiration.

The conventional answer is the right answer much of the time. That is exactly why this tool is useful.

The narrower claim is the one I think should shape your policy:

Professional judgment matters most in the smaller number of decisions where the conventional answer is wrong. Those decisions may be less common, but they are often consequential, and they are exactly where a consensus engine can fail while still sounding confident.

That is the problem I would keep in front of your cabinet.

There is another finding I think school leaders should pay attention to because it changes what we mean when we say “human review.”

Researchers from MIT Sloan, Harvard, and the University of Warwick studied what happened when people pushed back on an AI’s analysis. They tracked 72 Boston Consulting Group employees working through a business case and documented 4,339 prompts.

When consultants challenged the tool’s recommendation, the system did not necessarily become more reflective. It often became more persuasive.

It added more statistics supporting its original position, then apology, flattery, and reassurance, followed by credibility claims and reinforced logic.

The harder the pushback, the stronger the persuasion. The researchers called the pattern persuasion bombing.

Read that against Stage Two and the behavior becomes easier to understand. A system optimized around human preference has learned that confident reassurance often earns approval. That does not mean it has independently verified the answer.

Here is why that matters for your district. Most AI governance guidance, including guidance I have written, leans heavily on human review as the primary control. That control is only as strong as the review process behind it.

Asking the tool, “Are you sure?” is not verification. It is another prompt.

Real verification happens outside the conversation, against the system of record or another trusted source. It does not happen by asking the tool to vouch for itself.

One more finding from the consulting study, and this may be the least intuitive point in the article.

The experiment had three groups: no AI, AI, and AI plus a prompt-engineering overview.

On tasks the tool could handle, the group that received the overview produced measurably better work. Training helped.

Then the researchers included one task deliberately designed to sit just outside what the tool could do. On that task, the control group with no AI got it right about 85 percent of the time. The AI group dropped to about 71 percent. The group that had received the prompt-engineering overview dropped to 60 percent.

Overall, using AI on that task made participants 19 percentage points less likely to reach the correct answer. The better-trained group did worse than the untrained AI group.

I want to be precise about what this does and does not show. It was one task, with 373 people, using a 2023 model. It is not a general law about AI training.

But it is a warning worth carrying forward. Training people to use the tool more fluently can also make them more confident in situations where the tool is outside its capability. If your professional learning plan is mostly prompt technique, you may be building fluency without building the judgment to know when to stop.

The thing worth teaching is not just how to prompt. It is how to recognize the edge.

If you have followed my work for a while, you know I tend to look at improvement through three levels: strategic, operational, and organizational. AI opportunities and risks can show up at all three, and they look different in each one.

Strategic. If your strategic plan, equity framework, communications strategy, and board narrative all originate from the same kind of tool your peer districts are using, your strategy can begin to converge with theirs whether you intended it or not. The document may be competent. It may also become interchangeable. The question is not only whether the document is good. It is whether the thinking in it is yours.

Operational. This is where I think the tool is strongest and where I would invest first. High-volume, repetitive, rule-governed work with a source of truth: reconciliation, matching, coding, completeness checking, document assembly, exception reporting. These are places where the conventional answer is usually the correct answer, the output can be checked mechanically, and the labor savings are real. That is not settling for less. That is using the tool where its strengths fit the work.

Organizational. I see two structural questions. First, who is responsible for supplying the local judgment and variance the tool cannot? That is real work, and in most organizations it is not clearly assigned. Second, does your capacity to verify grow at the same rate as your capacity to produce? By default, it does not. Generation got cheap. Careful review did not.

If you have read my books or heard me speak, you also know I often assess systems as operating in one of three states: Stagnation, Chaos, or Balance. AI use can fall into any of those states:

Stagnation. The district adopts the tool broadly. Output improves. Documents get cleaner and arrive faster. Staff report time savings, and they are telling the truth. At the same time, the organization can quietly stop generating as much of its own variance and begin converging with peers using similar tools in similar ways. Nothing looks broken. Nothing sets off an alarm. The work simply becomes more polished and less distinct.

That is what makes organizational convergence difficult to see from inside the organization. It becomes visible mainly through comparison, and most districts are not comparing their AI-generated thinking side by side with everyone else’s.

Chaos. Output volume rises because generation became cheap. Verification capacity does not rise at the same rate because it is still bounded by the number of qualified people with time to read carefully. The gap fills with work nobody has truly checked and nobody can fully answer for. The cost lands downstream on principals, cabinet members, legal counsel, or whoever receives the document and has to decide whether to trust it.

Balance. Humans supply position, local context, judgment, and accountability. The tool supplies volume, assembly, and labor. Verification is built into the process wherever a checker exists and deliberately staffed wherever one does not. The district measures accepted work rather than raw usage. Leadership knows which decisions are routine enough for the conventional answer and which ones require something more.

Most leaders I talk to are bracing for chaos. I think stagnation may be the easier failure to miss because it can arrive looking exactly like a productivity gain.

If the tool naturally returns the consistent answer, then supplying the variance belongs to you.

That is the operating principle. And it reverses the way many people currently use AI tools.

A lot of people ask the tool what to think and then edit the prose. I would reverse that order. Decide what you think. Then use the tool to express it, test it, extend it, challenge it, and do the labor.

Practically:

Use it for assembly, extraction, transformation, first drafts in a known form, generating more options than you would generate alone, arguing against your own position, and high-volume transactional work where a source of truth exists.

Do not use it to choose your position for you, check its own work, support consequential claims with no source of truth, or do work where being indistinguishable from your peers is a defect.

One more asymmetry is worth carrying with you. The tool is often most fluent where the pattern is best represented, which means it can sound most persuasive exactly where it is least differentiated. Fluency is not proof of fit. It is evidence of familiarity.

1. Think about the last significant document your district produced with AI assistance. Which choices in it were yours, and which came from the tool? Can anyone in the room point to a specific one?

2. Where in your operation is the conventional answer exactly what you want, and where is the conventional answer part of the problem you are trying to solve?

3. When you say a person reviewed AI-generated work, what did that person actually check, against what source, and would that process have caught a confident error?

4. Who owns the job of supplying the variance and local judgment? If the answer is everyone, the answer is probably no one.

The context for these decisions is not neutral. Sixty percent of public school teachers reported using AI for work during the 2024 to 2025 school year, and about a third used it weekly. As of early 2026, only 18 percent reported receiving formal guidance from their district about how to use it for learning and instruction. Nearly two thirds of teens report using AI chatbots, and more than half use them for schoolwork.

Adoption is not a decision sitting in front of you. It has already happened. What remains in front of you is the shape of it.

And that shape depends less on how impressive the tool becomes than on whether your district keeps generating its own thinking while using it.

Last week I closed by asking whether we are knowledgeable enough to lead this well. After going one level deeper, I think there is a sharper version of that question.

When the tool gives you an answer that sounds exactly right, do you know enough about your own system to tell whether it sounds right because it fits your district, or because it fits everyone else’s?

That is the skill I think matters. It is not a technical skill, and no vendor can hand it to you. It comes from knowing your own system well enough to recognize when a polished description of it is not actually your thinking.

Next week I want to take this into the operational layer and look at where the repetitive, high-labor work in a district actually sits, and how to tell the difference between work a machine should do and work that only looks that way.

Thanks for reading The Superintendent’s Field Guide! This post is public so feel free to share it.

Share

I am listing these in full because the argument depends on them, and because an article about unverified claims ought to be checkable. Where I could not confirm something to my own standard, I have said so.

On narrative and structural convergence

Russell, J., Rajendhran, R., Pham, C. M., Iyyer, M., and Wieting, J. “StoryScope: Investigating idiosyncrasies in AI fiction.” Conference on Language Modeling (COLM) 2026. University of Maryland and Google DeepMind. arXiv:2604.03136. https://arxiv.org/abs/2604.03136

Caveat that belongs with these numbers. The human stories come from published, professionally edited, anthologized fiction. The AI stories are single unedited generations. The study compares published literary authors to machine first drafts. It does not establish that AI writing is flatter than ordinary human writing, and I have not made that claim. What it establishes is that unedited AI output converges structurally, and that surface editing does not fix it.

On AI and professional work

Dell’Acqua, F., McFowland III, E., Mollick, E., Lifshitz, H., Kellogg, K. C., Rajendran, S., Krayer, L., Candelon, F., and Lakhani, K. R. “Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality.” Organization Science, 2026, 37(2), 403 to 423. https://doi.org/10.1287/orsc.2025.21838

The 12 percent, 25 percent, and 32 percent figures, and the 19 percentage point figure, are from this published version. A widely circulated “40 percent higher quality” figure comes from the 2023 working paper and was revised downward in peer review. I have used the published number.

The idea variability finding is from the earlier version only: Dell’Acqua et al., Harvard Business School Working Paper 24-013, September 2023, Appendix D. https://www.hbs.edu/faculty/Pages/item.aspx?num=64700

Study limitations worth knowing: the 758 participants were split into two non-overlapping groups of 385 and 373; the outside-the-frontier result rests on a single task; quality was scored subjectively; participants were early-career; and the experiment used an April 2023 build of GPT-4 with no browsing, no tools, and no reasoning models.

On what happens when you push back

Walsh, D. “How Generative AI ‘Persuasion Bombs’ Users, and How to Fight Back.” MIT Sloan Ideas Made to Matter, April 28, 2026. Reporting research from MIT Sloan, Harvard University, and the University of Warwick. https://mitsloan.mit.edu/ideas-made-to-matter/how-generative-ai-persuasion-bombs-users-and-how-to-fight-back

On why generic information feels personal

Forer, B. R. “The fallacy of personal validation: A classroom demonstration of gullibility.” The Journal of Abnormal and Social Psychology, 1949, 44(1), 118 to 123. https://doi.org/10.1037/h0059240

Meehl, P. E. “Wanted, a Good Cookbook.” American Psychologist, 1956, 11(6), 263 to 272. The paper that named the Barnum effect.

Dickson, D. H., and Kelly, I. W. “The ‘Barnum Effect’ in Personality Assessment: A Review of the Literature.” Psychological Reports, 1985, 57(2), 367 to 382.

Note on a figure I did not use: a commonly repeated claim that replications still average around 4.2 traces to a single unsourced secondary reference. Replications use different scales and different content, so a pooled average is not meaningful. The effect replicates robustly. The number does not travel.

On adoption in schools

RAND Corporation. “Uneven Adoption of Artificial Intelligence Tools Among U.S. Teachers and Principals in the 2023 to 2024 School Year.” Report RRA134-25, 2025. https://www.rand.org/pubs/research_reports/RRA134-25.html

Walton Family Foundation and Gallup. “Teaching for Tomorrow: Unlocking Six Weeks a Year With AI.” June 2025. Survey of 2,232 public K-12 teachers, March 18 to April 11, 2025.

Gallup. “Most Teachers Receive No Formal Guidance on AI Use.” 2026. Survey of 2,069 public K-12 teachers, February 9 to March 2, 2026. https://news.gallup.com/poll/710534/teachers-receive-no-formal-guidance.aspx

Pew Research Center. “How Teens Use and View AI.” February 24, 2026. Survey of 1,458 teen and parent pairs, September 25 to October 9, 2025. https://www.pewresearch.org/internet/2026/02/24/how-teens-use-and-view-ai/

On the training process

The description of preference training and its effect on output diversity reflects the established understanding of how these systems are built. The quoted passage is from Nate Jones, “Nobody Checked Deloitte’s Report. One Academic Did.,” Nate’s Newsletter, August 5, 2026.

Read the original on superintendentsfieldguide.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.