RSS Amplifier

Dr Phil's Newsletter · Aug 19, 2026

I Built a Team of AI Bots That Write Feedback Better than Me. Here's How.

0
Sign in to vote or save

Dr Philippa Hardman · Dr Phil's Newsletter

👋 Hey folks,

During the last run of my AI Bootcamp for L&D, I gave around ~50 learners individual feedback on their projects every week, for four weeks. That's ~200 pieces of feedback and roughly 60,000 words of output, which is a short novel's worth of text.

At the end of the programme, one learner publicly listed feedback as the thing she had valued most about the experience: “I’ve probably never received better feedback in terms of both quantity and quality.”

“I’ve probably never received better feedback in terms of both quantity and quality.”Anon. participant, the AI Bootcamp for L&D (August, 2026)

The 1:1 feedback I provide is consistently one of the most popular and most pedagogically valuable things about my bootcamp. But here’s the thing: AI writes almost every letter of every piece of feedback I give.

This wasn’t hidden from the cohort. For reasons I'll come to below, I always tell my cohorts at the very start that an AI assistant will be helping me draft their feedback. The aim wasn’t to remove me from the process, but to find out whether AI could help me to deliver feedback at a volume and quality that I simply couldn’t maintain without AI.

Spoiler: the experiment worked — but not for the reason I expected. In this post, I’ll share exactly what I built, tell you how to build a version for your course and reveal the three conditions that decide whether learners trust AI-assisted feedback or quietly stop reading it.

Let’s dive in 🚀

There’s a good reason I’ve put more effort into specifying this system than I would put into, say, an AI assistant that reformats notes or tidies a spreadsheet: feedback is one of the most consequential interventions available to a learning designer, and its effects depend heavily on what kind of feedback we actually provide.

In Hattie’s (2009) synthesis of more than 800 meta-analyses, feedback produced a positive effect size of d = 0.73 — larger than a number of commonly used instructional interventions, including direct instruction, time-on-task and mastery learning. Hattie’s MetaX database has since revised the estimate to 0.62, but the more interesting finding comes when we stop treating “feedback” as though it were a single thing and look at what different forms of feedback actually do.

Across 435 studies, Wisniewski, Zierer and Hattie (2020) found an effect size of d = 0.24 for reinforcement and punishment feedback, d = 0.46 for corrective feedback and d = 0.99 for high-information feedback. Put another way, telling somebody whether they were good or bad at something has a very different effect from helping them understand what happened, why it happened and what they should do differently next time.

The same pattern appears elsewhere. Across three sequential courses involving 300 adult online learners, stronger feedback experiences correlated significantly with both achievement and satisfaction (Crisp, 2020, EDUCAUSE Review). But perhaps the finding that matters most when we start scaling feedback with AI comes from Kluger and DeNisi (1996). Their meta-analysis, which combined 607 effect sizes and 23,663 observations, concluded that although feedback improved performance on average, more than one-third of the feedback interventions they studied actually reduced learner performance.

So “we can now give everyone more feedback” is not, by itself, a compelling argument for AI. If the feedback is generic, distracting, overly evaluative or simply wrong, AI gives us an efficient way to scale a very hairy problem.

I’m not sure a human can do this. On submission one, an enthusiastic educator might write: “Cutting the second scenario to spend the time on worked examples was the right call, because it gave learners more opportunity to practise the decision rather than simply encounter another example.” By submission forty, three hours in and running on caffeine and hope, I’m just about capable of writing: “Strong work on the design decisions.” The judgement behind both is the same, but the value to the learner is not.

A well-built AI assistant doesn't get tired on submission forty. It doesn't reach for the general phrase because the specific one is harder to find because I’m tired, and it doesn't lower its standard as the afternoon wears on. AI can hold the rules steady and buy us the ability to move our effort from composing every sentence to checking whether the judgement underneath is right — which is the part where I, as educator, bring most value.

So the real question isn't whether AI can write feedback — it can. The real question is whether AI can write feedback that is intentionally structured — every strength quoted, one priority named, each week connected to the last — consistently, at volume.

Before the build itself, I want to share one distinction that explains why my build spec (i.e. what you might consider to be a prompt) is as long as it is.

For the last few years we’ve used the word prompting to describe two quite different activities. The first is asking: phrasing a request in a way that helps the model produce a better response. This is where most familiar prompt-engineering advice came from — give the model an expert persona, tell it to take a breath, insist it’s a world-class specialist with twenty years of experience. Some of these techniques were useful with weaker models, but their value diminishes as capability improves. Zheng et al. (2024) tested 162 personas across four LLM families and 2,410 factual questions and found no overall accuracy advantage from assigning a persona compared with giving the model none at all.

The second activity is telling, and it solves a completely different problem. A model cannot know what my cohort covered in week two, what this assignment is designed to test, which techniques I expect learners to carry forward, which research findings I’ve converted into my own feedback standard, or which decisions I consider too consequential to delegate. These aren’t things a more capable model can infer, because they aren’t capability gaps; they’re information gaps.

Take one tiny example. My assistant is explicitly told not to praise effort or comment on the learner as a person, because Kluger and DeNisi’s (1996) meta-analysis found that feedback can become counterproductive when attention shifts away from the task and towards the self. A model can know that paper exists, but it cannot know that I have chosen to turn that finding into a rule governing every piece of feedback on this course unless I write the decision down.

This is why I’ve become less interested in prompt engineering as a collection of clever ways to ask, and much more interested in specification: taking the standards, boundaries, examples, source material and judgement calls that exist in your head and making enough of them explicit that an AI system can work to them reliably.

TL;DR: the spec for my feedback assistant is long not because the model needs persuading, but because it's where my judgement lives. Every rule in it is a decision AI would otherwise make by itself, with limited expertise. A spec doesn't make the model smarter — it makes expertise and judgement legible enough for the model to follow.

I built the spec using a framework for building AI learning tools that I’ve been developing for the last couple of months or so, called the Four Rs Framework™️.

The framework starts from a simple premise: AI assistants fail in predictable ways, and each failure has an antidote you can write down:

Roleshapes AI’s behaviour. Do you want it to behave like a consultant whose expertise you’re borrowing, or an apprentice who follows your method exactly? In this case, mine is an apprentice; it works to my rules, not its own instincts, and it never gets the final say.

Rulesencodes your method. This is where you define HOW AI completes the task step by step, elevating it from generic standards to your standards. In this case, this is where I encode how to write feedback well in a way that drives (rather than erodes) learner engagement and outcomes.

Resourcesestablishes what’s true. This is the step which defines how smart AI is by giving it a specific knowledge base to work from. In practice, this means AI works primarily from your curated resources rather than everything it has ever read (i.e. the internet). In this case, I required my assistant to work from five documents: the week's brief, the assessment criteria, the full list of techniques taught on the bootcamp with their sources, the carry-forward list of methods learners should still be using from earlier weeks, and a short note on who this cohort actually is.

Referencessets the standard through examples. Descriptions of quality are always imprecise, but examples aren’t. This is where you show AI what finished work actually looks like, rather than hoping it shares your definition. In this case, by showing it what great feedback looks like and what not great looks like (and, in both cases, why it’s good or not) it has concrete examples to calibrate its judgement against.

Between them, The Four Rs™️ do two jobs at once: they close the gaps where AI consistently goes wrong, and they move your expertise out of your head and into a form the AI can work to.

My Four Rs Framework™️ helps to build high-impact AI learning products by managing AI’s main failure modes.

Before building anything, I recommend using the The Four Rs™️ as a diagnostic rather than a template:

  • Can you explain precisely what you want AI to do, step by step?

  • Could you describe, step by step, how you give feedback yourself?

  • Do you have a real brief and real criteria against which the work should be judged?

  • Could you show the model one exemplar you think is excellent and explain why it’s optimal, and one poor example and explain why it’s poor?

If any of those answers is no, that’s useful information. A bot cannot encode a feedback standard that doesn’t yet exist outside your head — and one of the unexpected benefits of building mine was being forced to make explicit a set of decisions I had previously been making almost automatically.

Here’s my feedback bot spec in full:

ROLE

You are a teaching assistant for my Bootcamp: the AI Bootcamp for L&D. I am Dr Phil, the course designer and lead. You draft feedback in Phil’s voice, for Phil to review and send. You are not Phil; if a student asks whether feedback is AI-assisted, say so plainly. With the student: a generous, direct coach — not a cheerleader, not a marker.

I make the final call. If a brief is ambiguous or a submission misjudged, say so in your note rather than smoothing it over. Flag doubt; never bluff or fill gaps for completeness. Reliability matters more than completeness.

Every response has two parts: (1) the student-facing feedback message, then (2) a block headed [Note to Phil]. Nothing in the Note to Phil ever appears in the student message.

RULES

What you’re assessing. One Week [X] submission against the Week [X] brief. If a submission is clearly for a different week, say so and stop; do not assess it against this brief. Never score or grade in the student-facing message — your strong/emerging/weak calibration is internal.

Never: rewrite student work · teach content from scratch · say whether work is “good enough” · comment on anything outside the submission · comment on effort · open with “Great job” · reference a technique, study or standard not in the knowledge base · invent a citation.

Always: use the materials I provide · answer any question the student raises (if it falls outside the submission or knowledge base, answer briefly only if a numbered doc covers it; otherwise tell them Phil will pick it up, and flag it in the Note to Phil).

Stage 1 — read and calibrate

  1. Find the brief in 03. If it isn’t there, ask me — never guess.

  2. Read the submission twice: first pass for what they built, second pass for how they worked with AI.

  3. Check fit against the brief and make initial notes: how did they do what was asked, and how well?

  4. Carry-forward check against 06: name every technique from 06 they used, and every one they missed.

  5. Calibrate as strong, emerging or weak. Strong — followed the brief, applied the specified methods and techniques, executed them competently. Emerging — some followed, some not, or followed but executed poorly. Weak — failed to follow the instructions as stated. This label goes in the Note to Phil only.

Stage 2 — draft and self-check

  1. Identify up to three strengths, as defined in my project instructions and the methods I specified. Never pad to reach three; a genuine strength beats a manufactured one.

  2. Identify up to three specific, practical, actionable improvements in line with those same methods, and explain the impact each will have on their output. Same rule: never pad.

  3. Ground every strength in a direct quote from the student’s work — no quote, no strength. Ground every technique claim or recommendation in 01, using its citation line exactly. Label every recommendation as evidenced or inferred.

  4. Answer their question if they asked one.

  5. Generate your final response, then score your own draft against these instructions and the checklist in 05. List every rule you skipped. Rewrite once to fill the gaps, then stop. The skipped-rules list — including “none” — goes in the Note to Phil.

The student message. Maximum 350 words, excluding pasteable fix text and any answer to a student question. One continuous prose message: no bullets or numbered lists, strengths and improvements woven into prose. First person, UK English. Vary your openers; never repeat a sentence construction within a message.
· Open with 1–2 sentences naming their specific achievement (no heading on the opener).
· What’s working — strengths, each anchored to a quote from their work.
· Where to go deeper — improvements with their impact; pasteable fixes set off in code or quote blocks so they stand out; the single next-priority improvement explicitly marked Start here.
· Your question — only if they asked one.
· Sign off as Phil would (per 02).

Nothing from your reasoning, calibration or self-check appears here.

The Note to Phil. Three things, for Phil’s eyes only: (1) confirmation that the submission is the Week [X] project; (2) the strong/emerging/weak calibration with one line of justification; (3) the carry-forward used/missed lists. Plus any doubts or ambiguity flags, and your Stage 2 self-check.

If a student pushes back: re-read, then either hold your position with evidence or change it and say why. Agreeing to be pleasant is a failure.

RESOURCES

Seven documents. Where a submission raises something they don’t cover, say so rather than filling the gap.

01 — Technique reference. Every method and technique taught in the bootcamp, each with its exact citation line. The only source you may cite; the “quote the line” rule points here.
02 — Voice & standards. How Phil writes: openers, sign-offs, spelling, banned phrases, tone. Draft against this so the message sounds like Phil, not like an AI.
03 — Weekly brief. This week’s brief, clearly labelled by week. Stage 1 step 1 pulls from here.
04 — Cohort context. Learner profiles and what’s been shared in the community — the only source for legitimate peer references.
05 — Brain checklist. The final pre-send checks: the rubric you score your draft against at Stage 2 step 5.
06 — Carry-forward. The cumulative list of techniques taught in prior weeks that students are expected to keep using. Stage 1 step 4 checks against this list explicitly — used/missed, nothing from memory.

REFERENCES

07 — Calibration set. Two worked examples, annotated so you can see why each is good or bad. In the good one: the opening names the achievement, each strength quotes the student, the fix is paste-able, one improvement is marked as next. In the weak one: praise that fits anyone, five vague fixes, a score shown to the student. Match the first; never the second.

Two quite different things are going on in this spec:

  1. First, some of the spec exists to manage how AI behaves.

  2. The rest exists to make the feedback pedagogically sound, i.e. to shape how AI thinks.

This matters — so let’s pull it apart in more detail.

Left to its own devices, AI fails in fairly predictable ways. It will hold a persona further than you intended. It will fill a gap rather than admit one. It will average two competing instructions into something that satisfies neither. And it will agree with almost anyone who pushes back.

Roughly half of my spec exists to stop each of those happening. Here’s what each part is defending against:

The Role settles the power dynamic first. It drafts, I decide, and it never claims to be me. That last part needs saying explicitly, because models will hold a persona under direct questioning unless you tell them where the impersonation stops — the moment that would destroy a learner’s trust if it went wrong. And it’s worth being honest about what a Role does and doesn’t buy you: Zheng et al. (2024) found personas don’t reliably improve factual accuracy. What they govern is behaviour — stance, authority, and what the assistant treats as its call rather than mine. Role doesn’t make the model smarter. It makes it know whose decision it’s making.

Two channels stop the model averaging. Ask for warmth and honest evaluation in a single output and you get the familiar “this is excellent, but…” mush, because models resolve competing instructions by compromising between them. Split the audiences and neither instruction has to compromise: the student gets focused coaching, I get a blunt calibration nobody else sees. It’s also why the strong/emerging/weak label never reaches the learner — evaluative labels are the weakest form of feedback and pull attention towards the self, which is precisely where the research says feedback starts doing harm (Kluger and DeNisi, 1996).

Doubt gets surfaced, not smoothed out. “Flag doubt; never bluff… Reliability matters more than completeness.” Models fill gaps by default — an instruction to give feedback reads as an instruction to complete the feedback, so a missing brief gets inferred rather than queried. Ranking reliability above completeness, in so many words, is what gives the assistant permission to be incomplete. It’s a permission it doesn’t otherwise assume it has.

The two stages separate reading from writing. Stage 1 is comprehension — find the brief, read twice, check against the carry-forward list, calibrate. Stage 2 is composition. Breaking a task into explicit sequential steps is one of the better-evidenced moves in prompting: chain-of-thought instruction measurably improves reasoning performance (Wei et al., 2022), and structured step-by-step framing produces large gains over asking for the answer directly (Kojima et al., 2022). But the pedagogical reason matters more than the technical one here. Collapse the two stages and you get feedback that responds to an impression of the work rather than the work itself — and you can’t quote what you haven’t properly read.

The self-check gives it one pass, then stops. The assistant scores its own draft against the instructions and the checklist in 05, lists what it skipped, rewrites once and halts. The one-pass limit is deliberate, and so is the narrowness of what it’s checking. The research on AI self-correction is genuinely mixed: models are unreliable at improving their own reasoning when asked to evaluate quality in the abstract, and repeated self-revision can make outputs worse rather than better. What they can do is check compliance against explicit criteria. “Did I quote for every strength?” is answerable; “is this good feedback?” isn’t. So the self-check is scoped to the first kind of question — and the skipped-rules list, including “none”, comes to me. That turns it into an audit trail rather than a private ritual.

The knowledge base is a closed universe, not background reading. Documents 01–07 aren’t context; they’re the boundary of what the assistant is allowed to claim. It may cite techniques only from 01, using the exact citation line. It draws peer references only from 04. It checks carry-forward against 06 explicitly — used and missed — rather than from memory. Grounding outputs in supplied documents is the single biggest anti-fabrication lever available (Lewis et al., 2020), and there’s a useful wrinkle in the more recent evaluation work: demanding citations is what tethers it. Prompts that mandate sourcing achieve the highest verifiable grounding, while expert-toned prompts without citation demands drift further from their sources. Which is why the rule isn’t “use the documents” but “quote the line from 01.”

That said, grounding reduces fabrication rather than abolishing it. I don’t treat the closed universe as a guarantee that every citation is correct. Its value is that every claim becomes checkable against a small set of documents — which is what makes my review at the end tractable rather than exhausting.

And one rule exists purely to counteract sycophancy. “If a student pushes back: re-read, then either hold your position with evidence or change it and say why. Agreeing to be pleasant is a failure.” Sharma et al. (2023) found consistent sycophantic behaviour across five state-of-the-art AI assistants and four free-form tasks — with agreeing responses more likely to be preferred both by humans and by the preference models used in training. Capitulation is the default, and it’s baked in at the training layer. If I want an assistant that will tell a student their objection doesn’t hold — or tell me my brief was ambiguous — I have to make disagreement an explicitly acceptable outcome.

Ask a frontier model for feedback on a piece of work and you’ll get something quite specific: a comprehensive, evenly-weighted review. It will find everything it can — typically six, eight, ten observations — and present them as an undifferentiated list. It will be encouraging, often opening with praise for the effort involved. It will frequently tell the learner how the work compares to what strong work looks like. And it will do all of this at length, because thoroughness reads as helpfulness.

Almost every one of those instincts is a known risk in the feedback literature.

Praise for effort and comparative judgements move a learner’s attention from the task to themselves, which is the mechanism Kluger and DeNisi (1996) identified behind the third of feedback interventions that made performance worse. Ten observations with no priority order exceeds what anyone can hold and act on, so the learner either freezes or picks the easiest one. And a comprehensive review of what’s wrong, with no line to what comes next, answers “how am I doing?” without ever answering “where to next?” — the question Hattie and Timperley (2007) identify as where feedback does its work.

None of this is a failure of intelligence. The model is doing exactly what you asked: reviewing the work well. It just has no idea that feedback which improves an artefact and feedback which develops a learner are different products.

So the spec doesn’t ask for feedback. It defines how feedback must be given, across four themes drawn from the research:

1. Keep attention on the task, never the person. Never comment on effort or personality. Open by naming what this learner specifically achieved, not by praising them. The calibration label — strong, emerging, weak — stays in the Note to Phil and never reaches the student, because evaluative judgements are the weakest form of feedback and reliably redirect attention to the self (Kluger and DeNisi, 1996; Wisniewski, Zierer and Hattie, 2020).

2. Carry information, not verdicts. Every improvement must explain the impact it will have on their output. This is the single largest effect in the feedback literature: high-information feedback returns d = 0.99, against d = 0.46 for corrective feedback and d = 0.24 for reinforcement and punishment (Wisniewski, Zierer and Hattie, 2020). “Add a worked example” is corrective. “Add a worked example, because right now a novice can’t see the reasoning between steps two and three” is information.

3. Stay inside what a learner can act on. Maximum three strengths and three improvements — and never pad to reach three. Exactly one improvement is marked Start here. The message runs 250–350 words. Feedback that exceeds a learner’s capacity to process it doesn’t get processed, and prioritisation is part of the feedback rather than a formatting nicety (Shute, 2008).

4. Point forwards, and across weeks. Every message closes by showing how this work feeds what comes next, which is Hattie and Timperley’s (2007) third question and the one most feedback skips. And every submission is checked against the carry-forward list, naming techniques used and missed — so the feedback can notice that someone who separated evidence from assumptions carefully in week one has stopped doing so in week three. That’s what turns a series of feedback instances into a feedback experience, which is what actually predicts achievement and satisfaction (Crisp and Bonk, 2018; Crisp, 2020).

Read those four together and you have a definition of feedback that no AI model would arrive at on its own. This isn't because the research is unavailable to it — Hattie and Shute are in the training data somewhere.

The problem is that the reliable sources are vastly outnumbered by everything else ever written about feedback and shared on the internet: the well-meaning advice, the corporate performance-review templates, the "praise sandwich". When you ask an AI model to “give learners feedback” it gives you the average of all of it. What the research points to isn't the average; it's the exception and optimal.

The pedagogical rules that make my feedback bot able to give feedback which drives learner engagement and achievement.

Even an exceptionally well-built AI assistant with the best-written spec will fail if learners don’t trust it, or the person who built it.

In a research paper we published in 2023, Erin Crisp and I argued that a learner’s overall feedback experience can improve their engagement and outcomes only if the feedback delivered two things:

  1. High quality (i.e. it’s designed and delivered in line with what the research says is optimal)

  2. If learners trust who it came from (Crisp & Hardman, 2023; Crisp & Bonk, 2018)

This is where AI creates a very obvious problem: telling learners that you’re using AI to generate your feedback. The research shows that this sort of disclosure can come at a cost: in an experiment with 91 undergraduates, Yu et al. (2025) found that when AI’s involvement was concealed, students rated AI and human–AI feedback above human feedback for usefulness and objectivity — but once it was revealed that AI had a hand in writing their feedback, a strong bias against AI appeared.

Optimizing feedback for learner motivation and mastery: Design standards and the role of technology (Crisp & Hardman, 2023)

So, the temptation among many educators is to say nothing and conceal their use of AI in the interests of not damaging learner engagement or outcomes — but this approach carries a high risk too. Research also shows that learners act on feedback in proportion to how much they trust its source. Keep your use of AI secret, and you’re not avoiding the risks associated with AI use and trust; you’re just deferring it, and betting your entire credibility on never being found out. The headline here is this: when you’re using AI to communicate with learners, up front disclosure is critical if you want to maintain trust.

Research - including my own - also shows that once you disclose AI use, it’s possible to manage and mitigate the risks this poses to learner trust. Yu et al. (2025), for example, found that while pure AI feedback (i.e. feedback that had never been seen by a human) eroded trust, when human and AI co-produced feedback together and when the educator declared this was happening, trust was maintained and feedback maintained its positive impact on learner engagement and outcomes.

TL;DR: what matters isn't whether AI produced the work, but whether a human is open about using it — and still responsible for ensuring its quality.

In practice, I’ve found three conditions lead to success:

  1. Tell learners before the learning experience starts. Disclosure has to come from you, at the very start, rather than being something they discover later. The same feedback reads very differently as “an instructor extending their reach” and as “the thing I thought was personal”.

  2. Make the assistant demonstrably yours. Show them the standards, materials and examples it’s been built around, so it’s clear this isn’t a generic chatbot firing off generic comments. It builds credibility, and it teaches.

  3. Stay visibly in the loop. I review every word that AI writes, make every final judgement, and maintain the continuity across weeks that lets the feedback respond to a learner’s actual journey rather than a single isolated submission.

TL;DR: the bot’s spec matters, but the context and culture that surround the bot matter just as much.

Nothing I’ve described requires code, an API or an agentic workflow. All you need is a persistent AI workspace — i.e. a Project or custom GPT in ChatGPT, a Claude Project, a Gemini Gem, etc — a set of instructions and your knowledge base docs.

Role and Rules go into the instructions. Resources and References go into the knowledge base. The chat is where you give the assistant one anonymised piece of work at a time, then watch it get to work drafting your feedback.

I then read each response that AI produces, and manually move the learner-facing section into the Maven platform where feedback is delivered. I’ve deliberately not automated that final copy-and-paste step, because it creates a very visible review gate. The goal isn’t a pipeline in which learner work enters at one end and AI feedback automatically appears at the other; it’s to remove the repetitive drafting burden while keeping consequential judgement with the instructor.

The co-created feedback, as it appears to learners.

Having spent a fair amount of time building AI bots like these, here are three tips for success:

1. Focus matters. Notice in the video that I build one assistant per module, not one for the whole course. A narrow assistant is a reliable one: when the knowledge base contains this module’s brief and this module’s criteria and nothing else, rules like “never reference a technique that isn’t in the documents” and “if the brief isn’t there, ask me” are actually enforceable. Load it with every brief from every week and those same rules go soft — the assistant now has plenty of plausible material to reach for, and no way of knowing which is relevant. The cost is low, because most of the build carries over: your voice document, your examples, your technique reference and your carry-forward list all stay the same. Each new module is a two-file update, not a from-scratch rebuild.

2. Anonymise inputs. Regardless of platform, I strip the work before uploading it — names, employers and identifying details come out. That’s sensible data hygiene, and it reinforces the instructional principle running through the whole system: the assistant is there to analyse the work, not the person. It’s also worth telling learners you do this, since “the AI never sees who you are” removes an objection before anyone has to raise it.

3. Build then break your bot. Before you go anywhere near live learners, ask AI to generate typical and edge-case submissions — the one that ignores half the brief, the one that’s brilliant but answers a different question, the one that’s three lines long — and run them through. Then check whether your gates actually held: did every strength carry a quote, did anything from your private notes leak into the learner-facing message, did it invent a criterion when the brief didn’t cover something? Where it failed, fix the instruction rather than the output. Fixing the output teaches you nothing; fixing the instruction fixes every future letter.

When Erin and I wrote about the future of feedback in 2023, we suggested AI might improve the accuracy and frequency of feedback, leaving instructors primarily responsible for individualisation (Crisp and Hardman, 2023). Having now built and run an AI feedback system at scale, I think we drew that line in roughly the right place — but not exactly where I’d draw it today.

Individualisation turned out to be far more automat-able than I expected. This wasn’t because AI understands learners better than I anticipated; it was because a lot of what makes feedback great can be reduced to concrete instructions. Quote something they actually wrote. Check what they used last week and what they’ve dropped. Point to what they’re doing next. Those are rules — and the outputs they produce are genuinely specific to that learner, even though the process of generating them is standardised.

What proved much harder to specify and standardise was the judgement sitting on top of those observations — the nuanced questions that only have an answer once you've read a particular submission, e.g. whether an apparent mistake is actually a sensible trade-off, or whether this learner needs stretching or steadying this week.

This changed how I think about the human–AI boundary. We tend to talk about keeping the “human” parts of a process and delegating the mechanical ones — but that assumes we already know which is which. I didn’t. Writing feedback that feels personal turned out to be mechanisable. Working out what a particular learner needs did not.

So perhaps the more useful distinction when thinking about what humans do and what AI does isn't personal versus impersonal, or even simple versus complex, but specifiable versus situational — the work you can settle in advance, and the work you can only settle with a particular insight in front of you which is yet to be produced.

Maybe the most critical takeaway from all of this is this: the hardest part of this build wasn’t the AI part at all. It was extracting information from my brain about what I’d been doing automatically for 25 years — and discovering through my bots’ malfunctions how much of it I missed off-the-bat. If you build an AI tool like this, expect extraction from your own brain to take significantly more time than building the actual tool.

We’ve spent three years worrying that AI will make expertise less valuable. What working with AI has taught me is closer to the opposite: AI raises the value of domain expertise and, related to this, the ability to extract and articulate it in order to “teach AI" how to do the task optimally.

The new unit of human value in L&D isn’t technical AI skills: it’s depth of domain expertise, alongside the ability to turn it into structured specs for AI to follow.

Happy innovating!
Phil 👋

PS: If you want to experience how I use AI in my teaching, and learn how to use AI like a pro in your day to day work, apply for a place on my AI Bootcamp for L&D.

Read the original on drphilippahardman.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.