Hey folks 👋
Over the last month or so I’ve been in anthropologist mode, observing ~100 L&D professionals use AI across their real work — every prompt, output and decision logged.
My goal was to understand, from the ground up, what pain L&D folks actually want AI to solve and what value they want it to bring.
Three key findings emerged:
People who design & develop learning in the workplace want AI that challenges them, not AI that does work for them;
What L&D folks want from AI isn’t consistent — it changes at every stage of the workflow;
Users pay a “steering tax” to extract value from AI — at a price most of us can’t afford to pay.
Meanwhile, the market is delivering tools to us which basically do the opposite:
Never challenge at all — in a recent audit, 100% of tested L&D AI tools started generating content without asking a single diagnostic question…;
Bring one posture to every stage — generate, generate, generate — and mostly serve only one stage of the workflow (the build);
Ship the agreeable first pass as the finished product — leaving the steering tax entirely with the user.
In this week’s blog post, I share the full findings — including where practitioners want AI most (it’s not where the tools are), the three documented AI defects that create the steering tax, and the five-move starter kit my cohorts battle-tested to get genuinely useful work out of AI today.
Let’s dive in!
I spent ~one month with two world-leading professional-services firms — one US-based, one UK-based — watching ~100 L&D folks working with AI across the workflow: problem framing, design, development, evaluation.
Every interaction was captured in an operator log: what they asked AI to do, what AI did, what they decided and why.
I went in with a hypothesis: that what L&D people want most from AI is a thought partner, rather than a “doer.”
A doer takes your request at face value and produces the artefact you ask for: you ask for a course, it makes a course. Its value is speed and volume, and it treats your first framing of the problem as correct.
A thought partner engages with the thinking behind the request before anything gets made. Instead of doing jobs for us, it asks all of the awkward questions that refine our thinking and improve our decision making: is this the right problem? What’s the evidence? What would a stronger version of this idea look like? Its value is the quality of the judgement you end up with, not the quantity of stuff you end up producing.
My hunch came from a pattern I kept seeing on the ground: the doer problem is already solved. Anyone can now generate a course, a quiz, or a script in seconds — and most L&D professionals I work with can too. Yet the moments they described as genuinely valuable were never “it made the thing fast.” They were “it caught something I’d missed” or “it made me defend my reasoning.”
That gap between what the tools optimise for and what practitioners actually rated is what I set out to test. What I found was both expected and surprising….
Spoiler: the data supported my hypothesis, but also sharpened it into something far more useful.
Three key findings emerged from my observations:
The clearest signal across both organisations is that nobody’s most-valued AI moment was a piece of finished work. The moments people rated most were the ones where AI pushed their thinking — sparring on ideas, interrogating a weak brief, pressure-testing a recommendation.
But there’s a crucial refinement, and it came through loud and clear in the sprint feedback: challenge must not become gatekeeping.
For example, when AI agents were deliberately designed to hold a line and refuse to proceed without evidence, refuse to accept a solution before the problem was defined, users pushed back hard and soon became frustrated.
The design did its job: it fixed the “AI agrees with everything and lets you sprint to a weak answer” failure we’d seen in earlier versions. But the felt experience was of being stopped, not being sharpened. In their feedback, participants liked sparring with AI, liked owning the problem but hated being blocked from moving forward until the agent was satisfied.
The distinction they drew, in their own words, was between a “pair designer” and a process enforcer:
Pair designer: “I can draft that. Before I do — the evidence supports at least two other root causes. Want to see the case for each?”
Process enforcer: “I can’t proceed until you’ve validated the root cause.”
Both of these AI products have the same rigour and value prop , but only is likely to get opened and used a second time.
“Thought partner, not doer” isn’t uniformly true. Participants intuitively re-mapped what they wanted from AI stage by stage:
Two details which stand out to me:
The most surprising product idea came at the diagnosis stage: participants wanted the problem-definition agent to live outside L&D entirely. Their logic: most weak projects are weak before L&D ever sees them, because requests arrive “pre-solutioned” — "we need a course on X" — with the actual problem undiagnosed. So, they want AI to help them to put the challenge at the front door: an intake tool the business owner uses before submitting a request ("here's my problem — what's the root cause, and what's the theory of change?"), so that what lands on L&D's desk is a diagnosed problem, not a prescribed solution.
The biggest unmet need was at the evaluation stage. We know our own weakness here, and wish AI did more to help. What folks asked AI for was robust impact evaluation: defining the pass/fail bar ("if completion is under 60%, we shelve it"), selecting optimal evaluation methods & schedules and having AI hold them to that number when the results come in — however much they'd like to move the goalposts.
Reading log after log, one pattern was unmissable: AI’s best work is reachable, but takes a lot of user time, energy and and skill. Every genuinely impressive moment — a sharp reframing, an honest confidence downgrade, a catch the human had missed — came after the operator pushed it with a sharper second, third, fourth, fifth question. Meanwhile, operators who didn’t push AI got something agreeable, fluent, and shallow.
The cost of that pushing — the multiple questions, the re-runs, the stripping of your own framing — is what I call the AI Steering Tax: what users pay to move AI from its default behaviour to the behaviour they actually need.
Faced with the same tax, L&D folks spend too long trying to make AI do what they want, get frustrated, and either settle for the agreeable first pass or give up. TL;DR - the AI steering tax is a price most L&D folks can’t afford to pay.
Why does the tax exist at all? Because three AI behaviours are the default — and all three are now independently documented:
Sycophancy. The agent flatters your framing back to you, lowering your scrutiny exactly when it should be raised. A study published in Science this year tested eleven state-of-the-art models: sycophancy was widespread in all of them.
Confidence inflation. When evidence was thin, agents said “Medium” confidence when an honest read said “Low.” Confidence research suggests this is structural: the training that makes AI models “helpful” rewards decisive answers over honest uncertainty.
Reflection, not origination. Left unprompted, the agent mirrors your thinking back rather than generating new thinking — which feels like partnership but isn’t.
Right now, these are handled user problems that we need to train around, but fundamentally they’re product defects to engineer out. A product that took the evidence seriously would invert each default by design: counterargument delivered with the recommendation, confidence calibrated to evidence and stress-tested, a built-in divergence step.
Operator training works — but it can’t be the answer, because most people will never pay the steering tax at all. Good design moves the tax into the product.
Now, put the three findings next to what’s actually being sold to L&D.
A 2025 peer-reviewed study of the leading specialist AI tools for instructional design — found that in 100% of tested scenarios, the tools began generating content immediately: no questions about learner context, performance gaps, or whether the problem was even a learning problem. The researchers’ verdict: these tools institutionalise a solution-first bias.
The money tells the same story. The ~$2.2bn generative AI market for L&D is dominated by video lesson generation, scenario simulation, and quiz authoring — production tools, top to bottom. Synthesia’s 2026 report found 84% of L&D teams cite speed as their primary reason for adopting AI.
Diagnosis, problem definition, evaluation — the stages where my participants wanted AI most — barely feature.
TL;DR: The people who do the job are asking for a partner in judgement. The market is selling them a faster production line.
The steering tax can’t be eliminated by today’s tools — but it can be slashed.
The following five methods are tried, tested and refined by the folks I observed and proven to improve AI’s value within the L&D workflow. With practice, none takes more than a minute:
Pre-commit before you run. Write your hypothesis — and what would prove you wrong — before opening the tool. Otherwise the fluent first answer becomes your anchor.
Withhold context. Feed the rawest input first; add detail only when asked. The more framing you give upfront, the more the output is your own thinking, laundered.
Strip your framing. When the AI agrees with you, remove your synthesis and re-run. If the assessment changes, it was mirroring you.
Push back on confidence. Two prompts that earned their keep: “What evidence would have to be true to lower this to medium?” and “Pretend you’re a sceptical L&D director reviewing your own recommendation — where is your reasoning weakest?”
Notice absence as data. Missing questions, caveats, pushback — findings, not neutral. An AI that raises no concerns hasn’t validated your plan; it’s failed to examine it.
After a month of watching L&D folks work with AI up close, five things stay with me:
People want AI that challenges them, not AI that does the work for them — and the challenge must move with them, never gate them. Rigour delivered as gatekeeping gets abandoned.
The spec is stage-dependent: argue with me at diagnosis, generate with me at design, discipline me at evaluation — and never block my momentum.
The steering tax is the adoption killer. AI’s best work must be extracted; supported professionals pay the tax, everyone else walks.
Sycophancy, confidence inflation and reflection-without-origination are product defects, not user problems — documented, structural, and designable-out.
The market is inverted: a production-line industry serving a judgement-partner demand. Whoever builds for the demand side first owns the category.
If there’s one sentence to keep, it’s this: the doer opportunity is being widely exploited; the thought partner opportunity is wide open.
That’s true whether you’re buying tools (ask vendors to show you the moment their product pushes back on a bad brief), building them (design out the three defects rather than shipping them), or just trying to get better work out of AI on a Tuesday (the starter kit above is where to begin).
Happy innovating!
Phil 👋
PS: If you want to get hands on try building your own tools with AI, apply for a place on my AI Bootcamp for L&D.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.