At 4:47 on a Tuesday afternoon, a junior analyst sends a board paper to her manager. It is the strongest-looking document she has produced in six months: clean structure, plausible numbers, crisp recommendations. An artificial intelligence assistant has summarised the source material, suggested the risk categories, rewritten the awkward sections and flagged what appeared to be the three most important issues.
Her manager is impressed. The paper takes half the usual time.
Then he asks a routine question: “Why did we exclude the fourth option?”
She pauses.
The answer is probably somewhere in the source pack. She remembers seeing the option, but not deciding against it. The model had compressed five alternatives into three and she had worked from there. She can reopen the documents, trace the reasoning and reconstruct the decision. It will simply take longer than producing the paper did.
Nothing has failed. No hallucination has been discovered. The output may even be right.
The unsettling possibility is that the organisation has gained a faster analyst before it has established whether it is still developing an analyst.
The question is not only whether artificial intelligence improves output. It is whether it improves – or at least preserves – the human who remains responsible for that output.
The productivity case for generative artificial intelligence is no longer hypothetical. In one of the strongest workplace studies, Erik Brynjolfsson, Danielle Li and Lindsey Raymond examined 5,172 customer-support agents using a conversational assistant. Access increased issues resolved per hour by about 15 per cent on average, with the largest benefits among less experienced and lower-skilled workers. Novices became productive faster, customer interactions improved, and the researchers found evidence that some gains reflected durable learning rather than mere dependence on the tool. During temporary outages, workers who had used the assistant retained part of their improvement. That is important evidence against a simple “AI makes people stupid” story (Brynjolfsson, Li & Raymond, 2025). [1]
Other experiments point in the same positive direction. Earlier controlled work found that ChatGPT reduced the time professionals spent on common writing tasks while improving average quality. More recent educational experiments show that artificial intelligence can support genuine learning when people use it to explain, search and augment their own reasoning rather than replace it. In a 2026 randomised study of undergraduates learning unfamiliar material, access to generative AI raised immediate knowledge-test scores by 0.27 standard deviations, and the gains persisted a week later. Students who used AI for explanation and exploration showed stronger delayed benefits than those who largely used it to generate text (Contractor & Reyes, 2026). [2]
So the issue is not assistance. It is:
what kind of assistance, replacing what kind of effort, at what stage of skill formation.
That distinction is easy to miss because organisations usually observe performance before they observe capability. A report is finished, a ticket is closed, code compiles, a student submits the assignment. The visible artefact gets better immediately, but the invisible stock of human skill changes slowly.
This produces a measurement trap: the short-run benefit and the long-run cost can sit on different timelines.
A useful distinction is between performance extraction and capability formation. Performance extraction asks: how much good work can this person-plus-system produce now? Capability formation asks: what is the person becoming able to notice, understand, challenge and do later? The two can rise together, but they can also diverge.
The divergence is clearest in learning research. Hamsa Bastani and colleagues ran a field experiment with nearly 1,000 high-school mathematics students. Students using an unrestricted GPT-4 interface performed much better during practice, but after the tool was removed they performed 17 per cent worse than students who had never had access. A second version was deliberately designed as a tutor: it constrained direct answer-giving and encouraged the steps needed for learning. That version preserved much of the immediate performance advantage while largely mitigating the later learning loss. The lesson is not that AI tutoring fails. It is almost the opposite: interface design can determine whether the machine behaves like a lift or a staircase (Bastani et al., 2025). [3]
Early workplace evidence suggests a similar mechanism. In a small randomised study of 52 mostly junior software engineers learning an unfamiliar Python library, researchers found that developers given AI assistance finished only slightly faster, but scored 17 percentage points lower on a subsequent mastery test. The largest deficit was in debugging – precisely the skill needed to recognise when AI-generated code is wrong. The small sample and short follow-up mean the result should not be generalised too far. More interesting than the headline was the variation inside the AI group: people who delegated code and debugging tended to learn less, while those who asked conceptual questions, sought explanations and checked their understanding performed better (Shen & Tamkin, 2026). [4]
This is a plausible mechanism for skill atrophy. Expertise is not simply a database stored in the head. It is partly a history of encounters with resistance: the bug that took an hour to isolate, the awkward client whose objection forced a better explanation, the source document that did not fit the summary, the calculation that had to be reconstructed from first principles. Some friction is waste. Some friction is practice.
Generative AI is unusually good at removing both.
That creates what we call delegation creep. A worker first delegates formatting, then drafting, then synthesis, then interpretation, then the first recommendation. No individual step feels reckless, and each saves time. But the boundary between “the machine helps me do this” and “the machine does the part through which I used to learn this” moves gradually. The worker may remain accountable for a process in which fewer of the capability-building steps remain human.
Recent survey research adds a psychological clue. A 2025 Association for Computing Machinery study asked 319 knowledge workers for 936 examples of real generative-AI use. Participants reported critical thinking in about 60 per cent of cases. Higher confidence in the AI’s ability to perform a task was associated with less reported critical-thinking activity; higher confidence in one’s own ability to perform or evaluate the task was associated with more. The study relied on self-report, so it cannot establish that AI caused reduced critical thinking. But its pattern is revealing: as work shifts from production to supervision, the ability to supervise depends increasingly on expertise that easy delegation may stop exercising (Lee et al., 2025). [5]
Preliminary 2026 experiments push the concern one step further. Across randomised studies involving 1,222 participants performing mathematical and reading tasks, AI assistance improved immediate performance, but participants performed worse and gave up more often after access was removed. The effects appeared after only brief exposure. These results came from a preprint, not settled evidence of long-term deskilling, and a ten-minute laboratory task is obviously not a career, but it identifies a mechanism worth watching: assistance can change not only what people know, but their expectation of how quickly difficulty should ‘go away’ (Liu et al., 2026). [6]
That matters because expertise requires persistence under ambiguity. A senior lawyer, engineer, clinician, investigator or policy analyst is valuable partly because they know what to do when the first answer is incomplete. If the interface constantly converts uncertainty into a fluent next step, users may get less practice inhabiting the uncomfortable interval between “I do not know” and “I have worked out how to know”.
The economic problem follows from the psychological one. Luis Garicano and Luis Rayo model a familiar professional bargain: junior workers historically perform lower-level tasks while learning from them and from senior review. If AI absorbs a growing share of those tasks, the work may become cheaper without preserving the pathway by which juniors become seniors. Their 2025–26 theory does not prove that apprenticeship systems will collapse; it identifies the conditions under which they could. The risk is structural: firms can rationally automate work that looks low-value today even when that work carries hidden training value for tomorrow (Garicano & Rayo, 2026). [7]
This changes how productivity should be understood. A firm can increase output per worker while consuming a stock of expertise it is no longer replenishing. The analogy is closer to drawing down capital than improving a machine. A 2026 theoretical model by Michael Caosun and Sinan Aral formalises that possibility: where AI displaces the practice through which expertise is maintained, short-run productivity can rise even as long-run capability falls. Their “augmentation trap” is a model, not an empirical estimate of how often this occurs, but it captures an incentive problem familiar to any organisation governed by quarterly targets: the productivity gain arrives now; the skill loss may become visible after the manager, worker or vendor has moved on (Caosun & Aral, 2026). [8]
There is another twist. The skills most vulnerable to neglect may be the skills required for meaningful human oversight.
Consider code. If AI writes more code, manual typing matters less. But code reading, architecture, debugging and knowing when a proposed solution violates an unstated constraint matter more.
Consider analysis. If AI can summarise 400 pages, summarisation itself may become less valuable, while provenance checking, spotting omitted alternatives and questioning the frame become more valuable.
In medicine or law, the same pattern can move the human role from primary production towards exception handling and judgement.
That sounds like an upgrade – and it can be. But exception handling is an expert function. We should be wary of designing career systems in which juniors are asked to supervise processes they were never allowed to practise.
This is where the “human in the loop” can become misleading. Human presence is not the same as human agency. A person who lacks the skill, time, source access or organisational permission to challenge an output is not meaningful oversight; they are a signature surface or moral scapegoat.
European law is beginning to recognise part of this problem. Under the European Union Artificial Intelligence Act, providers and deployers must support AI literacy among relevant staff, and deployers of high-risk systems must ensure staff involved in human oversight are appropriately trained. As of August 2026, the general AI-literacy obligation is enforceable, while key high-risk rules for areas including employment and education are scheduled to apply from December 2027 after the 2026 amendments. The law therefore treats human competence as part of AI governance, not merely a training benefit. But it does not solve the broader productivity problem: ordinary writing assistants, coding copilots and summarisation tools can reshape professional capability long before a system qualifies as “high risk” in law (European Commission, 2026). [9]
The deeper governance test is therefore not whether AI is used, or even how often. It is whether the workflow preserves the capabilities that remain necessary when the system is wrong, unavailable, conflicted, attacked or confronted by a genuinely novel case.
That suggests a different definition of uplift. Output quality and speed are only the first layer. A mature evaluation should also ask whether the person can still trace the evidence, formulate an independent view, detect a plausible error, perform critical parts unaided, and understand when responsibility cannot safely be delegated. If these measures deteriorate while throughput rises, the organisation has not necessarily become more capable. It may simply have moved capability from people into an external dependency.
Of course it would be a mistake to protect every old task simply because it once trained somebody. Spreadsheets displaced mental arithmetic; search engines displaced some recall; computer-aided design replaced manual drafting. Societies routinely stop practising skills that technology makes unnecessary.
The relevant question is not whether a skill declines. It is whether the declining skill remains necessary for judgement, recovery, accountability or future learning.
The evidence also shows that AI can compress learning curves rather than destroy them. In the customer-support study, less experienced agents acquired useful practices from the system and retained some gains during outages. In the 2026 university experiment, students who used AI as an explanatory tool showed persistent learning. In the high-school mathematics trial, changing the tutor design largely prevented the unassisted-performance penalty. These findings point away from prohibition and towards scaffolding: give enough help to extend capability, but not so much that the learner is removed from the cognitive work that capability requires. [10]
The industrial question was often how many units a worker and a machine could produce together. The conversational-cognitive question is more recursive:
What kind of worker does repeated collaboration with the machine produce?
The first risk is not unemployment but competence without depth. Workers may remain employed and highly productive while becoming increasingly dependent on systems they cannot fully inspect. This is especially consequential in professions where rare failures matter more than average throughput: security, engineering, medicine, finance, law, public administration and critical infrastructure. The normal case trains confidence; the exceptional case reveals whether that confidence was grounded.
The second is apprenticeship hollowing. Organisations may remove routine junior work because AI can perform it cheaply, then discover years later that those supposedly low-value tasks were where tacit judgement, domain vocabulary, exception recognition and professional norms were acquired. Garicano and Rayo’s model makes the intergenerational issue explicit: automation can alter not only the number of jobs but the economics of becoming qualified to hold the senior ones. [7]
The third is oversight theatre. Regulation, boards and customers may be reassured that a human approves the result even though that human increasingly lacks the independent competence required to disagree. This risk compounds because high confidence in AI can reduce reflective effort, while skills needed to detect errors may weaken when they are repeatedly offloaded. The danger is not that every worker stops thinking. It is that institutions continue to treat nominal human review as a stable control while the capacity behind that review changes. [11]
The fourth is inequality of agency. Strong experts may gain enormously from AI because they can interrogate outputs, recognise bad assumptions and use the tool as leverage. Less experienced users may obtain similar-looking output without acquiring the same underlying model of the problem. A very recent randomised experiment with 1,174 adults found that generative AI sharply narrowed education-based performance gaps while the tool was available, but a substantial gap returned when it was removed; follow-up performance improved most when intensive AI use was paired with sustained human effort. That is an encouraging result and a warning at once: AI can democratise performance faster than it democratises independent capability (Cruces et al., 2026). [12]
The fifth is trust decay after invisible dependency. When users discover that a professional, teacher or institution cannot reconstruct an AI-assisted conclusion without the tool, trust can fail abruptly. The problem is not simply that AI was involved. It is that responsibility was presented as human while the reasoning capacity behind it had migrated elsewhere.
Finally, there is a governance blind spot created by measurement itself. Organisations can count hours saved, cases closed and documents shipped. Skill retention, challenge quality, independent performance, source contact and recovery capability are harder to measure. What is easy to measure therefore becomes what is managed. The system can look increasingly successful while the organisation becomes less reversible.
Productivity without agency is not an inevitable outcome of artificial intelligence. The current evidence is too mixed, too domain-specific and too young for that claim. We have credible demonstrations of faster work, better novice performance and genuine learning.
We also have increasingly credible demonstrations that unrestricted answer-giving can reduce subsequent independent performance, persistence and mastery. The tension is not between “AI optimists” and “AI pessimists”. It is between two different definitions of success. [13]
The more defensible organisational response is therefore modest: measure both assisted performance and retained capability where retained capability still matters. Periodically test whether people can perform critical elements without the system. Preserve source contact and independent first-pass reasoning in high-consequence work. Design learning modes that explain and scaffold rather than simply complete. Treat a human approval step as meaningful only when the person has the competence and authority to refuse.
None of those measures requires pretending that manual work is morally superior. They simply recognise that some friction is the workshop in which judgement is made.
The conversational-cognitive revolution changes the productivity equation because the machine can now enter almost every stage of knowledge work: first draft, first explanation, first diagnosis, first recommendation, first check. Used well, that can give people capabilities they did not have before. Used carelessly, it can make the artefact stronger while leaving the author less able to stand behind it.
That is the line worth measuring.
The next article in this series will examine why institutions are poorly equipped to see it: when dashboards show speed, adoption and output, what tells us whether the human system underneath is becoming stronger or merely more dependent?
Bibliography
Brynjolfsson, E., Li, D. & Raymond, L. (2025). “Generative AI at Work.” Quarterly Journal of Economics, 140(2), 889–942. [14]
Noy, S. & Zhang, W. (2023). “Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence.” Science, 381, 187–192. [15]
Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö. & Mariman, R. (2025). “Generative AI without Guardrails Can Harm Learning.” Proceedings of the National Academy of Sciences, 122(26). [16]
Lee, H. P. H. et al. (2025). “The Impact of Generative AI on Critical Thinking.” Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. [17]
Shen, J. H. & Tamkin, A. (2026). “How AI Impacts Skill Formation.” arXiv:2601.20245. [18]
Liu, G., Christian, B., Dumbalska, T., Bakker, M. A. & Dubey, R. (2026). “AI Assistance Reduces Persistence and Hurts Independent Performance.” arXiv:2604.04721. [6]
Contractor, Z. & Reyes, G. (2026). “Experimental Evidence on the Learning Impact of Generative AI.” IZA Discussion Paper / arXiv:2607.08849. [19]
Cruces, G. et al. (2026). “Does Generative AI Narrow Education-Based Productivity Gaps? Evidence from a Randomized Experiment.” arXiv:2608.04198. [20]
Garicano, L. & Rayo, L. (2026 revision). “Training in the Age of AI: A Theory of Career Viability.” CEPR Discussion Paper 20634. [7]
Caosun, M. & Aral, S. (2026). “The Augmentation Trap: AI Productivity and the Cost of Cognitive Offloading.” arXiv:2604.03501. [21]
European Commission (2026). “AI Literacy – Questions & Answers” and “AI Act: Regulatory Framework.” [9]
Neural Horizons reference document context: Cognitive Susceptibility Taxonomy and Robo-Psychology Taxonomy
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.