RSS Amplifier

The Learning Dispatch · Jun 8, 2026

The Monthly Dispatch — What's New in Learning Science? — June 2026

0
Sign in to vote or save

Carl Hendrick · The Learning Dispatch

New research on retrieval practice shows it stalls under high cognitive load. Generic thinking skills again fail to transfer to related tasks. Removing a cognitive tool, it turns out, hurts more than adding it ever helped. And the most cited paper behind the claim that AI improves learning, Wang and Fan's meta-analysis, has been retracted by Nature.

One of the most consistent findings in learning science is that retrieval practice outperforms restudy. But what happens when the material being retrieved is genuinely demanding?

Redifer, Myers, Bae, Naas and Scott (Instructional Science) had undergraduates study a lengthy, conceptually demanding research article, comparing retrieval practice (free recall, practice quizzing and question generation) against rereading. Retrieval conferred no advantage on the final test. Crucially, cognitive load mediated the relationship between retrieval practice and final test performance: higher load was associated with poorer delayed recall, meaning that when working memory was already stretched by demanding content, the additional load of active retrieval cancelled out the benefit.

This is not a refutation of retrieval practice. As I have said before, the science of learning is a probabilistic enterprise, and this finding highlights a boundary condition, an important one. It adds to a growing list of boundary conditions accumulating around retrieval practice, joining the prior findings on feedback timing, material type and retention interval that have progressively narrowed the conditions under which the testing effect reliably appears. Retrieval practice works, but it is also rapidly becoming a lethal mutation, which is why careful implementation is key.

The mechanism is straightforward: retrieval practice works by forcing effortful recall, but if the content itself is already consuming all available working memory, there is no spare capacity for that effortful recall to operate. The implication for teachers is that retrieval practice needs to be calibrated to the complexity of the material rather than applied as a blanket strategy. For introductory content with manageable element interactivity, retrieval is powerful. For dense, interconnected material that students are encountering for the first time, other approaches may serve better.

A new paper by Sigayret, Parmentier and Silvestre (Frontiers in Psychology), suggests a further reason to be cautious, their paper failed to find any testing effect across two experiments run on Prolific. The authors are careful to say this is not a boundary condition on retrieval but a warning about the limits of crowdsourced platforms for cognitively demanding learning research, where disengagement and attrition swamp the effect.

Every school doing mandatory “Do Nows” as their primary retrieval strategy should ask: what exactly are students retrieving, and how complex is it?

Grimm and Richter published two preregistered experiments (N=304) in Learning and Instruction testing whether brief training in rational thinking would transfer to argument evaluation. Main finding is that the training worked: participants improved on the specific rational thinking tasks they practised, but those gains did not transfer to evaluating the quality of arguments, despite the two competencies being correlated.

This extends a long line of transfer failures which broadly shows that near transfer is a thing but far transfer is a much more difficult to capture. The correlation between two cognitive skills does not mean that improving one will improve the other. Transfer requires more than proximity; it requires explicit bridging, varied practice contexts, and often far more time than a brief intervention provides. This is another finding which challenges the generic “critical thinking” industry, which promises portable reasoning skills through standalone programmes.

The pattern across decades of transfer research is pretty clear: training a skill and seeing improvement on a test of that skill is not evidence that the skill has become generalisable. What transfers is knowledge, applied deliberately in new contexts, not abstracted competencies floating free of content.

Hong, Son and Kim’s study in Cognitive Research: Principles and Implications demonstrates something teachers innately know but rarely see quantified: brief exposure to information produced peak overconfidence that actually exceeded the baseline confidence of complete novices. Participants who had spent a short time with the material were more certain of their (often wrong) answers than people who had never seen it at all.

The danger here is not ignorance but the illusion of understanding that accompanies superficial exposure. For instruction, the implication is that partial coverage of a topic may be actively worse than no coverage, because it generates false confidence without the corrective feedback that deeper engagement provides. Students who have “done” a topic briefly are harder to teach than students who know they haven’t encountered it, because the former believe they already understand.

This connects to a broader pattern in the research: the conditions under which learning appears to have occurred (fluent performance, high confidence, rapid completion) are frequently the conditions under which durable learning has not.

Florean, Straga, Mantyla and Del Missier published a study in Cognitive Research: Principles and Implications demonstrating an interesting asymmetry in cognitive offloading: taking away a tool that students have been using damages performance far more than providing that tool in the first place improves it. Working memory capacity moderated the effect, with lower-WM individuals suffering the greatest disruption when the tool was removed.

This finding has immediate relevance to every classroom deploying AI tools, calculators, or any external cognitive support. The decision to introduce a tool is also, implicitly, a decision about what happens when it is withdrawn. If students build their understanding while relying on external support, they are constructing knowledge representations that depend on that support’s continued presence. Remove it for an exam, a different class, or a future year, and the architecture collapses asymmetrically: the loss is larger than the original gain.

As a general rule with tech, I would say that any tool introduced during learning should be introduced with its eventual removal in mind. Scaffolding should be designed to fade rather than to persist, and students should practise without the tool at regular intervals so that their mental representations do not become dependent on external support they will not always have.

The expertise reversal effect tells us that instructional methods effective for novices become ineffective or counterproductive for more advanced learners. Zambrano, Narváez-Rivera, Sayago-Heredia and Jácome-León tested this in a medical learning study published in Learning and Instruction, and confirmed the basic pattern: low-knowledge students learned better from direct instruction, high-knowledge students from minimally guided instruction.

But the critical finding was the moderator. The reversal only appeared when element interactivity was high. When the content was relatively simple, there was no crossover, regardless of prior knowledge.

This matters because the practical question teachers face is not just “when should I switch from explicit to exploratory?” but “when should I switch given this particular content?” This paper suggests the answer now has two variables rather than one: prior knowledge and content complexity. Simple material can be taught the same way to everyone without much cost. Complex, interconnected material is where the instructional method needs to match the learner’s existing schema. This is useful I think because it gives teachers a more actionable decision rule than “it depends on the student.”

Beege, Krieglstein, Wesenberg and Loibl published a study in the Journal of Educational Psychology that identifies when deliberately incorrect worked examples outperform correct ones. The answer: only with appropriate prompting, adequate scaffolding, and sufficient prior knowledge. Without all three conditions met, erroneous examples backfire.

When students have enough prior knowledge to recognise errors, when the prompting guides them to the right kind of comparison, and when the scaffolding prevents them from simply encoding the wrong procedure, the act of finding and correcting mistakes deepens understanding more than studying correct solutions.

The practical takeaway is conditional rather than categorical: erroneous examples are not a generic improvement over correct ones, but a powerful tool that requires three things to go right simultaneously. For teachers, this means using them later in a learning sequence rather than at the point of initial instruction, ensuring the errors are carefully chosen rather than arbitrary, and providing explicit prompts that direct attention to the critical features rather than leaving students to notice them independently.

Morin, Wichstrøm, Steinsbekk and colleagues published a birth cohort study (N=833) in Contemporary Educational Psychology with biennial assessments from ages 8 to 16. When within-person and between-person effects were properly disaggregated, the working memory–achievement link reflected stable individual differences rather than ongoing within-person causal effects.

This is not good news for the “train working memory to boost achievement” brain-training industry. The correlation between WM and achievement is real, robust, and has been replicated hundreds of times. But this study shows it is not telling the causal story most people assume. Children with higher WM capacity tend to achieve more, but changes in a given child’s WM over time do not reliably predict changes in that child’s achievement. In other words, the association seems to be trait-level, not state-level.

The practical implication is that interventions aimed at increasing WM capacity as a route to academic improvement are targeting a stable characteristic rather than a malleable cause. Resources spent on WM training programmes would almost certainly produce better returns if redirected toward reducing the WM demands of instruction itself.

A new study in Current Psychology using UK youth data from the BrainWaves Project showing that the already-small bivariate correlations between social media use and mental health outcomes (1–4% of variance) vanished entirely once theoretically relevant control variables were included. Social media accounted for effectively zero percent of the variance in depression, anxiety, social phobia, self-esteem, quality of life, mental wellness, and friendships.

This is not a new argument from Ferguson, he has been making essentially this case for over a decade, first about video games and now about social media. But the specific methodological point is important I think regardless of whether you find his broader position convincing. Most of the headline claims about social media harming youth rest on bivariate correlations: social media use correlates with depression, therefore social media causes depression. The problem is that social media use also correlates with dozens of other variables; sleep disruption, family conflict, pre-existing mental health difficulties, socioeconomic stress, and when you control for those, the unique contribution of social media shrinks toward zero.

The implication is not that social media is harmless. It is that correlational designs without adequate controls tell us very little about whether it is harmful, and that policy built on inflated bivariate associations is policy built on sand. For educators following the screen-time debate, this is a useful corrective to the more alarmist claims, though it is worth noting that cross-sectional data with statistical controls cannot establish causation in either direction, and Ferguson’s own framing is not neutral.

Seductive details can be described as interesting but irrelevant additions to instructional material, such as dramatic images, entertaining anecdotes or fun facts tangential to the core content. Colliot, de Pereyra and Boucheix published a study in Applied Cognitive Psychology showing that children with lower inhibitory control were disproportionately harmed by seductive details in instructional animations. The students most vulnerable to distraction were the ones least equipped to resist it.

This is important from an equity perspective. Seductive details are not just inefficient, they widen the gap between students who can filter out irrelevant information and students who cannot. The children with strong inhibitory control shrug off the distraction. The children without it are pulled away from the content that matters.

For teachers and curriculum designers, this means that the decision to include a fun-but-tangential video clip, a colourful-but-irrelevant image, or an amusing-but-distracting anecdote is not neutral — it is a decision that systematically disadvantages the students who are already struggling most with attention and self-regulation.

Powell, Rentzelas and Kambouri published a study in Learning and Instruction that should reassure every teacher who worries about giving honest feedback. Negative feedback per se did not harm performance. What harmed performance was giving mixed messages simultaneously: telling a student they are below average and that they have improved, in the same breath.

The mechanism is intuitive once you see it: mixed signals create interpretive uncertainty. The student has to resolve a contradiction rather than act on clear information. “You’re behind but you’re improving” sounds encouraging, but it forces the learner to decide which signal to attend to. Some focus on the “behind” and feel demoralised. Others focus on the “improving” and become complacent. Neither response is what the teacher intended.

The practical rule is simple: pick one frame and commit. If the message is “you need to improve,” deliver that clearly with specific guidance on how. If the message is “you’ve made progress,” deliver that clearly with specific evidence. Trying to soften hard truths by sandwiching them with encouragement does not produce the best of both worlds. It produces the worst of neither.

The single most important item this month is a retraction. Wang and Fan’s meta-analysis claiming ChatGPT boosted learning outcomes, published in Humanities and Social Sciences Communications (Nature), has been formally retracted. Commentators have pointed out that Wang and Fan lumped together studies with very different designs and small, heterogeneous samples, and coverage of the retraction notes that the authors did not respond to the journal’s correspondence. The formal retraction note itself is narrower, citing ‘discrepancies in the meta-analysis’ that undermine confidence in the validity of the analysis and its conclusions.

This matters because an large proportion of the “AI improves learning” claims circulating in education policy and edtech marketing either cite this paper directly or cite papers that cite it. When read alongside the new Liu, Xu and Xie meta-analysis in Frontiers in Psychology, which found a large positive effect of generative AI on intellectual outcomes (g = 1.096) but detected significant publication bias for exactly those outcomes, and the picture becomes clear: the GenAI-in-education evidence base is unreliable at the meta-analytic level. The positive results are there, but they are inflated by selective reporting and methodological weakness.

None of this means AI cannot support learning. It means the current evidence for the claim is far weaker than its proponents suggest.

A three-arm randomised controlled trial published in Computers and Education: Artificial Intelligence (N=275) tested AI assistance in a CS1 programming course. Students with AI support completed tasks more successfully and reported less stress. But they did not show deeper conceptual learning. Performance and learning dissociated cleanly.

This replicates the Bastani et al. Harvard tutoring finding in a different domain: AI makes the task easier without making the learner more capable. The students who used AI looked better on task metrics while learning less. This is exactly the kind of result that verification systems need to detect, because any measure based on task completion or course grades would show AI as beneficial. Only a measure of actual learning, tested independently of the AI assistance, reveals the gap.

The question is not whether AI helps students perform but whether it helps them learn, and those are increasingly looking like different questions.

Wong and Qiu published an RCT (N=196) in Educational Psychology Review comparing free ChatGPT use with a “think first” protocol that required independent ideation before AI access. Free ChatGPT use boosted initial creativity scores, but when the AI was removed, those gains collapsed. The think-first group produced more modest initial gains but retained their creative improvement independently.

This is one of the cleanest demonstrations of borrowed competence. The students who went straight to ChatGPT were performing with the tool rather than learning from it. Those who generated their own ideas first and then used AI to refine them internalised something durable. The mechanism maps onto the generation effect: producing an answer, even a wrong one, before receiving assistance creates a retrieval and elaboration event that strengthens learning.

The practical implication is not to ban AI but to sequence it: think first, generate first, struggle first, then bring in the tool. The cognitive work that precedes AI use is what makes AI use educationally productive rather than merely performatively impressive.

Oreopoulos and colleagues published an NBER working paper reporting results from a randomised controlled trial of Khan Academy in Indian schools. The gains were substantial. But the critical design feature was supervision: students used the platform during structured school time with teacher oversight, not independently at home.

This aligns with a growing body of evidence that technology-assisted learning works when it is embedded in a structured instructional environment rather than deployed as a substitute for teaching. The Khan Academy platform provided the content and adaptive sequencing; the teachers provided the structure, accountability, and human support that kept students engaged and on task. Remove either component and the effect would likely shrink.

For schools considering AI tutoring platforms, the lesson is that the technology is a complement to teaching rather than a replacement for it. The gains come from the combination, not from the software alone. This is expensive and labour-intensive, which is precisely why it works: there are no cheap shortcuts to effective instruction.

Telling isn’t teaching, even when it’s explicit. Sasson and colleagues claim that passive explicit instruction defined as clear explanations without guided practice, actually made students worse at question-formulating. The active version, where students did the cognitive work, produced gains. The label “explicit instruction” is covering two very different things.

Head-to-head on irregular words, and both methods assume phonics. Lithgow et al. published a direct comparison of mispronunciation correction versus the heart-word method in Scientific Studies of Reading. The key takeaway here is not which wins but that both approaches accept irregular words still have decodable parts, the debate has moved past “phonics vs whole language” to “which phonics?” technique.

ADHD plus reading difficulty is not the compounding disaster we assumed. Marks et al. found no evidence for an additive impact of co-occurring ADHD and reading disabilities. If this holds, schools may be over-differentiating their reading interventions based on diagnosis rather than the reading problem itself.

AI dependency runs through two routes, not one. Tian and Zhang’s dual-pathway model distinguishes an attentional route (engaged, evaluative use of AI; beneficial) from a reliance route (offloading cognitive work, harmful). Same tool, opposite effects, depending on the student’s stance.

Metacognitive instruction narrows the SES gap — with caveats. Maximino-Pinheiro et al. found that teacher PD focused on metacognition reduced the relationship between socioeconomic status and children’s metacognitive skills in early years. Hopeful, but 44 teachers is a small sample and this needs replication at scale.

Tyler Jagt’s essay in the Chronicle of Higher Education is worth a read, it’s sort of the university-level version of the literacy crisis that primary and secondary educators have been documenting for years. Jagt presents data on the collapse of student reading capacity at the undergraduate level. Paired with the growing evidence on AI-assisted cognitive offloading, this is a sobering portrait of where things stand.

Shalini Ramachandran's WSJ investigation is the most thorough piece of reporting on classroom screen use published this year. The headline number: a seventh-grader in Wichita accessed 13,000 YouTube videos during school hours in three months. A second-grader in New York watched 700+ in two months. YouTube now accounts for roughly half of all student traffic on school devices, and 60% of the K-12 mobile device market runs on Chromebooks,which is not an accident. Internal Google documents show the company explicitly targeted K-12 as an entry point for "lifelong brand loyalty" and identified children under 13 as the "world's fastest-growing internet audience."

Louise Eccles and Yennah Smart’s Sunday Times piece reports on a Centre for Social Justice survey of 1,000 GPs. The headline numbers: 75% agreed that clinical boundaries for ADHD and autism have expanded to include behaviours previously considered normal range. 66% said diagnoses are given out “too easily” where behavioural interventions would be more appropriate. 57% said financial entitlements (disability living allowance claims doubled from 420,000 to 900,000 children between 2016 and last year) strongly influence parental requests for assessment.

The OECD’s flagship education technology report is out. The headline finding comes from a field experiment in Türkiye: students given GPT-4 access improved short-term performance by 48%, but performed 17% worse once access was removed. The report distinguishes “fast AI”, namely cognitive outsourcing that boosts output while hollowing out learning, from pedagogically designed AI that keeps the student doing the thinking. 72% of teachers worry about academic integrity; only 37% of lower secondary teachers have used AI at all. This is sort of an institutional-level version of Bjork’s performance-versus-learning dissociation.

Deans for Impact published the second edition of their Science of Learning guide on 19 May 2026. The original was one of the most widely shared resources in evidence-informed education. The update incorporates research from the last decade, including newer work on retrieval practice boundary conditions, cognitive load measurement, and self-regulated learning. Free to download, and worth sharing with anyone entering the profession.

The Fordham Institute surveyed over 1,200 K-3 teachers and the results are a mixed report card for the science of reading movement. The good news: 82% have completed SoR-aligned training, Fountas & Pinnell usage has collapsed from 43% to 16%, and UFLI Foundations is now the most-used curriculum at 38%. The bad news: 30% of teachers still believe phonics and cueing are equally valid strategies, and 58% believe reading comprehension relies on transferable “skills” rather than background knowledge. Work to be done.

In related news, the Hechinger Report’s investigation into New York State’s $10 million phonics training programme is an implementation failure case study. The state gave the money to the teachers’ union to build SoR training for 20,000 teachers. The resulting course promotes balanced literacy, includes three-cueing, contains factual inaccuracies, and misrepresents researchers’ work. Meanwhile, 41% of New York fourth-graders score at the lowest NAEP level. Read alongside the Fordham survey above, one shows the movement’s progress, this one shows how easily that progress can be captured and diluted.

I was featured in this New York Times article about motivation.

Paul, Jim and I spoke about our book Instructional Illusions with the great Robert Pondiscio in a conversation with AEI.

I was interviewed for this article on AI in the TES.

I will be speaking at researchED in Houston on Sat 13th June.

Thank you for reading. If you’re new here, The Learning Dispatch is read by over 23,000 educators worldwide. If you find this useful, please consider sharing it with a colleague.

If you missed last month’s dispatch, you can find it here.

Until next month,
Carl

No posts

Read the original on carlhendrick.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.