RSS Amplifier

The Learning Dispatch · Apr 11, 2026

The Monthly Dispatch - What’s New in Learning Science - April 2026

0
Sign in to vote or save

Carl Hendrick · The Learning Dispatch

Does the sequence of retrieval and generative tasks matter?

If retrieval practice strengthens memory and generative activities like producing examples or explaining ideas support deeper understanding, then surely the order in which students do these things matters. A new laboratory experiment tested exactly this assumption, and found surprisingly little evidence for it.

University students studied short texts and then completed different sequences of learning tasks. Some generated examples before retrieving definitions; others retrieved definitions first. A third group restudied the material before completing generative tasks. Some students completed these activities immediately, others after a two-day delay. One week later, the sequence had made almost no difference. Students performed similarly regardless of whether they generated before retrieving or retrieved before generating.

What did matter was whether retrieval happened at all. Retrieval practice consistently produced better long-term retention than restudying, even when restudy was paired with generative activity. The implication is worth sitting with: teachers who agonise over the optimal choreography of learning activities may be solving the wrong problem. The critical variable is not the sequence but the presence of retrieval itself. Get students retrieving from memory, and the exact order of the surrounding activities appears to be a second-order concern.

Retrieval practice works but how much of this evidence actually comes from real classrooms, and does it work for all learners?

Tom Perry’s review of cognitive science in the classroom revealed an uncomfortable truth for many science of learning advocates: most research does not come from real classrooms. We know that retrieval practice is one of the most robust findings in cognitive psychology and over the past two decades, hundreds of studies have demonstrated what is often called the testing effect: attempting to retrieve information from memory strengthens long-term learning more than restudying the same material. A new perspective piece in npj Science of Learning analyses 23,850 publications on the testing effect and asks two important questions: how much of this evidence actually comes from real classrooms or from the laboratory? and does it generalise to all learners?

The answer is: mostly from the laboratory. Much of the testing effect literature comes from controlled experiments using short learning materials and university students. These studies provide strong evidence that retrieval practice works under tightly controlled conditions, but relatively few examine authentic classroom settings where instruction unfolds over weeks, involves multiple competing variables, and serves learners who are not psychology undergraduates. There is also limited research on diverse learner populations, including students with special educational needs or younger children, precisely the groups for whom the translation from lab to classroom is least straightforward.

None of this undermines the testing effect itself. The phenomenon is real and well-replicated. But classroom environments introduce what might be called a signal-to-noise problem: in a laboratory, the effect of retrieval practice is easy to detect because everything else is held constant. In a real classroom, that signal is diluted by differences in instruction, prior knowledge, motivation, and the sheer complexity of thirty minds working at different speeds on the same material. The review is a useful reminder that the distance between a robust laboratory finding and a reliable classroom practice is not a straight line, it is a translation, and translations always lose something.

Teaching children background knowledge before they read related texts does improve their comprehension, but the benefit is not equally durable for all learners.

This study set out to test a deceptively simple but practically consequential question: if you deliberately teach children the knowledge they need to understand a text, does that instruction actually improve their comprehension of it? Seventy-three Year 2 students (aged around eight) from a single regional school in Victoria were taught four topic-based knowledge units covering artefacts, fossil fuels, Faith Bandler, and Nelson Mandela, using an explicit instruction model over two four-week blocks. Students were then assessed on how much of the knowledge they had retained, and whether that retained knowledge enabled them to make inferences when reading related texts. A key additional manipulation was timing: some students read the related text immediately after instruction, while others read it after a four-week delay during which the knowledge was neither reviewed nor revisited. The study found that explicitly teaching a knowledge base did enable both skilled and less skilled readers to make more accurate inferences. However, less skilled readers acquired the knowledge base less thoroughly within the same instructional time, and they also showed significantly greater comprehension decline when there was a delay between learning the knowledge and reading the text.

The study supports the case for knowledge-rich curriculum design: deliberately teaching domain knowledge does appear to translate into improved comprehension of related texts, including for weaker readers. But the study also reveals a structural problem embedded within that approach. Less skilled readers, the very students a knowledge-rich curriculum is most intended to help, are less likely to fully acquire the knowledge base within a standard unit of instruction, and more likely to lose access to it if it is not regularly revisited. The metaphor that comes to mind is scaffolding that is removed too soon: the knowledge built through instruction may simply not be stable enough, in less skilled readers, to bear the weight of inference-making weeks later. This implies that effective implementation of a knowledge-rich curriculum is not merely a matter of content selection, but requires deliberate, recurring review built into the instructional sequence itself, what the authors call schema maintenance.

How students read multiple texts may depend less on general strategy skill and more on whether they know what strategic reading actually looks like.

A new study in Contemporary Educational Psychology examined the relationship between what students know about reading strategies and how they actually behave when reading multiple texts. The researchers developed a measure of strategic knowledge, not whether students can execute strategies, but whether they can recognise what good strategic behaviour looks like in the context of multi-text reading. They then tested whether this knowledge predicted actual strategy use and, in turn, comprehension.

The findings suggest that strategic knowledge is a meaningful but underexplored predictor. Students who could identify effective approaches to integrating information across texts were more likely to use those approaches themselves, and this translated into better comprehension outcomes. The implication is worth pausing on: much of the strategy instruction in schools focuses on practising strategies, but relatively little attention is given to building students’ declarative knowledge about when and why those strategies work. You cannot deploy what you do not understand. This echoes a broader pattern in the science of learning: metacognitive competence rests on a knowledge base, not merely on procedural habit.

Reading Comprehension Is Not a Skill

·

Mar 5

I taught English and reading comprehension for eighteen years. One thing I learned slowly, and against the grain of almost everything I was trained to do, is that when a student cannot grasp the main idea of a passage, the problem is almost never that they lack a “strategy.” The problem is that they do not understand enough of the words. Not a missing m…

Fostering creativity in schools: a systematic review of what actually works for younger learners.

A systematic review in Educational Research Review examined interventions designed to cultivate creative thinking in primary and early secondary classrooms. The review distinguished between explicit interventions — those that directly teach creative thinking skills or techniques — and implicit interventions that embed creative opportunities within regular instruction without labelling them as such.

The question the review addresses is one that education has struggled with for decades: can creativity be taught, and if so, how? The findings suggest that both explicit and implicit approaches can produce measurable gains, but that the evidence base is stronger for explicit instruction in creative thinking strategies. This is a finding that aligns with broader patterns in instructional research: making the target of learning visible and explicit tends to produce more reliable outcomes than hoping it will emerge as a byproduct of open-ended activity. For those tempted to dismiss creativity instruction as inherently incompatible with structured teaching, the review offers a useful corrective. The most effective creativity interventions in the review were not unstructured. They were highly structured, but structured around the right things.

Do we actually understand what makes feedback effective, or are we still describing what feedback looks like?

Two companion papers in the March 2026 issue of Contemporary Educational Psychology take stock of the feedback literature and arrive at a shared conclusion: the field knows more about the forms feedback takes than about the processes by which it works.

Daumiller and Meyer argue that feedback research has been disproportionately focused on surface-level features — timing, specificity, source — while paying insufficient attention to the cognitive and motivational processes that determine whether feedback is actually used. Their call is for a shift in focus: from what teachers do when they give feedback to what happens inside the learner when they receive it. Koenka, McManus, and Nicolai make a related but distinct point. They argue that feedback’s conceptual foundations remain underdeveloped, and that the field needs a more coherent theoretical framework to explain why the same feedback intervention can produce dramatically different effects depending on the learner, the task, and the context.

For practitioners, the message is both validating and cautionary. Teachers have long suspected that giving feedback is not the same as students learning from it. These papers confirm that suspicion and locate the gap precisely: the problem is not a shortage of feedback, but a shortage of understanding about what happens between the moment feedback is delivered and the moment it either changes thinking or is quietly ignored. If feedback research is going to be useful to teachers, it needs to move beyond describing inputs and start explaining mechanisms.

Teachers keep learning throughout their careers and 15 years in, their pedagogical knowledge is still growing.

A longitudinal study published in Contemporary Educational Psychology tracked teachers’ pedagogical and psychological knowledge over 15 years in the profession. The finding challenges a common assumption embedded in much of the professional development literature: that teacher knowledge plateaus early and that experienced teachers are merely refining habits rather than acquiring new understanding.

The study found measurable growth in pedagogical knowledge sustained well into mid-career. This is not the same as saying all teachers improve, or that experience alone is sufficient. But it does suggest that the profession’s knowledge base is not a fixed endowment determined by initial training. For school leaders, the implication is that experienced teachers may be significantly undervalued as sources of pedagogical insight — and that professional development models which treat all teachers beyond five years as essentially equivalent may be missing real differences in accumulated expertise.

Which teacher characteristics actually predict instructional quality? A meta-analysis finds the answer is more nuanced than “years of experience.”

A meta-analysis in Learning and Individual Differences synthesised the evidence on which teacher attributes most strongly predict effective instruction and student learning. The study examined a range of characteristics — content knowledge, pedagogical knowledge, beliefs, self-efficacy, personality — and assessed their relationships to independently measured instructional quality.

The findings resist easy summarisation, which is itself informative. No single characteristic dominates. Content knowledge matters, but its relationship to instructional quality is weaker than many assume. Pedagogical knowledge shows stronger and more consistent associations. Teacher beliefs and self-efficacy show modest but reliable effects. The pattern that emerges is of a complex, multi-dimensional skill that cannot be reduced to any single variable. For those designing teacher selection or development programmes, this is both a warning against silver-bullet thinking and an argument for broad-based approaches that develop multiple dimensions of teacher expertise simultaneously. What makes a teacher effective is not one thing. It is many things working in concert — and isolating any single factor risks mistaking a thread for the tapestry.

An AI tutoring system that personalises the order and difficulty of practice problems produced meaningfully better exam outcomes than a standard fixed sequence across 770 high school students over five months.

Most AI tutoring platforms are reactive: they wait for students to ask questions and then respond. The research team argues this paradigm misses something fundamental about how learning works. Drawing on mastery learning, they designed a system that actively guides what problems students attempt next. The core innovation is a reinforcement learning algorithm that estimates each student’s knowledge state using not just answer correctness but the quality of student-chatbot conversations and the nature of code editing behaviour. This richer signal feeds into an algorithm that dynamically selects the difficulty level of the next practice problem. Across ten Taipei high schools over five months, 770 students were randomly assigned to either this adaptive sequencing condition or a fixed easy-to-hard sequence. The adaptive group outperformed the control group by 0.15 standard deviations on an in-person, unassisted final examination, a pre-registered finding that held across multiple regression specifications.

For educators, the implications are encouraging. The study does not simply show that AI tutoring works; it shows that a specific mechanism, personalised problem sequencing, adds meaningful value over and above a generic AI chatbot that all students had equal access to. The control group was not deprived of AI support; they had the same chatbot and the same content, just in a fixed sequence. This is an important comparison condition, and it matters. Crucially, the gains were not achieved by making students do more work or spend longer on the platform in absolute terms. The mediation analysis suggests the benefit was carried through increased engagements, students in the adaptive condition spent more time on task and made more attempts per problem, and they interacted with the chatbot in qualitatively more productive ways. For school leaders considering AI tutoring deployments, the practical implication is that what you sequence matters as much as what you provide, and that keeping students productively challenged, not overwhelmed, not bored, appears to be the active ingredient.

Does AI-powered note-taking help students learn from video lectures, or does it quietly do the cognitive work that would have built understanding?

A study in Contemporary Educational Psychology examined what happens when students use AI-powered note-taking tools while watching STEM video lectures. The researchers assigned students to different conditions — some took notes with AI assistance, others without — and measured both learning outcomes and students’ perceptions of how helpful the tool was.

The study adds to a growing body of evidence that suggests the relationship between AI assistance and learning is not straightforward. The concern is one that will be familiar to anyone who has followed the cognitive offloading literature: when a tool reduces the effort required to process information, it may simultaneously reduce the depth of encoding that makes learning durable. The note-taking context is particularly telling because note-taking has always been understood as a generative activity — its value lies not in the notes produced but in the cognitive processing required to produce them. If AI handles the processing, the student receives a better product but may undergo a weaker learning experience. The question for instructional designers is whether AI note-taking can be configured to augment rather than replace the generative work — or whether the efficiency gains and the learning losses are inseparable features of the same mechanism.

When people know content was produced by AI, they trust it less — even when it is competent. What does this mean for AI-assisted learning?

A study published in Computers in Human Behavior: Artificial Humans examines what the authors call the “AI Penalty”; the systematic reduction in trust, perceived authenticity, and knowledge uptake that occurs when people are told content was generated by artificial intelligence. The paradox is that transparency about AI involvement, which we might expect to build trust, actually undermines it. Identical content is rated as less credible, less authentic, and less persuasive when attributed to AI rather than a human author.

For education, this finding cuts in two directions. On one hand, it suggests that student scepticism toward AI-generated feedback or explanations may be a real barrier to learning, even when the content is accurate. On the other, it raises an uncomfortable question about what happens when AI assistance is not disclosed. If students learn equally well from AI-generated content when they believe it was written by a human, but less well when they know its provenance, the disclosure paradox creates a genuine dilemma for educators committed to both transparency and effectiveness. The study does not resolve this tension, but it names it precisely — and that is a useful contribution.

From “AI Psychosis” to “Brain Rot” — how pseudo-diagnoses may be crowding out genuine understanding of technology’s effects on the mind.

A review in Computers in Human Behavior argues that the proliferation of informal diagnostic labels for technology-related mental states; “AI Psychosis,” “Brain Rot,” “doomscrolling disorder”, is actively impeding legitimate psychological research. The authors contend that these labels, amplified by social media and news coverage, create an illusion of scientific understanding where none exists. A catchy name is not a diagnosis. A trending term is not a construct.

The relevance for education is direct. The history of educational psychology is littered with pseudo-scientific concepts that were adopted because they were intuitively appealing rather than empirically grounded — learning styles, digital natives, the 10,000-hour rule. Each acquired diagnostic authority long before the evidence warranted it, and each proved remarkably difficult to dislodge once embedded in professional practice. The authors’ concern is that the same pattern is repeating with AI-related psychological claims: we are naming phenomena before we understand them, and the names themselves are shaping how we think about the underlying reality. For educators navigating the AI conversation, the paper is a useful reminder that the impulse to pathologise is not the same as the work of understanding.

Can large language models actually teach creativity? New research suggests cognitive gains but affective gaps and the distinction matters.

A study in Computers in Human Behavior: Artificial Humans investigated whether LLMs can enhance creative thinking. The researchers measured both cognitive outcomes, divergent thinking, idea fluency, originality, and affective responses, while also tracking neural patterns during the creative process. The headline finding is that LLM interaction did produce measurable improvements on standard creativity metrics. Students generated more ideas, and those ideas were rated as more original by independent evaluators.

But the affective data tells a different story. The subjective experience of creating with AI assistance felt qualitatively different from unassisted creation. The emotional engagement, the sense of ownership, the felt difficulty of the work, all were diminished. This is a finding that resonates well beyond the creativity literature. It echoes the pattern visible in the cognitive offloading research: AI can improve measurable outputs while simultaneously degrading the quality of the experience that produces them. Whether this matters depends on what you think education is for. If the goal is to maximise the quality of the product, AI assistance is clearly beneficial. If the goal is to develop the capacity of the person, the answer is considerably less clear — because the struggle that AI removes may be precisely the struggle through which creative capacity is built.

Most adaptive learning systems adapt to generic signals like accuracy and speed. A new editorial argues this misses the point: what matters is the structure of the domain being taught.

An editorial introducing a special issue of Learning and Instruction makes a deceptively simple but consequential argument: adaptive learning systems cannot be truly adaptive if they are domain-blind. Strohmaier, Depaepe, Nickl, and Obersteiner argue that most current approaches to adaptivity — whether in AI tutoring platforms or human-led differentiation — rely on generic learner metrics: accuracy, response time, time on task. These signals tell you that a student is struggling. They do not tell you why, and the why is almost always domain-specific.

The authors identify three dimensions where this matters. First, what to measure: the relevant learner variables differ by subject. A meaningful error in fraction arithmetic — confusing the denominator with a whole number, say — is categorically different from a meaningful error in source evaluation during a history lesson. A system that treats both as “incorrect” has already discarded the most useful diagnostic information. Second, how to measure it: the same observable behaviour carries different meanings in different domains. A long pause before responding might signal productive mathematical reasoning in one context and a complete vocabulary gap in another. Third, how to respond: the instructional move that follows from a diagnosis should be shaped by the conceptual structure of the domain, not merely by a generic adjustment to difficulty level.

The special issue includes nine empirical studies across different subjects that collectively illustrate both the promise and the difficulty of building domain-specific adaptivity. For educators evaluating AI tutoring tools, the editorial offers a useful diagnostic question: is this system adapting to what my students are thinking, or merely to whether they are getting things right? The difference is the difference between a thermostat and a diagnostician. One adjusts to a signal. The other understands a system.

AI Brain Fry, Workslop and the Ironies of Automation

·

Mar 14

In 1983, a cognitive psychologist named Lisanne Bainbridge published a four-page paper in an engineering journal that almost nobody outside her field has ever read. It concerned the automation of industrial processes: nuclear plants, chemical refineries, flight decks. Its tone was measured, almost dry but it would prove to be unerringly prophetic for th…

Not all motivation fluctuates equally: achievement goals shape not just how motivated students are, but how much their motivation swings from day to day.

A study in Learning and Individual Differences moves beyond the question of average motivation levels to examine something more granular: daily variability. Using intensive longitudinal methods, the researchers tracked students’ intrinsic motivation across multiple days and examined how different achievement goal orientations — mastery goals, performance goals, and their various subtypes — predicted both the level and the stability of that motivation.

The finding that matters most for practitioners is that some goal orientations produce motivation that is not only higher on average but more stable over time, while others produce motivation that is volatile — high one day, absent the next. For teachers, this is a useful reframe. It suggests that the question is not simply “how do I motivate my students?” but “how do I produce motivation that persists?” A student who is intensely motivated on Monday and entirely disengaged on Wednesday may average out to “moderately motivated” on any standard measure, but the instructional challenge they present is fundamentally different from a student who is steadily, modestly engaged all week. The study suggests that the type of goal a student pursues may matter as much for motivational stability as for motivational intensity.

Adolescent self-regulated learning is not a single dimension and the profiles that emerge predict achievement in ways that simple averages miss.

A study in Learning and Individual Differences used latent profile analysis to identify distinct patterns of self-regulated learning among adolescents. Rather than treating self-regulation as a continuum from low to high, the researchers looked for qualitatively different profiles: students who might be strong on planning but weak on monitoring, or effective at metacognition but poor at time management.

The profiles that emerged were meaningfully different from one another, and they predicted academic achievement more accurately than any single self-regulation measure alone. They also showed associations with sociodemographic factors, suggesting that self-regulation profiles are not purely individual traits but are shaped by the contexts in which students learn. For teachers, the practical implication is that “teach students to self-regulate” is too blunt an instruction. Different students may need support with different components of the self-regulatory process, and a student who appears to lack self-regulation may in fact have a specific deficit in one area while being competent in others. The profile approach suggests that diagnosis should precede intervention — and that the diagnosis needs to be more fine-grained than a single score on a self-regulation questionnaire.

Does metacognition actually drive improvement in maths over time, or does it just track who is already doing well?

This longitudinal study followed 334 junior high school students in China over two and a half years, measuring both their metacognitive development and their mathematics achievement every six months. The results show a surprisingly uneven pattern. Metacognition rose in the first year, then declined across the following two years. More importantly, students clustered into three distinct developmental groups: those with consistently high and slowly rising metacognition, those with moderate but declining levels, and those with low and sharply declining levels. These trajectories were mirrored, in complex ways, in their maths attainment.

The key finding is subtle but important. Students with higher metacognition tended to have higher maths achievement at the outset, but changes in metacognition did not strongly predict changes in achievement over time. In other words, metacognition seems to be more of a marker of who is already doing well rather than a clear driver of improvement. For teachers, this challenges the common assumption that simply “teaching metacognition” will automatically raise attainment. It suggests that metacognitive skill sits alongside knowledge and achievement rather than acting as a straightforward lever to accelerate it.

No posts

Read the original on carlhendrick.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.