We lost something when assessment became measurement outside of the learning process.
The word assessment comes from the Latin assidere, meaning “to sit beside,” or “sit with in counsel or office.” The original sense was someone sitting alongside you, counseling you, helping you work and see where to go. Somewhere along the way, that meaning collapsed into something narrower. Assessment became the thing that happens at the end for a grade. It became measurement.
This article is about what we lost when that happened.
In 2008, Jeffrey Karpicke and Henry Roediger published a paper in Science titled: “The Critical Importance of Retrieval for Learning.” College students learned forty Swahili-English vocabulary pairs under different conditions. The students who repeatedly tried to recall the pairs from memory remembered about 80% of them a week later. The students who simply restudied the pairs after getting them right once remembered about 36%. The two distributions had little-to-no overlap.
This finding has been replicated, refined, and extended for two decades. Adesope, Trevisan, and Sundararajan’s 2017 meta-analysis of 188 experiments found an average effect size of g = 0.51 for practice retrieval compared to restudy, and g = 0.93 compared to no practice at all. Those are much larger effects than we see in most educational interventions we currently spend big money on in schools worldwide.
The classroom replications are the most impressive part of this research - it proves it’s true in practice, not just laboratory settings. In one sixth-grade social studies classroom in Illinois, students scored 79% on end-of-semester items that had been quizzed during the unit, versus 67% on items that hadn’t. In an eighth-grade science classroom, the effect ranged from 13 to 25 percentage points on unit exams. Pooja Agarwal’s 2021 systematic review of forty-nine classroom studies found that 57% showed medium-to-large benefits, across grade levels and across subjects.
This is one of the more robust effects in the entire learning literature and we just don’t talk about what it means or what to do with this information.
The mechanism matters because it changes what we should make of the finding.
When a student tries to pull something out of memory, the memory itself changes. The retrieval cues become more diagnostic and the information becomes easier to find next time. The act of reaching is the learning event, not the act of being told the right answer.
This is also why the pretesting effect works. Richland, Kornell, and Kao found that having students attempt to answer questions before they’ve been taught the material improves retention later, even when the pretest answers are mostly wrong. Wrong answers generated under genuine effort enhance learning of the correct material. The reaching produces the cognitive conditions under which the eventual correct information becomes meaningful.
Last week I argued that grading punishes the errors learning requires. This week’s argument is the same one rotated ninety degrees: accurate retrieval requires errors. The act of trying, failing, and trying again is what makes the memory durable, it’s what makes learning stick.
If retrieval is so effective, why don’t students do it on their own?
Karpicke, Butler, and Roediger surveyed 177 college students about their study habits. About 84% reported rereading their notes or textbooks. Only 11% mentioned self-testing as a strategy. When given the explicit choice between rereading and self-testing for an upcoming exam, most chose rereading.
The reason is more of a metacognitive misjudgment than laziness or lack of information. When a student rereads, the material starts to look more familiar, which produces a strong feeling of fluency, and so the student concludes s/he knows it. The feeling is wrong, but it’s persistent, and it leads people to keep using the strategy in pursuit of that feeling rather than the strategy that produces the learning.
Last week I wrote about the same phenomenon on the teaching side: Teachers and administrators confuse smooth lessons with effective lessons because smoothness produces the feeling of learning happening. Now, the student version of this problem is that students confuse familiarity with knowing. Both errors share the same root: the feeling of learning and the fact of learning are not the same thing, and we have built schools that systematically reward the feeling.
It’s worth directly addressing the question: if retrieval practice is so good, why don’t we just give more tests?
The answer is: The retrieval practice that produces these effects is low-stakes, ungraded, frequent, and embedded in the flow of teaching. It is not the high-stakes graded tests most think of when they hear “more assessment.” We could call it ‘formative assessment’ if that hadn’t been co-opted for measurement also.
The data on anxiety back this up. Agarwal and colleagues surveyed 1,408 middle and high school students whose teachers used low-stakes retrieval practice in class. 72% said it made them less nervous about exams and 92% said it helped them learn. The mechanism is intuitive: Anxiety thrives on novelty and unpredictability, while a classroom where students regularly practice the experience of being uncertain in a low-stakes setting, with an adult on their side, gives them the confidence, the calm, and the repetition to reduce test anxiety.
The retrieval practice that helps is the kind that doesn’t go in the gradebook.
Let’s talk limitations because retrieval practice isn’t a magic bullet. The research is robust, and it clearly works, but, as always, there’s nuance.
Retrieval practice strengthens what students are learning, but it is not a substitute for high-quality first instruction. It works best on material students already have some baseline understanding of, which is why brain dumps don’t work as well before instruction as they do after a partial introduction to the material.
The effects of retrieval practice on transfer to different applications are real, but modest. Retrieval practice helps students remember what they retrieved and produces some transfer to related material, but it doesn’t quite generalize to novel contexts. That question, what makes knowledge usable in places we didn’t practice it, is for next week.
This research and this instructional strategy should make more administrators interested, but we don’t need any evangelists.
The deeper point is the one sitting in the title. Assessment was supposed to be a sitting-beside, a way of helping students see where they were and where the work was leading. But, we built systems for the measurement of learning and forgot to build systems for the assessment of it.
The good news is that we have a clear, evidence-based version of the assessment-as-learning strategy. The brain dump at the start of class is assessment. The pretest before a unit is assessment. The two-minute exit ticket asking what students remember from today is assessment. None of these need to go in the gradebook to do their work.
For teachers, the implication is immediate. Open class by introducing the lesson and then immediately lead a brain dump. Pretest a unit with hard questions before teaching. Use exit tickets framed as retrieval rather than checks for understanding. None of this requires permission, a new curriculum, or additional time.
For administrators, the work is different. Distinguish between assessment for learning (frequent, low-stakes, ungraded) and assessment of learning (periodic, graded, summative) in your assessment policies. Make sure the retrieval practice tools you adopt don’t automatically push scores into your student information system. Teach the metacognitive trap explicitly in professional development, because teachers who haven’t grappled with the gap between the feeling of learning and the fact of learning won’t understand why to implement these practices.
We have built systems for measurement. We can rebuild systems for assessment too.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.