Hello, how are you, and thank you for subscribing to The Coaching Letter. You rock. It is surprisingly cool here in New England today; at my house right now it’s trying to decide whether to rain or not—this is a state my mother would have called “swithering”. I know that many people reading this are preparing to start school again, or may even have already started—I was talking to a high school teacher in Rhode Island last week who was reflecting that every year he wonders how he’s going to manage, but every year he does, and I know that feeling really well. And, truth be told, I don’t miss it. The start of school was always a time of high anxiety for me, so if you’re in that boat right now, I feel for you, but don’t worry, you’ll figure it out.
This Coaching Letter is a bit geeky, so if you’re busy trying to get school open, ignore it, you won’t miss it1. But if you are in a situation where “the research” gets cited as if there is a clear-cut case to be made for or against a specific practice (BTC comes to mind2), then it might be useful. It’s about what Bill McCallum calls “The Science of Reading References”. Here’s the ChatGPT-generated TL;DR version:
Don’t assume that a confident claim followed by a citation is the same thing as a claim supported by evidence. Research is routinely simplified, selectively quoted, stretched beyond its boundary conditions, or treated as settling questions it never actually examined. Current debates about explicit instruction, cognitive load, retrieval practice, and productive failure are far more nuanced than their advocates and opponents sometimes admit. That matters for practices such as Building Thinking Classrooms that do not fit into tidy instructional boxes. Research on productive failure provides a credible explanation for why giving students carefully designed opportunities to grapple with problems before explicit instruction can improve what they notice, understand, and subsequently learn. The point is not that BTC—or any other approach—is always right. It is that we should read the research closely enough to ask what it actually establishes, for whom, under what conditions, and compared with what.
And this issue has been on my mind a lot this summer as I have been reading books and articles and blogs while on ferries and trains (not to mention the entire day I spent at Schipol). It seems to me there is a fair amount of swithering going on in the blogosphere, where pundits of various types are pointing out that perhaps what has been promulgated about a particular practice—either in favor or against—is perhaps overly simplistic. As Bill McCallum puts it in one of his posts (and of course, drawing from The Princess Bride) I don’t think it means what you think it means. Or at least, it doesn’t mean everything you think it does.
This is a lesson I learned decades ago when I was a parent representative from my kids’ school to a district advisory committee—one of those advise and consent things that don’t really have any power—maybe those have been abandoned by now. Anyway, the district was going all in on a massively expensive consulting arrangement with an expert on co-teaching to implement a massively expensive staffing arrangement—teachers are expensive (as they should be—I would pay them more) and putting two teachers in a room is a massive investment of resources, so if you’re going to do that you should be pretty very extremely sure that those resources wouldn’t be more effectively leveraged some other way. Given the stakes, this is a huge enormous consequential deal. And I didn’t really buy it. My relevant experience (with special education, with Title I co-teaching, with budgets, with scheduling, and so on) told me that we were being sold a story. So I bought the consultant’s book, and I started looking up the research he was citing, and in the first three pages of the introduction, NONE of the references that he cited actually said what he said they did. What did I do? I wrote emails to complain and the superintendent was snarky to me and then he was fired over a sexual harassment case and we moved to Connecticut (these are not necessarily connected events), so I don’t know how the story ended… I’m not arguing against co-teaching, per se. But I did learn a lesson. Don’t take the citations at face value. (If you want a ChatGPT-generated summary of what the current state of the research on co-teaching says, here you go.)
Bill McCallum’s framing—I don’t think it means what you think it means3—is on the kind end of the spectrum, and suggests that people who fail to read (or who misrepresent) what the writers they are citing are citing are mostly harmless. If that is true, what is going on? First, there is motivated reasoning. This is a kind of cognitive bias when someone, frequently unconsciously, is invested in protecting or promoting a viewpoint they already hold. We all do this, obviously—we all interpret data in light of our current knowledge, beliefs and mental models. You read a paper that seems to agree with your take on something, and you think, “See! Told you so!” and you are not necessarily motivated to read for nuance or complications4.
But sometimes this tips over into something a bit more disingenuous. Bill McCallum uses the example of one of Greg Ashman’s blogs, in which Ashman argues that conceptual understanding is not a real thing. This might come as a surprise to a lot of people, so bears checking out. And indeed, one of the sources Ashman quotes as suggesting that there is no difference between conceptual and procedural understanding actually says no such thing when you read the couple of sentences immediately preceding what Ashman quotes. That feels a little dodgy.
It pays, then, to read the citations, because you might be surprised at what they actually say. And, which is where it gets really tricky, you’ve got to spend enough time with all this to be able to notice what people don’t cite, which is why I’m always feeling like I’m running just to stand still—and why I spent so much time this summer reading this stuff. I think there’s enough active debate about several issues going on right now that some prominent folks are sounding notes of caution that we should pay attention to. So here is a not-random, but not in any particular order, list of things I’ve read recently related to this issue of being careful about making claims—or inferring conclusions—about the research that don’t hold up when you pay attention to what they are actually standing on—or choosing to ignore.
Explicit Instruction Works. Education Schools Just Won’t Admit It by Zach Groshell. Here’s a great example of a piece that’s not wrong, exactly, it’s just not complete. It’s true that there is a large corpus of research validating the effectiveness of explicit instruction, no argument there. But in rejecting other approaches, Groshell overlooks research showing that novice status is not a fixed characteristic of children: it reflects how much relevant knowledge a learner has for a particular domain or class of problems, and therefore changes as the learner acquires knowledge (Kalyuga, 2007; Kalyuga et al., 2003). The instructional support that is helpful early in learning may consequently become unnecessary—or even counterproductive—later. This is known as the expertise-reversal effect: instructional techniques that benefit learners with little relevant knowledge can lose their effectiveness, or impose redundant cognitive load, as learners develop expertise (Kalyuga et al., 2003). Research on fading worked examples similarly supports a transition from example study toward increasingly independent problem-solving rather than a permanent choice between the two (Renkl et al., 2002). Children need explicit teaching and opportunities to solve problems; the interesting question is how those experiences should be designed and sequenced. Beginning every lesson with explicit instruction is not, therefore, an irrefutably correct instructional choice5.
Will you forgive me if I get a little more snippy than is even typical for me? I find it laughable that Groshell complains that explicit instruction has been caricatured in the literature, when he does a bang-up job of caricaturing all forms of instruction that are not explicit instruction. And I think many people who cite Project Follow Through have not read it—neither have I, but I know enough to know that it’s very likely that you’ve never even heard of the other approaches that Direct Instruction beat out, so I’m not sure where the victory is. PFT was conducted more than 50 years ago. Groshell presents it as though it were a valid head-to-head contest that proved, once and for all, that Direct Instruction is the best instructional approach ever. This is a little like saying that a study conducted in 2006 proved once and for all that the BlackBerry was the best mobile phone and anything that came after is sure to be a dud. The problem is not old research—as many of you know, I like to point out that much valid instructional research has very old roots. The problem is that this particular research is historically situated and cannot establish the superiority of Direct Instruction over approaches that had not yet been developed6.
Retconning the Curriculum: Why The Science of Learning Has a Serious Design Problem by Carl Hendrick. This is a great example of a researcher who is, unlike some bloggers, very keen to make sure that people pay attention to the limits of the research (e.g., “Over the last 10-15 years, retrieval practice (and its symbiotic twin, spacing) have become an article of faith but increasingly I feel they are being used in ways which are not supported by the evidence.”) I always appreciate when someone makes me think differently about something that I had assumed or taken for granted, and here it was the idea that retrieval practice (being asked to recall what you have, in theory, already learned, but not the last thing you learned) should be built into the curriculum. This is framed as part of a broader statement that a lot of what we think of as instruction should actually be built into the curriculum—this is a point that we make over and over again. Task design, for example, gets treated as though it is a normal part of a teacher’s lesson planning responsibility, but task design is difficult and time consuming, and teachers are not trained to do it—plus the idea that all teachers should be coming up with their own tasks is just depressingly inefficient. I loved the example of the textbook study that showed that only 10% of the problems in the books required that students “retrieve” anything from previous chapters—not least because this just reinforces the point that teachers cannot possibly be expected to, individually, track the problems that they assign over the course of the year to ensure that students are given enough retrieval practice. This should be a curriculum function.
That Carl Hendrick post got us to this report on Cognitive Science in the Classroom, and I had ChatGPT give us a summary, which is super useful.
How Much Complexity Can One Theory Take? by Michael Pershan. This one made me feel like an idiot. I talk about Cognitive Load Theory all the time. It makes sense to me, because the research and the theory hang together nicely, and also because they comport with my own experience—it makes sense to me that we shouldn’t play music while people are being asked to think because it is a distraction, and I know that I don’t work well when my surroundings are cluttered. But I did not know that the history of the concept was so complicated, and this is where I learned about the expertise reversal effect, and that productive failure is not what I thought it was…
Productive Failure by Manu Kapur. This is a book, not an article. And I want to preface this part by saying that Tom and I have been talking about productive struggle for weeks now, and have a plan to write at least one joint post on the topic. And maybe the concept of productive failure deserves its own CL… Anyway, it’s another idea I didn’t completely understand. I had been influenced by Greg Ashman7, who writes about productive struggle and productive failure and conflates the two, uses them interchangeably, and is universally disparaging. But the book, and the underlying research, are really useful in conveying why it might be really useful to put students8 in the position of not being able to solve a problem at first blush, so to speak. In the book, Kapur argues that unsuccessful problem solving can produce forms of preparation that traditional Cognitive Load Theory (CLT) largely ignores. For example, students may:
notice critical features of the problem,
discover why naïve strategies fail,
differentiate important variables,
activate relevant prior knowledge,
become aware of gaps in their understanding,
develop richer questions
be motivated to know, and understand, the right answer.
And all of these might be useful in making subsequent explicit instruction more useful to them. In other words, and my experience totally backs this up, and despite what the CLTists insist cannot possibly be helpful, proponents of productive failure in some instructional circumstances are not trying to argue against explicit instruction—they are trying to make explicit instruction more productive. Imagine!9 And for those of you interested in BTC, I think what the author is describing in terms of instructional practice comes closest to describing BTC. I realize there’s some danger here, that someone might infer that BTC does not care about the right answer. That’s not what I’m saying. I’m simply pointing out that a set of practices that are often questioned as not being in line with “the science of learning” (not totally clear what that is…) or cognitive science or CLT does appear to closely resemble valid, peer-reviewed research showing that it is frequently beneficial to students to give them the opportunity to grapple with problems and ideas that they are not immediately going to be able to answer easily or smoothly10.
Maximizing Opportunity to Learn—our book! It has a chapter on Potentially Problematic Practices that tries to dig into the nuances around several practices that tend to be treated as though they are always beneficial but actually are nuanced and don’t always accomplish what people think they do. These are:
Teacher questioning and wait time
Differentiation and scaffolding
Grouping
Learning intentions and success criteria
Feedback
Again, we are not arguing that you should not employ these practices, merely that the devil is in the details and it’s worth knowing more about them.
I fear I am running out of steam—this has taken me a lot longer than I thought, it’s now dark and the thunder has passed. So before I do, three plugs for other Substacks you should follow. First, the careful reader of research whose blog post inspired this one, Bill McCallum. He definitely has invested a lot of effort understanding many strands of research, and deserves to have a bigger following. My colleague Tom, who writes short, tight summaries of research that scratch below the surface. Here are a few you should check out if you haven’t already: #18 on small groups, #10 on instructional time, #15 on student discourse, and #5 on group v individual work. Finally, Marc Smith writes a curiosity-filled blog about the research on learning, Dynamic Learning.
And just to round out the geeky theme, Rydell, Andrew, Tom and I have a chapter in the newly published book Teaching and Learning for Collaborative Continuous Improvement in Education. Our chapter is called “The Core Practice Test Kitchen: A Model for Collaborative Professional Learning through Small, Facilitated Instructional Improvement Teams”. It’s the best description we have, so far, of the work that we do with our various networks and partner districts to improve instruction by asking teams of teachers to systematically test core instructional practices (“recipes”). I haven’t read the other chapters, because I don’t even have a copy of the book yet, and I know that you’re unlikely to want to buy a book for one chapter (although the code MEP25 will get you 25% off!), but if you do read it, I would love to know what you think.
I feel like this was overly ambitious and I didn’t really pull it off, so I will circle back. But it’s important to me to at least get these ideas out there, in the hope that we will all do a better job of thinking twice before making categorical statements about what is wrong or what is best. In the meantime, if I can do anything for you, please let me know. Best, Isobel
I finally figured out that Substack will let you use footnotes! This is just your reminder that all posts are available at isobelstevenson.substack.com which is searchable, although not always easily so. And if you have trouble finding what you need, you can always email me.
If you are mostly here because of your interest in BTC, then you should DEFINITELY read Productive Failure, by Manu Kapur. I write about it in the body of this Coaching Letter, and I plan to write about it some more, and it really helped me to understand what’s really going on, cognitively, in the Task portion of the lesson.
For those of you who do not have Princess Bride memorized, the line “I do not think it means what you think it means” is used by Inigo Montoya to Vizzini, because the latter keeps saying “Inconceivable!” when what he means is that he did not think it could happen—the two are different, but Vizzini conceives of himself as intellectually superior, and so if he did not think it could happen then it must be, ipso facto, inconceivable. On the one hand, this could be construed as a series of simple mistakes by Vizzini. On the other, you could see him as engaging in an ongoing defensive reaction to a challenge to his self-perceived infallibility. He certainly doesn’t learn from the consequences of his over-confidence. So actually, I don’t know whether Bill McCallum is using it more kindly (a benign failure to do due diligence) or to call out disingenuousness or motivated reasoning.
Interestingly, I have found one of the most useful uses for ChatGPT is to check me on this. I will give it a research article, what I have written about, or citing, this research article, and ask it to tell me if what I have written is justifiable. And, unsurprisingly, it always comes back with the same response, to wit: Yes, you are broadly correct, but you overstate your claim thusly… What would be a more accurate reflection of what the research actually supports is… And I am duly humbled and I take the feedback and make corrections. So yes, I have asked ChatGPT to check this Coaching Letter for possible overstatement of my case.
Here are the complete citations:
Kalyuga, S. (2007). Expertise reversal effect and its implications for learner-tailored instruction. Educational Psychology Review, 19(4), 509–539. https://doi.org/10.1007/s10648-007-9054-3
Kalyuga, S., Ayres, P., Chandler, P., & Sweller, J. (2003). The expertise reversal effect. Educational Psychologist, 38(1), 23–31. https://doi.org/10.1207/S15326985EP3801_4
Renkl, A., Atkinson, R. K., Maier, U. H., & Staley, R. (2002). From example study to problem solving: Smooth transitions help learning. The Journal of Experimental Education, 70(4), 293–315. https://doi.org/10.1080/00220970209599510
And to throw a bone to Zach Groshell, he does a nice job of explaining why extraneous cognitive load is a bad thing for learners in the first chapter of his book, which is called, somewhat predictably, Just Tell Them!
Just so you know, I’m a paid subscriber of Greg Ashman’s, because I use his work occasionally in our workshops, and I think that if you’re using someone’s stuff you should pay for it. Some of you may have seen his post on differentiation: Where is the evidence to support differentiation? which we use as part of a jigsaw on differentiation. So I don’t think he’s wrong about everything. But he seems to be inexplicably determined to condemn BTC, even though I don’t think he understands the first thing about it. Since this is a footnote that no one’s going to read anyway, I’m just going to say that I find this anti-intellectual, hubristic, and puzzlingly ironic for someone who’s always complaining that people deride Direct Instruction when they don’t fully understand it. So here’s his main anti-BTC post, Building Thinking Classrooms is still nonsense, and here’s my (snarky) response to him and others like him who don’t seem to want to understand before they broadcast their opinions: Coaching Letter #198: On understanding before opining. And also, perversely, I am grateful for being forced to do the research. I “knew” that he was wrong, but I couldn’t fully explain why, and now I think I am a lot closer.
and adults—it only occurred to me much later that at least two of the activities we use in workshops all the time are exercises in productive failure. No, of course I’m not going to tell you which ones. But it does mean that I can personally testify that people in the situation of not being able to find a right answer do not despair, they do not yell or cry in frustration, and they do not give up. They do not give any appearance of floundering (a word that shows up all the time in the posts of the CLTists). They persist. And we also have evidence—because they tell us—that they learned more from the exercise than they would have otherwise.
The CLTists, however, don’t want the help. I imagine them, Monty Python-like, standing on the parapet, waving their swords and yelling, “Begone, thou foul and pestilential creatures! Take thy vile and sordid ideas and drown thyselves in yonder lake! We would rather die than accept thy pathetic and ill-conceived offers of assistance!” Sorry, I find the idea of footnotes rather too freeing.
And because we are being pedantic here, I’m going to add what ChatGPT told me I should: This does not mean that research on productive failure validates every BTC practice or proves that BTC will work under all conditions. It does mean that the central idea of having students grapple with a well-designed problem before receiving full explanation has a much stronger research basis than BTC’s critics sometimes acknowledge.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.