RSS Amplifier

Edtech Insiders · Jul 10, 2026

The Three Jobs of Teaching a Child to Read, and Where AI Fits In

0
Sign in to vote or save

Sarah Morin, Alex Sarlin, Ben Kornell, Lin Ler, Jen Lapaz · Edtech Insiders

Thanks to our Presenting Sponsors Hire Education, Tuck Advisors, and Cooley, for making Edtech Insiders possible.

To learn more about becoming a sponsor, email us at info@edtechinsiders.org

Sponsored by Overdeck Family Foundation

Lin Ler is a Stanford MBA and MA in Education candidate working at the intersection of business strategy, learning science, and frontier AI. With a background spanning management consulting, EdTech startups, and the social sector, Lin translates complex AI capabilities into practical insight for the future of learning.

On the 2024 Nation’s Report Card, 40% of American fourth graders scored below “basic” in reading, and the share of eighth graders below “basic” was the largest in the assessment’s 32-year history. Into that emergency walks AI, carrying the playbook that worked in math: tight feedback loops, hints before answers, a tutor that never tires — the design that lifted math scores by a third of a standard deviation for five dollars a student in Ghana. Whether that playbook transfers to reading is a genuinely open question, and the early trials split. In 2024, a randomized trial in Beijing found a chatbot helped kindergartners understand stories as well as a trained human did. A year later, a randomized trial across seven English schools found teenagers who studied texts with an LLM learned less than classmates who took notes. Two studies can’t settle anything, but the split between them is instructive, because the two chatbots were doing entirely different jobs. Teaching a child to read involves three jobs. It’s worth auditing each one and how AI fits into it.

Watch a parent or teacher work with a young reader and you’ll see three jobs performed so fluidly they look like one.

The first is talking. The adult pauses mid-page and asks, “Why do you think the bear is hiding?” The child guesses, the adult builds on the answer, asks a follow-up, connects it to the child’s own life. The child does most of the talking; the adult steers. This dialogue around the text is where comprehension and vocabulary grow.

The second is listening. Now the child reads aloud, and the adult sits beside them tracking every word: catching the stumble on “through,” noticing that words with silent letters keep causing trouble, deciding whether to correct now or let this one go. The child produces; the adult diagnoses.

The third is supplying. Before any of that can happen, someone chose this book — hard enough to stretch this child, easy enough not to defeat them — and decided which questions were worth asking about it. The child never sees this job. It happens the night before, in the teacher’s planning hour.

Every AI reading product is a bet on at least one of these jobs. Chatbot reading companions bet on talk. Speech-recognition reading tutors bet on listening. Content generators bet on supply. The results diverge sharply enough that “does AI teach reading?” has no single answer. So let’s audit, job by job, what the research shows.

In reading, the conversation is the curriculum. Four decades of research on dialogic reading — the questioning technique Grover Whitehurst formalized in 1988 — shows that children’s language and comprehension grow when an adult asks questions during shared reading: prompting the child, evaluating the answer, expanding on it, repeating the prompt. The technique works. The bottleneck has always been the adult: many parents and teachers never learn it, and the children with the least access to trained questioners are the ones who fall behind first.

The news is that a machine can now carry that conversation. In the Beijing trial, published in the Journal of Experimental Child Psychology, 148 kindergartners from less-resourced families were randomly assigned to share a picture book with either a chatbot or a trained human, using either dialogic questioning or plain narration. These children are pre-readers; the book was read to them, and what the trial measured was story comprehension — how well a child understands a story they’ve heard, the foundation that reading comprehension is later built on. The dialogic technique drove the gains in comprehension and word learning, and it worked just as well from the chatbot as from the human. On one measure, learning four-character idioms, the chatbot-led sessions actually came out ahead. The control condition is just as instructive: the chatbot that simply narrated the story, no questions asked, produced the lowest comprehension scores of all four groups. The technology without the technique is inert.[1]

This is a different mechanism from the one our math article found. In math, the effective AI is a gatekeeper: it withholds the answer the student is trying to extract, forcing the productive struggle. In storytime there is no answer to withhold. The effective reading AI is an initiator: it asks questions the child was never going to ask, and the child does the cognitive work of answering. The principle underneath is the same — learning happens when the child does the thinking — but the design move is the opposite. Math AI succeeds by refusing to talk. Reading AI succeeds by starting the conversation.

The product implications are already visible. StoryBuddy, a CHI-published system from the same research lineage, generates dialogic questions for any storybook rather than a preset catalog, and lets parents curate which questions get asked. Parents in non-English-speaking households saw a use case the researchers hadn’t led with: a grandparent who speaks no English can now supervise English storytime.

The finding stands as the first result of its kind in this series: an AI matching a trained human at a core teaching practice. The machine didn’t replicate the adult’s warmth. It replicated the adult’s questions. That turns out to be the part that teaches.

Here is the trap in Job One: there is a second kind of “AI reading companion” that looks identical from across the room — a student and a chatbot, talking about a text — but runs the conversation in reverse. In Beijing, the machine asked and the child answered. In the products teenagers actually use, the student asks and the machine answers. That reversal flips the learning result.

The English schools trial, run by Cambridge University Press & Assessment and Microsoft Research with 344 students aged 14 and 15, tested exactly this setup: students studied history passages with an LLM chatbot available to answer their questions, against classmates who took notes. Three days later, tested without any tools, the note-takers won on every measure — literal retention, comprehension, and free recall. The students hadn’t used the LLM foolishly; 90% asked for elaboration and deeper explanation, exactly the prompts you’d hope for. They lost anyway. Answering a question about a text forces you to select, connect, and put ideas into your own words — that effort is what makes reading stick. When the machine does the answering, the machine gets the practice.

The students couldn’t feel it happening. They rated the LLM easier and more helpful than note-taking, and preferred it 42% to 27%. One in six pasted the LLM’s output into their notes nearly verbatim. Out of 4,929 prompts across the study, exactly six questioned whether the LLM was right. Readers of the first article in this series will recognize the shape: cognitive debt in a reading jacket — the tool that feels most helpful teaching the least.[2]

So the audit of Job One comes down to one observable: which way the questions flow. When they flow toward the child, the child does the thinking, and the learning follows. When they flow toward the machine, the machine does the thinking, and the learning goes with it. After ten minutes with any reading tool, count who has said more about the text — the child or the machine.

The less glamorous job in reading instruction is sitting beside a child, hearing them read aloud, and catching the stumbles. It is also the job AI has been learning for the longest. Project LISTEN, launched at Carnegie Mellon in 1990, built a Reading Tutor that displayed stories, listened to children read aloud through speech recognition, and intervened the way an expert tutor would. In a seven-month comparison across 178 students in grades 1–4, students using the Reading Tutor outgained peers doing sustained silent reading on word identification, comprehension, and fluency, with the largest effects in first grade — and the lowest-proficiency students often gained most. The lineage predates chatbots by a quarter century, and it has been quietly compounding ever since.

The current commercial version of the idea looks like BuddyBooks, and its design is worth dwelling on, because every element maps to something reading science already validated. The platform co-reads: the machine reads a short section aloud with the words highlighted — modeling fluent pacing — then the child reads the next section, and speech recognition scores their accuracy and fluency in real time. Text is chunked into short segments to build stamina without overwhelm. Sections the child struggled with get flagged for Review Mode, where the child hears a fluent model, hears their own attempt, and rereads — which is repeated reading, the classic fluency intervention, automated. The machine never sighs, never rushes, never looks disappointed, which matters for struggling readers who have learned to dread reading aloud. In a 2024–25 study of 107 students with dyslexia in grades 2 through 5 in a Texas district, students who engaged consistently — at least ten focused minutes in most weeks — outgrew less consistent peers on i-Ready and posted nearly double the growth on the state reading test. The study is correlational, so it suggests rather than proves. But its design lesson stands on its own: total minutes didn’t track gains, and neither did total weeks. Only regular, focused practice did.

The capability underneath is improving fastest where trained assessors are scarcest. A 2025 study in South Africa fine-tuned speech models to score early-grade oral reading in Xhosa and reached roughly 91% diagnostic accuracy with a few hundred labeled recordings per test item; the same task given to a state-of-the-art multilingual model off the shelf scored 6.9%. Child speech in noisy classrooms breaks generic AI, and modest local investment fixes it. A growing research stack — from child-speech LLMs to automated fluency scoring in Tamil and Malay — is working the same problem across languages. For countries that assess early reading through oral tests administered by scarce trained examiners, this is a genuinely new capacity.

The honest summary of Job Two: the modern studies are descriptive, and no recent randomized trial certifies that today’s AI listeners cause reading gains. But this is the job with working products, a validated mechanism, and a thirty-year pedigree. Before AI can teach a child to read, it has to be able to hear them. In English, that problem is largely solved. Everywhere else, it is solvable and cheap.

The job AI looks most obviously built for is generating the material: passages, comprehension questions, leveled versions of the same text for a mixed classroom. This is where the audit turns up the quiet failure. Generation is solved. Pedagogical targeting is not.

Consider what happened when researchers at ETS asked GPT-4o to generate reading comprehension questions targeting specific inference skills — the diagnostic precision a struggling reader actually needs. Expert reviewers judged 93.8% of the generated questions good enough for operational use in grades 3 through 12. Impressive, until the second number: only 42.6% tested the inference skill they were asked to test. A third of the output drifted into literal, fact-lookup questions requiring no inference at all. The items were polished. They just weren’t measuring what they were supposed to measure, and a question bank that can’t target a skill can’t diagnose a struggling reader.

Text leveling shows the same signature. A 2025 evaluation had three LLMs simplify sixty 12th-grade passages down to 8th-, 6th-, and 4th-grade reading levels, across 2,315 outputs. Simplifying to 8th and 6th grade worked. At 4th grade, not one model-and-prompt combination reliably hit the target; the texts stayed too hard. And the further down the models reached, the more meaning they shed along the way. The pattern is exactly backwards from the need: the students requiring the deepest simplification — young readers and English learners furthest behind grade level — are the ones current models serve worst.

The mechanism is worth naming because it will outlast these particular models. LLMs control the form of educational content — clean questions, simpler words, fluent prose — far better than they control its cognitive demand, which is the property that makes content teach. In both studies, performance was rescued the same way: human structure injected into the pipeline, expert-written examples in one case, expert-chosen keywords in the other. That is not an argument against AI-generated reading material. It is an argument that “the AI wrote it and it reads fine” is the wrong acceptance test, and every reading product built on auto-generated content is currently relying on a human check it may or may not be doing.

Score the three jobs against the evidence and the market looks upside down.

Talk earns a split grade: the strongest causal result in the set when the AI asks the questions, delivered through a designed protocol — and a causally demonstrated harm when it answers them for an unsupervised teenager. A for dialogic design; F for the open chatbot.

Listen earns a B, incomplete. Deployable products, a solved capability problem, real-world state-test correlations, a thirty-year evidence lineage — and no modern causal study to certify it.

Supply earns a C minus. Unlimited fluent output, unreliable targeting, and a mandatory human in the loop that the economics of content generation quietly encourage everyone to skip.

Buyers ask whether a reading product “uses AI.” The audit says that’s the wrong question, and it supplies three better ones. Who does the talking? If a session ends with the child having produced fewer words about the text than the machine, the tool is doing the impostor version of Job One. Does it listen? A tool that never hears the child read aloud is not touching the job where AI is most proven. Who checks the text? If no human verifies that generated material hits the level and the skill it claims, the product is running on fluency and hope.

For builders, the collision is the same one the math evidence exposed, transposed into a new subject: the products that demo best — instant answers, instant content — are the ones the evidence grades worst, while the designs the evidence rewards are slower to build and harder to sell. The difference in reading is that the winning design isn’t about withholding anything. A child learns to read by talking about books with someone who listens. The best AI so far is the one built to remember that.

1. The Beijing chatbot was a scripted, purpose-built system that used speech recognition to classify children’s answers — not a generative LLM. What the trial licenses is structured dialogic prompting delivered by a machine, not handing a kindergartner ChatGPT. Outcomes were measured immediately after a single session; how long the gains last is untested.↩︎
2. The trial had no reading-only control, so it shows the LLM losing to an evidence-based study strategy, not to doing nothing. And the pattern may not hold for older learners: in a German laboratory experiment with 334 university students, unrestricted AI-tutor access during textbook study improved unaided test scores, and students kept reading anyway — with gains concentrated among students with strong self-regulation. Teenagers, midway through developing exactly that capacity, may be the population most exposed.↩︎

Special thanks to Overdeck Family Foundation for sponsoring this article in our AI & Efficacy Editorial Research Series diving into key research findings from Stanford’s AI Hub for Education Research Repository (a project by Stanford’s SCALE Initiative).

You can support our work by becoming a paid subscriber, or email info@edtechinsiders.org to learn more about partnerships and sponsorships.

A new independent benchmark, KORA, evaluates how popular AI tools perform on child safety, with leaderboards currently focusing on both models and apps.

Learn more here.

Our friends at Old Soul have launched a new Edtech Job Board! Peruse open roles or submit your open roles to be included.

Learn more here.

New York City Public Schools has paused all new edtech purchases while it reevaluates procurement policies in response to the rapid rise of AI. The move reflects growing efforts by districts to strengthen oversight around privacy, safety, and instructional value as AI becomes increasingly embedded in classroom technology.

Learn more here.

As outcome-based contracts become more common and every investment faces greater scrutiny, Edtech organizations are under growing pressure to demonstrate real impact, not just usage. The companies that win renewals aren’t waiting for ARR and NPS to slip, they’re reading the early signals that predict outcomes and changing the trajectory before it’s too late.

The Educator Index (EI) is a research-based report that surfaces those early, educator-grounded signals of adoption, implementation health, and long-term impact, giving organizations actionable insights before challenges show up in renewal conversations. Thanks to a generous supporter, up to five EdTech companies will receive a fully funded Educator Index Report, including insights that can inform implementation strategy, customer success efforts, QBRs, board discussions, and strategic planning. Apply for a sponsored Educator Index Report and discover implementation risks before your next renewal cycle!

Learn more here.

We recently had Smita Saxena and Alexia Lewis on The Edtech Insiders Podcast!

Smita Saxena is the Founder of Maestro, an AI coaching platform helping K-12 STEM teachers improve instruction and student outcomes. Alexia Lewis is an 8th Grade Algebra I teacher at KIPP Polaris and a Maestro coach whose students achieved remarkable gains using the platform.

  1. How AI coaching saves teachers hours every week.

  2. Why human coaching + AI outperforms either alone.

  3. How AI supports differentiated instruction.

  4. Using AI to turn assessment data into action.

  5. Practical ways AI can reduce teacher burnout.

Listen Here

Thanks for reading! Support our work by becoming a paid subscriber, or email info@edtechinsiders.org to learn more about partnerships and sponsorships.

Read the original on edtechinsiders.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.