RSS Amplifier

How We Frame Machines · Aug 3, 2026

Measuring How We Reason: The New Blueprint for AI-Era Learning

0
Sign in to vote or save

Mike Kentz · How We Frame Machines

My friend Nneka McGee and a stellar group of co-authors (Candace Thille, Ikkyu Choi, Kadriye Ercikan, Isabelle Hau, and reviewer Matthew Johnson) recently released a white paper out of the Stanford Accelerator for Learning and ETS titled Responsible Assessment in the AI Era.

It comes out of a January convening that brought together a hundred top minds across research, technology, K-12, and higher ed. If you care about the future of learning, I recommend carving out some time to read it.

Nneka and her co-authors have handed the education field a vital roadmap. For anyone struggling to figure out how to evaluate actual human learning in a world where AI can fake any static assignment, this report provides a clear lens for where assessment must head.

Rather than treating assessment as a single test score or an essay inspection, the paper frames assessment as a continuous system of evidence gathering. Conversational learning and interactive simulations emerge not as the sole answer, but as one of the most promising spaces for capturing how people actually think, reason, and adapt.

Below is a breakdown of why this paper matters, how it diagnoses the breakdown of traditional grading, and 7 key examples from the report that give us a new framework for evaluating process over product.

The paper starts by discussing something we have all felt in the classroom: traditional assessment is fundamentally broken.

Generative AI did not break it; it simply exposed the crack that was already there. As the authors highlight when citing OECD research on AI and assessment, grading only final products risks measuring a student’s technological fluency rather than their actual human capability.

Even worse, Stanford’s Cindy Mazow points out the evaluation trap: the moment a grade gets attached to a task, students stop taking risks. They hide their confusion. We spend all our energy trying to build safe spaces for messy thinking, only to destroy that safety the second we evaluate the final artifact.

So what is the alternative? Process data.

The paper defines process data as the record of actions, steps, and timing a learner takes while working through a task. Instead of inspecting a finished product after the fact, process data captures how a student navigates a challenge as it unfolds.

This is where conversational learning and interactive simulations become so relevant. When a student works through an assignment by engaging in a live dialogue with an AI system, every prompt, reframe, and response generates a rich stream of process data. The resulting transcript becomes a map of their thinking rather than just a polished, potentially ghostwritten final draft.

That distinction between evaluating a static final product and evaluating an interactive process changes the entire learning dynamic.

Because a conversation varies naturally every time a student runs through it, process-based assessment inherently supports iteration. A student can engage in a simulation, receive targeted feedback on how they reasoned through the dialogue, adjust their strategy, and try again. Evaluation ceases to be a single, anxiety-inducing judgment event and becomes an ongoing learning loop. We have been running trials around this at AI Friction Labs, and watching how students navigate an interactive conversation reveals cognitive habits that an essay could never surface.

Framework from "Responsible Assessment in the AI Era" (McGee et al., Stanford/ETS).

Share

Section 2.3.b.ii of the paper offers an inventory of tools researchers are exploring to measure human-centered skills. What makes this section useful for practitioners is that it expands our definition of evidence far beyond written answers.

Whether you are designing classroom tasks, building edtech tools, or setting school policy, these 7 examples from the report highlight how institutional research is operationalizing interactive and process-based evaluation:

  1. Dialogue-Based Evidence Systems (Diego Zapata-Rivera, ETS): Using Evidence-Centered Design (ECD) to build AI agents that help learners demonstrate what they know through ongoing dialogue rather than static, one-way testing.

  2. AI as an Interactive Role-Play Partner (Patrick Kyllonen, ETS): Deploying AI as a conversational partner to elicit and measure real-time social dynamics, communication, and interpersonal problem-solving.

  3. Real-World Conversational Scenarios (Patrick Kyllonen, ETS): Utilizing AI to generate dynamic scenarios (drawing on proven capability assessment models from adult professional learning) where students navigate complex situations.

  4. Interactive Behavioral Signals (Jesse Sparks, ETS): Reading authentic signals of persistence and resilience from how learners interact within open-ended environments like block-based robotics.

  5. Operationalizing Interaction Rubrics (Sparks et al., 2025): Establishing formal psychometric rubrics that make raw interaction logs legible and valid to researchers. This is the exact foundational work needed to evaluate process data at scale.

  6. Recognizing Construct-Irrelevant Variance (Kadriye Ercikan, ETS): Identifying key measurement traps, such as ensuring AI scoring does not accidentally reward pure verbal fluency over actual domain reasoning during an interactive task.

  7. Keeping Humans in the Loop (Section 4.3.c / OECD Frameworks): Advocating that interactive evidence should empower educators rather than fully automate high-stakes decisions, ensuring assessment remains human-centered.

"We need instruction that produces adaptive learners and assessment that can tell." - Stanford GSE Dean Dan Schwartz (p.10)

When a hundred researchers convene at Stanford and agree that static, single-event testing is no longer serving us, it is a huge signal for where the field is heading.

Conversational simulations are not the only tool for the future, but they are a uniquely powerful one. They give students an adaptable space to practice, receive feedback, and try again, while unlocking a rich stream of process data that shows how people actually reason and adapt.

Nneka and the team at Stanford and ETS have provided a vital bridge. They have connected cutting-edge exploratory work in interactive learning with the rigorous psychometric theory required to make it credible.

This report gives the entire field a shared foundation to build on, and I could not be more excited to see where this conversation goes next.

Doan Winkel and I are hosting free virtual launch event for AI Friction Labs on Wednesday, August 19th at 1pm EST. We’ll demo the platform and you can engage in one of our conversational simulations yourself. Register here: [AI Friction Labs Faculty Demo Event Registration]

Read the original on mikekentz.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.