The next generation of AI tools must help learners build the skills the Age of AI demands
Much of the current conversation about AI in education is, understandably, focused on performance. Can the model follow instructions? Can it generate accurate answers? Can it respond appropriately across languages, cultures, and learner profiles? Can it avoid harmful or biased outputs?
These are important questions and education systems should not adopt AI tools that are unreliable, unsafe, or only effective for dominant languages and privileged learners.
But beyond these technical foundations, the central question for education is not simply whether an AI system performs well. It is whether it improves learning and teaching.
This distinction may sound obvious, but it is often missing from how AI tools are designed, benchmarked, and also funded, and consequently, evaluated. Many current benchmarks assess whether a model can produce a correct or acceptable response. Fewer ask whether interacting with that model helps a learner understand more deeply, think more clearly, persist for longer, collaborate better, or transfer knowledge into new contexts.
Take the example of an AI tutor that may produce fluent explanations. Or a chatbot that may answer questions quickly. A writing assistant may generate polished text for students. But if these tools do not strengthen learners’ knowledge, confidence, reasoning, creativity, or agency, then their educational value remains uncertain.
Beyond basic test scores
Schools are places that cultivate knowledge where understanding is nurtured over time, friendships are built, social skills are developed etc. through an interactive and iterative process of attention, struggle, feedback, reflection, social interaction, and meaning-making.
AI systems will only serve education if they can support the full process of learning.
This is why the question of evaluation is so important: if we benchmark AI systems mainly on whether they produce correct responses or increase currently measured outcomes, we risk incentivising tools that look educationally impressive but do not necessarily deepen learning. We risk confusing performance by the machine with progress by the learner.
Currently, when educational outcomes are discussed, they are often reduced to standardised measures of literacy, numeracy, or subject-specific attainment. Foundational learning remains urgent, and any serious EdTech evaluation must attend to whether tools improve core academic outcomes.
In parallel, the case for AI in education is often made on broader grounds — we are told that learners need to be prepared for an AI-shaped future: a future requiring critical thinking, creativity, collaboration, judgement, ethical reasoning, and adaptability.
If that is the justification, then our evaluation frameworks need to reflect it… We should be asking whether AI tools help students ask better questions, not only answer existing ones. Whether they encourage learners to compare evidence, challenge assumptions, and revise their thinking. Whether they support collaborative problem-solving rather than isolated task completion. Whether they help teachers understand learners more fully, rather than merely automate assessment. Whether they expand opportunity for students who are currently underserved, rather than intensify existing inequalities.
These future-ready skills are harder to measure but with current research and tools, we must rise to the challenge. How many AI-enabled EdTech tools do you know of that are developed to optimise for these skills? We would love to see more strong examples… Because if AI in education is optimised for traditional outcomes, we risk building the future while leaving the next generation unprepared to live in it.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.