In Short:
To be educated once meant you could produce the work — write the essay, solve the problem, sit the exam. A machine can now produce all of it. So exam season asks the question underneath: what can this student actually do, and how do we know?
Generative AI did not remove the need to prove skill; it removed our cheapest proofs. The answer is not better detection — that game is unwinnable — but redesigning assessment so the student is visible again.
Three principles do that: assess raw human performance with the AI stripped away, test often and in the open, and judge take-home work across several artefacts rather than one. Let AI in on the way to learning, and keep it out of the moment we measure what a student can do.
Education has always graded proxies. An essay stands in for can this student think; a problem set for can this student reason; an exam for does this student know. We trusted them because producing the work required the skill — you could not write a good essay without being able to think.
Agents sever that link. An excellent paper can now exist without a student who could have written it. Calculators, Wikipedia and Google Translate each broke a single proxy, and each time we moved assessment somewhere the tool could not reach. Agents are different: they break every proxy that can be produced on a computer.
Rather than a moral crisis around plagiarism, this is a design problem. If our old proofs of skill no longer work, how do we design new ones? Not by chasing detectors — Andrej Karpathy and a lengthening list of universities have concluded that fight is lost — but by rebuilding assessment around the one thing agents cannot fake: the student in front of you.
Until now, we assumed a student’s work was their own unless shown otherwise. It’s time to reverse it: assume anything taken home may have been produced by an agent, and grade accordingly — this is fairer than an AI ban, which only serves to punish students who follow the rules, and reward those who are good at prompting AI for plagiarism. AI can still be kept out of the hall exam, the viva, the live defence, in environments that the school can control, but many traditional proxies for student skill have been broken, and we need to react accordingly.
Students will use AI — and they should — to revise, to make flashcards, to generate practice papers. None of that changes the expectation that they can still perform without it, because that personal capability is the wall everything else leans on. A randomised study this year had 52 programmers learn to use a new software library; those using AI were no faster, yet scored 17% lower on a mastery test afterwards. You cannot catch a hallucination in a subject you never learned, and advanced AI systems mask a students’ actual capability.
The most reliable way to see that capability is to remove the AI and watch. Closed-book sit-down exams are expensive and can be problematic, but they are still among the best assessment instruments we have against AI plagiarism. The same holds for anything performed live — presentations, debates, vivas, a speech prepared in thirty minutes and delivered from the podium. It can be recorded, timed and questioned on the spot, and no agent can sit it for the student. When a Yale undergraduate was placed on disciplinary leave this March, what settled the case was re-sitting the work without AI, where he did worse on exactly the questions AI would have answered.
A course built on a midterm, a final, and one essay rests on three data points. That was always thin, and now that AI agents exist, it is reckless. Teachers need more data points with which to evaluate whether a student’s submission actually reflects their subject matter knowledge, or if it’s from a Claude or ChatGPT subscription. The quickest fix any teacher can implement now is more formative assessment — frequent, in-class checks that you can trust because you can see them work in front of you. It’s good practice anyway, with Black and Wiliam’s review finding that strengthening this kind of frequent feedback produces some of the largest gains in the education literature, and Roediger and Karpicke showed that testing itself deepens memory rather than merely measuring it.
Recall the many rows over suspected AI plagiarism? If the teachers and students had a weekly record of the student’s quiz performance over the term, that may have diffused much suspicion. Nor need it add to a teacher’s workload: AI can write quiz questions from your slides, mark short answers against a rubric, and flag the student whose classroom performance and homework do not match.
However, it doesn’t mean we have to cut AI out of all homework and assessment. By the time our students enter the job market, no company would hire workers who are AI illiterate, so assessment needs an appropriate ratio between authentic human-only work, and AI-augmented output. For example, basing 80% of course grade on authentic in-person assessment, formative snapshots, and other human-only work, while 20% belongs to homework, where AI use is assumed (and welcome). I’d adjust the ratio over time also, giving older students (high school and above) more space to use AI, while the youngest students may not need to have AI in their curriculum at all.
But what does that AI-enriched homework look like? To gauge understanding of a topic, I’d want the student to submit several artefacts in different forms: a presentation, a piece to camera, an infographic, so that I can triangulate among them to see hints of the student’s actual understanding. Here’s an example of vibe-coding game design as assessment:
In this course, students are invited to create games for each other to play, with the objective that through playing this game, their peers would learn something about a UN SDG of their choice (you can play the games by clicking on the links). Heritage Hunt is a geoguessr game over Hong Kong’s 263 graded historic buildings, which is fun but doesn’t reflect a lot of subject matter understanding. Hong Kong Waters is a marine conservation simulation set in Hong Kong waters, where you weigh budgets and policy as the ecosystem responds.
You do not need the source code to know which student understood more. Pairing this with a presentation, infographic, report, or simple Q&A will let you triangulate: It’s hard to gauge a student’s understanding based on a single piece of work, but if a concept is misapplied across formats, that’d reveal far more than any single tidy submission. This does ask more of students, but they’d be doing this with powerful AI systems on their side. The same systems are on teachers’ sides too, to help us assess and monitor student progress across all the quiz answers, presentation transcripts, debate videos, pitch decks, essays, and other artefacts that combine to paint a rich picture of the students’ real proof of skill.
When an agent can produce any single artefact a student hands in, trust in teachers’ judgement become essential for holding the integrity of our assessment system together. The teacher who can tell who actually understands — who is bluffing, whose eyes light up when the question gets harder — becomes the thing the whole system depends on.
For a century, to be educated meant you could produce the work. That definition is quietly ending, because producing the work is exactly what machines now do. What remains is harder to fake and easier to recognise: what a student genuinely understands, what they can defend unaided, and the judgment to know when to lean on the tools and when to put them down. That is what assessment has to look for now — and what it means to be an educated person is changing under our feet.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.