RSS Amplifier

Kelly Webb-Davies · Aug 4, 2026

Validity before Categories

0
Sign in to vote or save

Kelly Webb-Davies · Kelly Webb-Davies

Categories. AI-use category systems everywhere. I did hope we’d be past this by now 🚦

A close-up of a traffic light against a cloudy sky. The red light at the top and green light at the bottom are unlit, while the middle amber light is illuminated and contains the word "AMBIGUOUS" in bold black capital letters, visually suggesting uncertainty rather than a clear stop-or-go decision.
Image generated by ChatGPT

I understand why they are everywhere, and I don’t think anyone using them is acting in bad faith. Staff and students are desperate for clarity. Institutions need policy language that survives a committee. Course teams need something they can actually paste into an assessment brief on a Tuesday afternoon, and nobody wants a thousand competing interpretations of what students may and may not do with generative AI. So we get categories. AI prohibited, AI minimal, AI permitted, AI integrated, red, amber, green, levels one two three four five. Or some other vocabulary that tries to carve AI use into cute little communicable chunks.

But I think we are doing it backwards and trying to give people colourful clarity about the wrong thing.

My worry is that AI-use categories are a very visible way to point people firstly towards permissions. But no amount of precision about whether a student is allowed to use AI to proofread, summarise, rewrite or polish gets at what is actually going wrong.

The first question should actually be: what capability is this assessment really trying to evidence? And then: what would make that evidence trustworthy?

An AI-use category cannot answer either question. It can tell a student what they are permitted to do. It cannot tell us whether the assessment still supports the claim we want to make about that student’s learning. A declaration can ask a student to report what they did. It cannot show if the student did the relevant learning work. A policy can require staff to communicate permitted use more consistently across a faculty. It cannot make the assessment valid.

A category-first approach lets us feel as though we have done something that looks like assessment redesign when what we have actually done is AI-use communication. It’s like assessment redesign cosplay. And while communication is very important — students do need to know what is expected of them, staff need shared language, institutions need consistency — I am arguing that shared language is most useful when it points people towards the right problem in the first place.

What is the capability? Where is the evidence? Under what conditions was it produced? Can we even know for certain the identity of the person who completed the assessment? Could this final product have been produced by someone who does not have the capability we are claiming it demonstrates? And if the answer to that last one is yes, what are we going to do about it?

That is validity. It needs to come before categories or permissions.

We are overwhelmingly research institutions. We basically have the best research methodology skills out there. We could be applying them to assessment.

I feel like categories are partly an attempt to make this simple, to hand busy people an accessible way into work that can feel overwhelming. And while I do understand the instinct, I do not think we need the shortcut as much as we assume, because we should already have this knowledge. Every research paper we write makes a claim and then shows its working: what data supports the claim, what threats to validity exist, what design choices make the findings defensible. Peer reviewers would reject a study that stated its conclusion and offered no controlled evidence for it. This is what we do and what we teach. An assessment makes a claim too, that this student can do this thing, and we should be perfectly capable of designing one that would stand up to the same scrutiny. But too often instead we state the capability, set the task, receive the product, and infer the learning, too often without controlling for the things that might threaten that inference.

An AI-use category is the assessment equivalent of a drug trial that instructs participants not to take other medication but never checks whether they comply. We would flag that as a threat to validity in a study. But a drug trial at least has statistical power: a large sample can absorb some individual non-compliance without destroying the population-level finding. Assessment does not have that cushion because we are not making a claim about an entire cohort. We are certifying that this individual student has this specific capability. It is n=1. Every uncontrolled variable lands directly on the inference, so a rule that cannot be verified is not a methodological control.

It is faith. And we are awarding degrees on the basis of it.

Even before generative AI many of our assessments were already carrying far more evidentiary weight than they could bear. Generative AI just took assumptions about assessments that were already fragile and kicked the scaffolding out from under them, which is why I don’t think we can afford another year of polishing categories while believing it’s redesign.

This work should have started two years ago, or more. Where it hasn’t started, it needs to start now.

The other problem is that categories look much clearer than they are. They feel student-friendly, because labels feel like help. But those labels hand extremely difficult boundary decisions to the people least equipped to make them: students, who are still learning the very skills those decisions require.

Borges wrote that “there is no classification of the universe that is not arbitrary and conjectural.” Categories about writing such as planning, drafting, editing and translation can look as though they are clean natural divisions. But they are not. Where does planning end and drafting begin? When does editing become rewriting? When does translation reshape argument? When does feedback become co-composition? When does “improving clarity” change meaning? When does a prompt become conceptual work? We are imposing a framework on a complicated, nuanced process, then asking students to behave as though the boundaries were real. Borges was not arguing that classification is useless; we do need ways to organise and make sense of the world, and they have driven much of our scientific progress. The danger comes when we forget that they are provisional, and mistake the category for the thing itself.

Writing is mediated, iterative, recursive, social, embodied, hard and very very messy. We think through notes, speech, drafts, reading, feedback, false starts and half-sentences, often sideways and in the wrong order. I spend most of my professional life thinking about writing, AI and assessment, and I absolutely cannot draw those boundaries cleanly in my own work. I could not honestly instruct a student to use AI for “this part” of writing but not “that part”, as though writing were a sequence of discrete containers. Generative AI writing assistance touches all of this and it does not stay in some magic box that we assign it.

So when we manage assessment through categories of AI use, especially in writing, we are asking students to police boundaries that we struggle to define ourselves, under time pressure, with their degree classification attached, and with every incentive to resolve unclear boundaries in their own favour.

That is not clarity. That is just more ambiguity with a label slapped on it.

I have long been a supporter of the two-lane approach from The University of Sydney. And while it also puts labels on assessments, the key difference is what those labels are for. In practice, AI-use categories are often applied as permission labels: what a student may do. A lane labels validity: whether the evidence a task produces can be trusted at all.

Lane 1 means the assessment produces structurally secured evidence of student capability. That evidence might come from an exam, or equally from a viva, a live practical, an observed studio task, or a stage that captures reasoning before a written submission exists. The point is not the format. The point is the evidence that the format results in.

Additionally, it is worth being clear that secured does not have to mean no-AI. What defines Lane 1 is that the conditions are controlled enough to trust the evidence, which means AI can be part of the task, used in the open and under observation, rather than banned by a line in a brief. A supervised session where students work with AI and have to account for it is still Lane 1. An unsupervised take-home that simply states “no AI” is not.

Lane 2 means the assessment is open. Students have access to AI, tools, examples, templates, translations, explanations, and everything else that exists in the world, because that is often educationally valuable and professionally authentic. It just has to be designed taking that into account.

AI categories classify permissions. Two-lane classifies evidentiary conditions. That is a different kind of clarity, and it moves the conversation from “how do we control every possible AI use?” to “what does trustworthy evidence of capability actually look like and where does it come from?”

Lanes work in one direction: you already know what the assessment evidences and how secure that evidence needs to be, and the label follows from what the assessment already is. An AI-use category approach on the other hand will not do that work for you. Whatever category framework you begin from, the discipline-specific design still has to happen: what counts as trustworthy evidence in this subject, on this module. If that work has to be done either way, it is worth asking why we so often start with the permission category rather than just with the design itself.

Students are desperate for clarity, and we have mostly read that as a demand for clearer rules. However, a colleague at the University of Sydney showed me a different kind of clarity. She told me that a major benefit of the two-lane framework is that, because students are no longer the ones deciding how much AI use is acceptable on a given task, they find it much clearer. The decision is already built into the assessment, so there is nothing left for them to adjudicate.

I think trying to control the whole workflow, or asking students to do so in good faith, is a losing game, and it creates a miserable place to teach and study. Staff become rule-writers and suspicion-managers. Students become compliance interpreters. Everybody ends up litigating whether something counts as brainstorming, editing or improper assistance, and in the end there is less space to talk about learning.

I do know why institutions are moving carefully. Staff are exhausted. Approval cycles are glacial. Assessment regulations are clunky and often older than the problem. External requirements exist. Nobody wants to stand in front of a course team and announce that every assessment needs rethinking by September.

But the gentle approach becomes dangerous when the easier approach occupies the space where the necessary work should be. Category-first policy can function as institutional displacement activity: visible, documentable, circulatable, but also structurally beside the point. We build elaborate permission architecture around assessments that still might not even evidence what they claim to evidence.

I would much rather we spent that energy on validity-focused assessment redesign, and then on talking to students about what they need to learn and how to show us that, because we would finally be measuring it properly. What capability is this task designed to develop? Why does it matter in this discipline? Where might AI help you practise, explore, check or extend your understanding? Where might it do the learning work for you? What will you need to explain, defend, verify, perform or adapt later, without help? And how is this assessment going to give us evidence of that?

That is a much more productive conversation than “here are the rules, don’t cross the line.” It also shifts the frame from responsible AI use, which slides too easily into rule-following, to effective use: whether the AI actually served the student’s learning. That kind of design takes real time to get right, but it is the conversation worth having.

This is the progression I would want to see: first, we give students the disciplinary knowledge AND the study skills to make good decisions about their own learning. That is our responsibility as educators. Second, we trust them to use that knowledge, including making decisions about when and how to use AI. That is where the trust sits, and it is trust we have always extended in some form. Third, we design assessments that actually check whether the learning happened and whether those decisions were good ones. That is what secures the trust, rather than simply assuming it was well placed.

So yes, communicate AI use, but put it where it belongs: last. Before that, identify the capability. Decide what evidence would support the claim that a student has developed it, on the assumption that generative AI already exists in the environment they are working in. Decide where that evidence needs to be secured. Decide what AI-assisted mediation might be appropriate around it. Then, and only then, communicate the AI-use expectations, with a category if it genuinely helps your context. Do that design work first, and then do not be surprised if you barely need categories or rules at all.

If we start with permissions, we have framed the problem as student behaviour. If we start with evidence, we have framed the problem as assessment design. That puts the onus on us, as it should be.

Assessment design is the work. It is harder than writing categories and slower than adding a line to a brief. It needs disciplinary judgement, programme-level thinking, staff time, accessibility work and institutional nerve. Do that hard work first, though, and everything downstream gets easier. The assessment is valid. The evidence is secure in the way it needs to be. There is no long list of rules to hand a student, because a student just needs to do the assessment. What is left over is the easier and much more important conversation: about learning.

Validity before categories. Evidence before permissions.

Start with the capability.

Figure out the evidence.

Then write the rules. You will likely need fewer of them than you think.

Opinions expressed are the author's own.

This post was originally published on LinkedIn on the 28th of July 2026.

Challenge failed on this post :( I don’t want to edit the text either because it was published on LinkedIn first.

No posts

Read the original on kellywebbdavies.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.