Imagine you are a teacher of English for Academic Purposes (EAP). A student you have been teaching for a term submits a piece of academic writing several levels above anything you have ever seen him produce. You have watched him hesitate over articles and prepositions and still get them wrong. You’ve discussed ideas with him that he could not yet write down accurately. Then a text arrives that has none of that. Clean, idiomatic, confident, and completely disconnected from the person you have been teaching.
This happened to me regularly when I was an EAP tutor. And it happened from the moment I began that role in 2019.
Which was three years before ChatGPT.
Language teaching has been dealing with this problem for well over a decade, because machine translation got good while everybody was still making jokes about how bad it was. And I think unfortunately we squandered that head start. We banned it. We wrote Google Translate into academic integrity policies and told students they weren’t allowed to use it.
It did make sense. The construct we were protecting was real: if the thing you are certifying is a person’s ability to use English, then obviously a text produced by a machine that produces English is not evidence of competence. The problem was that prohibition was basically the only thing we tried for unsecured written assessments, and it did not work.
What we could have done instead was treat it as the question it actually was: what are we really trying to certify, and how would we know a student could do it once the text could no longer tell us? Treating it solely as an integrity problem to be policed left us without the expertise to share when generative AI arrived and created the same problem at scale in every other discipline.
Machine translation was easy to dismiss for years because it was genuinely not great, but the dismissals kept circulating for years after that stopped being true. You could spot it. It could not handle idioms. It fell apart on technical vocabulary. All of those criticisms were true, until they increasingly weren’t, and a great many language teachers were still confidently repeating the old version to students who could see with their own eyes that the output was now really very accurate.
Generative AI is running that exact same curve, and at orders of magnitude faster. Almost every current criticism is a claim about a past state of the technology rather than a claim about the technology right now or in the future. The overly formal language. The em-dashes. The invented references. The triads and the flattened prose and the words nobody uses in real life (except that they do). Whatever you object to on grounds of quality or accuracy right now, you are arguing with a moving target, and it is in the interests of the tech companies developing the technology to keep developing it to address the criticisms. It is always the worst it will ever be.
That rules out two of the education sector’s usual responses at once. It rules out detection whether by style, because the stylistic tells are a temporary artefact, or by the statistical probabilities underneath digital detectors, because those signals will always shift with the technology too. And it rules out any assumption that AI output will remain recognisably worse than what a good student produces. Those ships have sailed. If your position still depends on either, it is already out of date.
I have always approached generative AI from linguistics and language teaching standpoint, because that is my training and because generative AI is language technology.
So, I thought: if it is clear to me that there is a difference between having access to and knowledge about a language versus being able to use one fluently…
Then it follows: having access to disciplinary knowledge is not the same as being able to think with it.
What students need to develop, I want to suggest, is disciplinary fluency. By that I do not mean speed, smoothness or polished delivery. I mean disciplinary knowledge that has become available for responsive use: knowledge a learner can retrieve, connect, judge, apply, explain, respond with and adapt as the context changes.
Fluent English produced by machine translation is obviously not English proficiency. A convincing disciplinary argument produced by an LLM is disciplinary performance, not evidence of disciplinary fluency. So, while generative AI has made content production easier, it has not made the capacity to evaluate, apply and reflect on knowledge any less necessary, nor has it made that fluency any easier to acquire.
Working out what students actually need to develop now that information output is cheap and easy is already under way in the sector. A great example is a recent TEQSA-commissioned resource, Assuring Quality Learning in a Gen AI-Integrated Future: The Role of Adaptive Capabilities. It draws together research on the gap between how much you feel you understand and how much you have actually learned, then proposes a framework of adaptive capabilities built on deep disciplinary knowledge as the foundation the rest sits on.
Disciplinary fluency is my language-learning frame for that foundation: knowledge that has stopped being just held and started being usable. Linguistic competence has never simply been a matter of knowing vocabulary and grammar alone. It is knowing what to say to whom and when, the many variations of how to say it, and the deep cultural and contextual knowledge needed to know why you’re choosing to say things the way you are. Disciplinary fluency is the same thing: not just the information a discipline contains, but the ability to think and act responsively with it appropriately in context. I propose that while disciplinary competence is the wider repertoire of knowledge, practice and judgement a learner builds, and disciplinary performance is the observable product (which AI can now generate without being evidence of competence), fluency is how much deep disciplinary knowledge a learner can actually mobilise in context.
Doesn’t everyone know someone who has watched hundreds of hours of subtitled television and still cannot order a coffee, or racked up a thousand-day streak on Duolingo and cannot hold a conversation? Comprehensible input is necessary but it is not sufficient. You have to use the language, badly, repeatedly, in real contexts with a human who responds.
I think that the disciplinary version of this is not only understanding, but retrieving, applying, explaining, being challenged, revising, and then applying again under conditions different enough that you cannot just reproduce the memorised pattern and have it work the same way. Reading the textbook is comprehensible input. Watching an engaging lecture is comprehensible input. So is reading the confident explanation ChatGPT just gave you (and that one is particularly seductive, because it feels like a conversation you had rather than a text you read). All of them feel productive, and none of them are sufficient. Fluency is what repeated meaningful use leaves behind: knowledge you can reach for without being told where to look, and reshape when the situation does not match the example.
But how do you tell fluency apart from a convincing surface-level performance? Not all types of evidence can answer that question equally. The 4Ps framework makes a useful distinction here between product, process, performance and practice as different proxies for learning, each supporting different kinds of inference. Product evidence tells us what was created, but a polished product produced in unobserved conditions does not tell us whether it is evidence of the student’s thinking or a machine’s. Process evidence tells us something about how it was created, but not necessarily what has become cognitively available. Their idea of performance evidence gets closer to what I am proposing: what the learner can actually do with knowledge in real time, responding and adapting as the situation unfolds. If the capability we are trying to infer is disciplinary fluency, this kind of responsive performance may give us much better evidence of it than a finished product or a record of process.
Education very often treats writing as the privileged evidence of learning and thinking. I have written about my problems with “writing is thinking” before, but there is an additional issue here, which is that perhaps writing actually became the default evidence not because it was the best window onto disciplinary reasoning but because it was the most administrable. It produces a portable artefact. It can be marked asynchronously, at volume, with a paper trail, by a marker who has never even met the student. It is mode that we could scale, and wow did we take advantage of that.
Writing can of course genuinely be a mechanism of thinking, and I would not argue otherwise. It is solitary and slow in useful ways. You can stop, look things up, restructure, discover that your argument does not hold, and rebuild it. For a great deal of disciplinary work, that is certainly a capability we want. But being solitary does have its downsides. While you are doing it, there is no one there and no exchange the way there is in dialogue, so the thinking itself is less social, even when it is good thinking.
Writing is, among other things, also a technology for offloading cognition (it allows you to return to ideas you thought then forgot), and spoken conversation leaves far less room for that. It requires enough of what you know to be cognitively available that you can retrieve and reshape it as the exchange unfolds, under time pressure and in response to what the other person just said, with them right there in front of you, waiting. That is what makes conversation hard, and I suspect it is also what might make it better evidence of disciplinary fluency.
The silver lining is that this forces the issue: writing has been carrying much more evidential weight than it can bear. Accepting that AI can take over some of the textual expression work could be the push to move assessment off the page and into spoken expression and conversation, where I would expect the evidence of fluency to be more visible.
The gain is not only evidential. Leaning back towards dialogic interaction brings relational, human contact back into the classroom, which is worth wanting for its own sake, and it’s exactly why disciplinary fluency doesn’t just mean oral exams for everyone all the time. It means more seminar discussion, talking to your students, and getting them talking to each other, as well as verbal checks in class, structured peer explanation, and feedback the student has to respond to rather than simply receive, much of it in teaching rather than assessment. It works for the same reason conversation practice works in language learning: it creates a manageable amount of disequilibrium, the discomfort of finding that what you know does not quite cover what you are being asked to do, and forces the learner to repair, reformulate and adjust in real time. Using the knowledge like this is both how the learning happens and how you see it.
There are two objections I can predict, and I think the malleability of spoken language addresses both somewhat: a conversation adapts, a fixed page cannot. A blagger can prepare around a written assessment’s fixed questions, but in conversation you follow up on exactly the point that felt thin, which they cannot see coming. And marking spoken fluency does risk privileging the confident, prestige-accented speaker over the hesitant multilingual one, except that the same malleability — plus an assessor who can hear past hesitation to focus on the content — lets you find another way in when one phrasing does not land. Speaking is not less accessible than writing, it is differently biased, and conversation lets you work around that bias better in the room.
Students already understand language fluency. They just need to apply it to their other disciplines. Nobody believes that running their French homework through Google Translate makes them a French speaker. The person who cannot order a coffee has not learned Spanish, whatever their essay looked like. And nobody thinks you can call yourself a violinist without spending years actually playing the violin, however much music theory you know. Give them this same fluency frame for their history essay and I expect that most of them understand before you even finish the explanation.
I think where institutions keep going wrong is that they reach for rules first. Rules about AI use are usually unenforceable, they change between modules, they contradict each other across a degree, and they teach compliance rather than judgement. A student who has learned twelve different sets of permissions has not learned much except that the university is anxious and scrambling for a solution. A reason works where a rule cannot, because the student can reapply it themselves in a situation nobody wrote a rule for, and the reason here is entirely in the student’s own interest, which is what makes it worth making as clear as possible.
Using AI well in your field takes deep disciplinary knowledge, and the fluency to put it to work, because working with AI is not so different from working with a person: it says something, and you respond. Catching where it is confidently wrong, telling a plausible answer from a right one, knowing when to check and when not to: none of that is possible without knowing the field yourself. Without it you are not using the tool well. You are simply trusting it.
The question every student can ask themselves is the same one every assessment should be built to answer: if the tool were taken away, what would be left?
A language only ever becomes yours through use. And so must a discipline.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.