The idea of digital exams feels like a no-brainer. Anthing described as “digital”, “modern” and “efficient” is often perceibved as inherently good. Handing students paper and asking them to write on it with anything as primitive as a pen might, to some, seem ridiculously out of step with the world we live in.
In an interview with Tes, AQA chief exec Colin Hughes argues that the system lacks the necessary urgency. AQA, he says, would be ready to deliver digital exams “tomorrow morning” if asked, even if the practical transition would need to be phased in over several years. Hughes claims students are “slightly baffled” that exams remain on paper, resistance to onscreen assessment “just doesn’t stack up” and, on the overwhelming majority of fronts, he sees “almost no downside”.
The strongest case for digital exams is that the current system may prevent some students from demonstrating what they know as accurately as they could. For some students, typing may remove a barrier that has nothing to do with the subject being assessed. In such cases, could digital assessment improve validity by reducing the extent to which an exam measures handwriting fluency alongside subject knowledge?
The accessibility case extends beyond typing. Digital platforms can support enlarged text, alternative colour settings, screen readers, adjustable layouts and other assistive features more easily than paper. Properly designed, onscreen assessment could offer some students a more direct route to demonstrating what they know, while reducing the need for cumbersome or inconsistent adjustments.
Hughes suggests that controlled assessment will become more important in the age of LLMs. Coursework completed beyond direct supervision is increasingly difficult to authenticate, and teachers are left to determine how much of a final submission reflects the student’s own thinking. Timed, secure assessment offers a clearer basis for judgement and may become increasingly necessary if qualifications are to retain public confidence.
Young people are often described as digital natives, yet familiarity with social media tells us very little about whether they can type efficiently, manage files, use spreadsheets, navigate complex software or work productively on screen. Access to these skills is uneven, and students from more advantaged backgrounds may gain them informally through regular access to laptops and other devices. A more deliberate approach to digital fluency could therefore be defended as an attempt to reduce inequality rather than deepen it.
There are substantial administrative benefits as well. Fully digital delivery could reduce printing, transport, scanning and physical storage, while allowing papers to be distributed instantly and responses to be captured securely at the point of submission. It could possibly provide richer information about how students interact with questions, allow problems to be identified more quickly and give awarding organisations tighter oversight of the assessment process.
Questions could include dynamic sources, audio, video, simulations or interactive data. Some assessments could become adaptive, allowing the level or sequence of questions to respond to a student’s performance. Accessibility features could be built into the platform rather than added afterwards. From this perspective, onscreen assessment isn’t merely a cheaper way of delivering the same exam but a possible route to assessments that are more flexible, inclusive and responsive.
The central question shouldn’t be whether digital exams are possible or popular but whether they allow us to make more valid, reliable and fair judgements about what students know and can do. But on that question, boosterism runs well ahead of the evidence.
The most striking part of Hughes’s argument is AQA’s finding that students using typed short-answer boxes tend to give better answers, with particular benefits for lower-performing students. If typing allows students to reveal knowledge that handwriting would otherwise obscure, this is potentially important. An assessment shouldn’t reward the physical ease with which a student can form letters when handwriting is irrelevant to the subject being assessed.
One difficulty is that higher scores don’t, by themselves, tell us that an assessment has improved. All this tells us is that changing the response medium has changed performance. Whether that change is desirable depends on why it occurred. Perhaps students type better answers because handwriting imposes an irrelevant burden. Perhaps editing is easier, composition is faster and ideas can be revised with less friction. Equally, students may perform better because they’re more fluent with keyboards, more familiar with digital interfaces or more adept at navigating the particular design of the response box. The phrase “better answers” conceals all of this. Better in what sense? Longer, clearer, more accurate, more developed or simply more highly rewarded? Unless AQA can show that typing improves the validity of the inference rather than merely raising scores, the finding proves less than Hughes suggests.
As yet, independent research has not supported what Hughes would like to be true. A 2024 Cambridge review found that typed and handwritten responses can differ in scores, length, language, planning and revision. The direction of those differences isn’t consistent, and comparability can’t simply be assumed where both modes are treated as equivalent. Ofqual’s 2025 review reached a similarly cautious conclusion. Students are often positive about onscreen assessment, but working on screen can create greater cognitive demands, particularly for reading. The review found no clear and consistent overall pattern in performance because mode effects vary by subject, task and student characteristics. It also noted that much of the available research doesn’t come from live, high-stakes English examinations.
A separate study reported in 2025 from Broc et al found that students typed substantially more and achieved higher scores than when writing by hand. That may indicate that typing removes a barrier, but it also raises an obvious question about whether the two modes are assessing the same thing. The study’s recommendation that schools may need to teach touch typing also exposes the potential curricular consequences of changing the medium of testing.
AQA’s experiments may therefore be promising, but ‘students tend to give better answers’ is a long way from being conclusive. However, even if it wasn’t, students whose handwriting places them at a significant disadvantage can already apply for access arrangements that allow them to type their exams. The existence of a genuine barrier for some students is therefore an argument for making appropriate accommodations easier and more consistent, not for requiring all students to use a keyboard.
Assessment researchers have a name for the problem: construct-irrelevant variance.1 This occurs when differences in scores are caused by something other than the knowledge or capability the assessment is intended to measure.
In a digital history exam, for instance, performance might vary partly because students differ in typing speed, familiarity with the platform, ability to navigate between windows or confidence reading extended material on screen. The result would still reflect historical knowledge, but it would also reflect competence with the medium. This doesn’t mean any difference between paper and screen invalidates an assessment. Every exam places incidental demands on students. Handwriting, reading speed, working memory and familiarity with examination conventions already affect performance to some extent.
The question is whether these demands are necessary to the thing being assessed and whether they affect some groups more than others. If keyboard fluency changes a student’s history grade despite keyboard fluency forming no part of the intended history construct, the resulting variation is irrelevant to what the examination claims to measure. Short-answer text boxes may reduce construct-irrelevant difficulty for students whose handwriting is slow or effortful but they could equally create construct-irrelevant difficulty for students who think well but type slowly, navigate clumsily or struggle to review writing on screen.
Digital exams shouldn’t be adopted wholesale regardless of the subject being assessed. A format that helps with a short science response or a maths paper may hinder an English literature essay, a geography source question. The case has to be made subject by subject and task by task.
Hughes says that “being able to use keyboards is a really important skill that we should be supporting students with from a much earlier age”. There’s nothing unreasonable about teaching students to type, but typing is just one narrow aspect of digital competence. One student may be able to type quickly while being unable to evaluate a source, verify an output, protect personal information or adapt to unfamiliar software. Another student may type slowly but exercise excellent judgement and learn new systems with ease. Reducing digital fluency to proficiency with a QWERTY keyboard mistakes familiarity with one current interface for preparation for the future.
Keyboards may remain common for decades or they may go the way of the mouse or the floppy disk. The history of computing offers little reason to assume that today’s dominant input device will retain its position indefinitely. Speech recognition is improving rapidly, stylus input allows handwriting to be captured and converted digitally, predictive systems increasingly complete or transform what users produce, and emerging interfaces combine voice, touch, gesture and visual input. The likeliest future isn’t one in which the keyboard is simply replaced by another universal device. It’s one in which people move between several modes according to the task. They may speak a first draft, type a precise correction, annotate with a stylus, manipulate information by touch and rely on software to translate between formats. Building the future of national assessment around keyboard fluency risks confusing preparation for one current interface with preparation for a digital world defined by constant changes of interface.
Digital fluency ought to involve judgement, adaptation, verification and purposeful use. Keyboard speed is narrower and more contingent. It may be worth teaching, but making it consequential in unrelated examinations is another matter entirely.
Once keyboard speed influences examination performance, schools will respond rationally. They may introduce touch-typing schemes, digital readiness audits and extra rehearsal of the examination platform. Departments may be told to build screen navigation, digital annotation and onscreen proofreading into curriculum plans. Assessment creates incentives which inevitably backwash into the curriculum. Lessons spent preparing students to operate the assessment platform are lessons not spent on teaching subject content.
Hughes might reply that this is precisely what schools ought to be doing. If students can’t use keyboards or screens properly, he says, the system has failed them. There’s some force in that argument, but it still confuses two different questions. Digital fluency may be worth teaching, but it doesn’t follow that a student’s GCSE grades should depend on it. First aid is useful. Financial literacy is useful. Cooking is useful. Their value doesn’t make them legitimate hidden components of every GCSE.
A digital exam that requires schools to teach typing creates a new condition of success and then instructed schools to compensate for the inequality it introduces. And, of course, the cost won’t fall evenly. Well-resourced schools can provide regular access to devices and sustained practice. Others may offer a few sessions in a computer room and hope for the best. Students with regular access to laptops may arrive with a substantial advantage in familiarity with the medium.
That isn’t merely an implementation problem. It is part of the validity problem; familiarity with the assessment hardware will influence the grades awarded.
Hughes sees the teaching of digital fluency as a means of promoting equality. Although it’s right that leaving existing differences untouched advantages students with better access to technology, universal digital exams would expose these differences to high-stakes assessment.
Fair digital assessment would require much more than placing a computer on every desk. Students would need sufficiently comparable keyboards, screen sizes, software response times, accessibility settings and opportunities to practise. Schools would need dependable networks, replacement devices, charging capacity and technical support.
This would also create an enormous financial problem. Ofqual’s proposals rule out students using their own devices, which means schools would have to provide sufficiently consistent devices. A secondary school with a cohort of 200 Year 11 studentswould probably need at least 220 exam-ready devices once contingency capacity was included. Even at a relatively modest £350 to £500 per machine, that represents £77,000 to £110,000 before secure storage, charging, network upgrades, software, technical support and modifications to examination rooms are considered.
A cautious estimate would therefore place the initial cost for a typical secondary school somewhere between £120,000 and £230,000. Across England’s state secondary sector, the capital bill could easily fall between £400 million and £800 million, with colleges, special schools, independent schools and other examination centres pushing the total higher. Devices would then need replacing every few years, while networks, licences, security and technical staffing would create continuing costs. Existing school devices would reduce the additional expenditure in some centres, but ordinary classroom provision isn’t necessarily sufficient for simultaneous, secure, high-stakes examinations. A school may own hundreds of laptops without possessing hundreds of identical, charged, reliable and examination-ready devices available at the same time.
The precise figure will vary enormously, but the order of magnitude is clear. Fully digital exams wouldn’t be a minor technical adjustment. They’d require a national infrastructure programme costing hundreds of millions of pounds, much of it paid for by schools unless government or awarding organisations provide substantial additional funding. Digital exams would add another source of inequality which schools were responsible for addressing.
Paper’s advantages are easy to miss, largely because we’ve stopped noticing them. Students can underline, annotate, sketch, cross-reference, turn pages and spread materials across a desk without first learning an interface. They can see where they are in a paper, move backwards and forwards and use space in idiosyncratic ways.
Of course, paper isn’t cognitively effortless — handwriting and physical navigation impose demands of their own — but these demands are familiar, stable and already built into existing examination practice.
Digital systems can reproduce many of paper’s functions, but reproduction isn’t equivalence. Scrolling changes how information is located, split screens reduce usable space and moving between a source and an answer box requires navigation. Digital annotation tools must be learned, while reviewing several pages of writing on a screen isn’t the same as seeing them laid out on paper. None of these difficulties is insurmountable, but each consumes attention. That may make little difference in a short-answer test, but it becomes more significant when students are comparing sources, interpreting diagrams or sustaining a complex argument.
Ofqual’s research makes precisely this point. Mode effects can arise during the presentation of material, the thinking required, the production of a response and the process of reviewing it. However well designed they are, digital exams introduce mode effects, technical fragility, inequality and cognitive demands that have nothing to do with the subject being assessed.
Hughes is also right that artificial intelligence creates serious problems for unsupervised coursework. If students can generate or substantially improve work using AI, it becomes harder to know whose knowledge and capabilities are being assessed. Timed, controlled conditions may therefore be fairer than work produced beyond direct supervision. But this is an argument for controlled assessment, not digital exams.
A paper examination is every bit as timed and controlled as an onscreen one. Moving an exam onto a screen doesn’t solve the problem of AI-assisted coursework because the important feature is secure supervision rather than the response medium.
Hughes still believes coursework remains “very much the best way” to assess some aspects of history and English literature, but this now feels more like wishful thinking than a workable position. Supervised drafting, staged submission, oral defence and greater reliance on work completed in class may make malpractice harder, yet each would add substantial complexity and workload without eliminating the central problem: once students can use AI between stages, it becomes increasingly difficult to know which parts of the final submission are genuinely theirs.
Coursework may theoretically assess some capabilities better than timed exams, but its validity now depends on being able to authenticate authorship. Unless a system can do that with reasonable confidence and without creating an administrative circus, calling coursework “the best way” is no longer enough. It may be the best method in theory while becoming progressively less defensible in practice. Replacing a valid form of assessment with a less suitable one because the latter is easier to secure would be an administrative solution to an educational problem.
The strongest case for digital exams is accessibility. Some students will find typing easier than handwriting, while screen readers, adjustable text and assistive software may remove barriers that paper preserves. However, this is an argument for flexibility rather than universal compulsion. Where technology allows a student to demonstrate knowledge more accurately, it should be available. Existing arrangements for word processors and reasonable adjustments can be improved, broadened and made more consistent.
Universal digitisation would solve some problems while creating others. Some students may find screens tiring or difficult to navigate, while others may plan and review their thinking more effectively on paper. There is no reason to assume that the format best suited to one group should become the compulsory norm for everyone.
An accessible assessment system should respond to meaningful differences. It shouldn’t impose one medium and treat any resulting disadvantage as a temporary inconvenience.
Mixed-mode assessment would, of course, create its own comparability difficulties. AQA acknowledges this in its response to Ofqual, which is another reason the transition can’t plausibly be described as having “almost no downside”.
The strangest feature of the debate is how little ambition it displays. Digital technology could allow us to rethink how complex performances are assessed, yet the flagship proposal appears to be placing a traditional examination paper on a screen and asking students to type into boxes.
Comparative judgement offers a more interesting example of what digitisation might make possible. Instead of asking examiners to assign marks by working through increasingly elaborate criteria, judges compare pairs of responses and decide which is better. Repeated comparisons can then be used to place scripts on a quality scale.
Comparative judgement isn’t magic. Its validity depends on the quality of the task, the expertise of the judges and what they attend to when making decisions. It may suit some subjects and performances better than others, while translating a rank order into defensible grades raises further questions.
AQA’s 2023 account of the pros and cons of comparative judgement feels out of step with the progress being made by outfits like No More Marking. Fixating on questions about scaling, validity, feedback and the number of comparisons required feels very behind the curve.
Comparative Judgement addresses a genuine assessment problem. We find relative judgements of complex work easier and more consistent than assigning independent scores against lengthy criteria. Digital systems can distribute scripts, collect many judgements and combine the results efficiently, and research has found comparative judgement capable of producing reliable and valid assessments of complex performance in suitable contexts. Why, then, is placing old exams on screens treated as urgent while more ambitious attempts to improve the quality of judgement remain peripheral?
Comparative judgement isn’t necessarily the answer (even though it’s undoubtedly cheaper, easier and more valid and reliable than traditional assessment) but its approach uses technology to improve the process of assessment, rather than merely changing the mechanism through which students enter text. Where digitisation amounts only to typing the same responses into the same boxes, it’s stationery reform rather than assessment reform.
The strongest argument for digital exams is administrative efficiency. Digital delivery may reduce printing, transport, scanning and storage costs. It may allow richer data, faster processing and tighter control over the examination process.
These are real benefits, but have largely already been secured. Students write on paper, their scripts are scanned and examiners mark them onscreen. Responses can be distributed electronically, marking can be monitored centrally and scripts can be stored and accessed digitally.
This hybrid system keeps the simplest and most familiar technology in front of students while digitising the administrative process behind the scenes. It captures most of the administrative prize without requiring schools to build and maintain a national digital examination infrastructure.
Of course, scanning doesn’t capture every possible advantage. Fully digital assessment can support adaptive pathways, new item types, instant capture of responses and integrated accessibility features. Where these improve validity or access, they deserve serious consideration. But those benefits have to be demonstrated rather than inferred from bandying about the word “digital”. Conventional questions typed into conventional answer boxes offer a meagre educational return.
Whether or not students find paper exams “baffling” (I’d like to see the research on which Hughes bases this claim!) or whether AQA is able deliver digital exams tomorrow is besides the point. (Schools definitely are not ready.) The case for digital exams stands or falls on whether they produce a more valid, reliable and fair assessment of a particular body of knowledge.
Where digital delivery improves assessment, we should be using it. Where adaptive testing, simulations, comparative judgement or assistive technology offer a genuine advance, exam boards should explore them seriously. But, where digitisation merely transfers costs and introduces new sources of irrelevant difficulty, we should remain sceptical.
To think that there “almost no downside” to moving to digital exams seems spectacularly ill judged, especially when the devices, training, infrastructure, curriculum time and new sources of inequality will appear on someone else’s balance sheet. Exam boards may gain efficiency, but schools and students will be burdened with the costs. The question isn’t whether exams can be put on screens — plainly they can — the question is whether doing so improves the reliability and validity of assessment. Until that case is made subject by subject, task by task and group by group, digital exams remain an expensive way of confusing novelty with progress.
In assessment theory, the construct is the knowledge, skill or capability an assessment is intended to measure. Construct-irrelevant variance occurs when scores are also affected by factors unrelated to that construct. For example, if a history exam is intended to assess historical knowledge and reasoning but results are influenced by typing speed or familiarity with a digital interface, those differences introduce construct-irrelevant variance. Along with construct under-representation, where an assessment fails to capture important aspects of what it claims to measure, it’s considered one of the principal threats to assessment validity. See Messick, S. (1994), ‘The interplay of evidence and consequences in the validation of performance assessments’
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.