In our previous post, we introduced our four-layer framework (Design, Deter, Detect, Defend) that underpins how Podium’s assessment integrity technology approaches the challenge of AI-assisted cheating in remote cognitive ability testing. We had hinted at two studies, so now, let’s get into what we actually found.
Study 1: Can we catch an AI-assisted test-taker?
The first question was fundamental: does behavioural monitoring actually work? The Confidence Score tracks how candidates interact with the test: how often they switch away from the assessment window, whether their response times are irregular, and how inconsistent their answer rhythm is from question to question. While it sounds sensible in theory, did tracking these behaviours actually help discriminate between candidates who are working independently and those who are getting help?
To find out, we ran a controlled experiment with 415 participants recruited through Prolific, an established online research platform. Everyone was pre-selected for AI tool experience to ensure that these were regular LLM users, not novices. All participants took a 36-item version of the General Cognitive Ability Test (GCAT) under high-stakes framing, with a real cash bonus on offer for the top 5% of scorers to make the motivation genuine.
Participants were randomly assigned to one of two conditions: the Integrity group (n = 122) were asked to complete the test independently, told that AI monitoring was active, and informed that the bonus would be forfeited if assistance was detected. The Assisted group (n = 293) had all restrictions explicitly removed, were openly encouraged to use any AI tool available, and coached on how to do so effectively within the full-screen environment.
After the test, every participant completed a brief anonymous survey asking how many items they had sought help with, and it is that self-report that became our ground truth.
What we found
The Confidence Score (CS) clearly distinguished between the two groups: Honest candidates averaged a CS of 85 out of 100, while assisted candidates averaged 49. That 36-point gap represents a very large effect, highlighting nearly twice the typical difference between university graduates and those with secondary schooling only on a cognitive ability test.
In terms of practical accuracy: 93% of honest candidates were correctly cleared, and 91% of flagged sessions were genuinely assisted. Hence concluding: if you flag someone, you are almost certainly right, and if you clear someone, you can be confident in that too.
One finding stood out as particularly important for practitioners: among the assisted group, AI proficiency had no relationship with CS scores - the correlation was essentially zero (r = .015). Candidates who were more experienced and skilled at using AI were no better at evading detection than those who were less experienced. The reason is straightforward: what the CS measures is the physical act of switching away from the test window. You can write the most sophisticated prompt in ChatGPT, Claude, or an AI app of your choice in seconds, but you still have to leave the assessment window to do it.
However, that left this study open to one clear limitation: Participants who used a second device to photograph the screen with a phone and getting answers from a multimodal AI on a separate device,were invisible to the CS. Their primary browser looked completely normal, because it was. This is not a flaw in the algorithm, but a structural boundary of session-level monitoring, and it is precisely what Facial Validation is designed to address.
Study 2: Does it hold in the real world?
A controlled experiment tells us whether something can work. A field study tells us whether it does work in real organisations, with real candidates, under real hiring pressures.
Study 2 drew on over 107,000 GCAT sessions across 34 client organisations, 6 countries, and 6 years of data (2019–2025). Because this is observational data rather than a randomised controlled trial, we took careful steps to make the comparisons fair: 1:1 demographic matching on gender, age, education, and country and 1,000 Monte Carlo iterations, meaning we ran the entire analysis 1,000 times with different random draws to confirm the results weren’t a statistical fluke.
We also used Abstract Reasoning as a within-test control throughout. Because Abstract Reasoning requires figural pattern recognition rather than language or numerical computation, it is substantially harder to assist with using a text-based AI. So when our integrity interventions reduced Verbal and Numerical scores but left Abstract flat, that is the fingerprint of cheating being removed rather than a different cohort or seasonal effect.
Does a simple message make a difference?
At an unproctored client (no camera, just the Assessment Integrity Message introduced before the test), Verbal scores dropped from an average of 10.8 to 10.6 out of 20 after the message was introduced, and Numerical from 10.1 to 9.8. Abstract barely moved. The most suspicious sessions, ie. those with a Confidence Score below 40, dropped from 3.3% to 1.2% of all sessions. A two-thirds reduction in assumed cheating cases, just from a text message that costs nothing to deploy.
Does a camera change scores?
After matching 7,358 pairs of monitored and unmonitored candidates at two public sector clients, unmonitored candidates averaged 0.7 points higher on Verbal Reasoning and 1.5 points higher on Numerical Reasoning than their demographically identical monitored counterparts. The gap on Abstract was smaller, as predicted. This result held across 100% of 1,000 Monte Carlo iterations.
Importantly, the Confidence Score showed almost no difference between the monitored and unmonitored groups. Candidates with a camera present behaved identically in terms of their session telemetry: same blur patterns and timing profiles. But what changed is that they stopped consulting external tools, confirming that the camera removes the cheating, not how people take the test.
So… how many candidates are actually cheating on cognitive assessments?
This is usually the question practitioners most want answered. Survey data says roughly one in three candidates self-report using AI during pre-hire assessments, but that spans a wide spectrum, from a passing query that doesn’t affect the score to systematic answer generation across every item.
Using the field data to estimate the proportion receiving assistance that materially changes their score, we arrived at a figure of 5–9% of unmonitored candidates. This is consistent with a prior independent field study that estimated 6–8%. In a pool of 1,000 candidates, that is 50–90 artificially inflated scores entering your shortlist, and at competitive selection ratios, those scores are displacing genuinely capable candidates.
Although our numbers are smaller than the self-report surveys suggest, it is large enough to matter particularly in high-stakes selection contexts where cognitive ability is a meaningful predictor of performance and the consequences of a poor hire are significant.
What this means for your practice
Podium’s anti-cheating technology provides the Confidence Score and Assessment Integrity Message as part of its standard platform. Facial Validation is available for clients where the stakes warrant it, and the evidence suggests the effects are 2–3 times larger than the message alone.
The practical starting point is simple: deploy the Assessment Integrity Message before every test. It is zero cost, takes no additional administration, and the field data shows it works. For higher-stakes decisions, add the camera. And treat every CS flag as the beginning of a conversation, rather than an automated gate. Human review is non-negotiable, both psychometrically and under HPCSA Booklet 20.
Our third and final post in this series looks at the question we know South African practitioners care about deeply: is any of this fair? We go into the equity and adverse impact findings, and what they mean for practice under the Employment Equity Act.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.