RSS Amplifier

Psytech SA's Substack · Aug 7, 2026

Is Your Online Cognitive Ability Test Still Valid?

0
Sign in to vote or save

Psytech SA · Psytech SA's Substack

Not long ago, the biggest integrity concern with online cognitive ability testing was candidates sharing questions in WhatsApp/Reddit groups. While inconvenient, this was still somewhat manageable. However, the landscape has changed significantly, and faster than most practitioners realise.

In 2023, GPT-4 scored below the 20th percentile on a standardised quantitative ability test. By late 2024, OpenAI’s o1 reasoning model scored at the 95th percentile on the same test. That jump happened in under 18 months, and now, roughly one in three candidates now self-report using AI tools during pre-hire assessments (Robie et al., 2026).

The consequences run in two directions: For candidates, inflated scores displace genuinely capable people further down the shortlist and for organisations, the predictive signal that cognitive ability testing is valued for, is quietly degraded. You may be hiring people who perform worse, leave sooner, and deliver less value than their scores suggested.

So what do we do about it?

First, the instinctive response: and why it doesn’t always work

The obvious answer is to detect AI-generated content and tools like DetectGPT have received considerable attention. But for cognitive ability testing specifically, content-based detection fails in three important ways:

  1. It’s model-fragile: a detector trained on GPT-4 output struggles with o1, and every new AI release potentially requires a new detector.

  2. It’s demographically biased: research has shown these tools disproportionately flag writing by non-native English speakers as AI-generated (Liang et al., 2023): a serious concern in the South African context.

  3. Many cognitive ability tests are multiple-choice, meaning there is no text to analyse. The entire cheating interaction happens off-screen and leaves no trace in the assessment content itself.

The paradigm shift we need is to stop asking what was answered, and start asking how the test was taken.

A framework built around how cheating actually happens

Cheating on an assessment doesn’t happen in just one way, and no single intervention can catch all of it. What’s needed is a layered approach that addresses different threat vectors at different stages of the assessment process. Podium’s assessment platform is built around exactly this logic, with four layers working together: Design, Deter, Detect, and Defend.

Design: build security in from the start

The first layer is the assessment design and experience itself. For example, the General Cognitive Ability Test (GCAT) uses Linear On-the-Fly Testing (LOFT), which draws from a vast item bank to generate millions of unique item combinations., meaning practically, every candidate receives a different version of the test. This makes item sharing and answer key harvesting essentially pointless, since there is no stable “answer” to find and circulate. In this way, we can say that security begins before the candidate even opens the test.

Deter: make the monitoring expectation visible

The second layer is surprisingly simple and effective. An Assessment Integrity Message (AIM) is a brief, non-threatening statement shown to candidates before the test begins, and makes clear that their session is being monitored for behavioural integrity. It doesn’t threaten consequences in detail, but removes the rationalisation that “nobody is watching.”

The deterrence literature is clear on this: people respond more to the perceived probability of being caught than to the severity of consequences (Steger et al., 2021). In our field study of real candidate data, introducing the AIM at an unproctored client reduced the most suspicious sessions by two-thirds - all from a text message that costs nothing to deploy.

Detect: watch behaviour, not content

The third layer is where the technology becomes more sophisticated. Podium’s anti-cheating technology includes the Confidence Score (CS), which is a behavioural monitoring metric that runs silently in the background during the assessment, and tracks three signals: a) how often a candidate leaves the test window (blur score), b) whether their response times are unusually irregular compared to their own average pace (time score), and c) how inconsistent their answer rhythm is from question to question (variability score). These signals are combined into an overall Confidence Score from 0 to 100.

The key insight is that the CS doesn’t care which AI tool a candidate used, or how skilled they are at prompting it, but it captures the physical act of switching away from the test, an act that happens regardless of how efficiently the candidate formulates their query or how fast the AI responds. In a controlled experiment with 415 AI-experienced participants, the gap between honest and assisted sessions was large enough to be statistically meaningful, and candidates with higher AI proficiency were no better at evading detection than those with less experience (More on that study in our next post, so remember to subscribe!)

Defend: add a visible layer for high-stakes decisions

The fourth layer is Facial Validation, which in short, is periodic camera snapshots that confirm a single candidate is present and engaged throughout the session. It is important to note that this is validation, not verification: the system is not building a biometric profile or confirming identity against a database; it is simply checking, at intervals, that the right person is there and that no one else is feeding them answers.

Facial Validation addresses the candidate who photographs their screen with a phone and gets answers on a completely separate device, leaving no telemetry trace on the primary browser at all, one of the limitations of looking at Confidence Scores alone. In our field data across more than 107,000 real candidate sessions, clients using Facial Validation showed consistently lower scores on the subscales most amenable to AI assistance, compared to matched unmonitored groups. The effect was explained by what candidates stopped doing once a camera was present, rather than who they were.

Layers that work together

What makes this framework coherent is that each layer covers what the others cannot: Design prevents item compromise, Deter disrupts the rationalisation before testing begins, Detect catches on-device AI use during the session, and Defend addresses off-device cheating that telemetry cannot see.

We believe that no single layer is sufficient on its own, and the research further supports that. In our next post, we go into the evidence in detail: what the controlled experiment showed, what 107,000 field sessions revealed about how many candidates are actually receiving score-changing assistance, and what the findings mean for how you run your assessments.

Psytech SA is the South African distributor of the GCAT (General Cognitive Ability Test), developed by Podium. This post draws on research presented at the SIOPSA 28th Annual Conference, 2026.

No posts

Read the original on psytechsa.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.