It’s alive! It’s alive!... or at least 57% alive, claim fearless Hendrycks and his powerful colleagues in a new preprint: A Definition of AGI. Just in time for Halloween, a Frankenstein’s monster has been released into the streets of AI policy and evaluation.
The authors take a swing at the perennial “what is AGI?” question by reanimating the Cattell-Horn-Carroll (CHC) model of the structure of human intelligence: a century of psychometric research stitched together into ten latent cognitive-ability domains. They claim that, when AI models outperform human standards on test operationalizations of all the CHC ability domains, it can be said that AGI has been achieved. To that end, they subjected GPT-4 and GPT-5 to such tests and ultimately report these models as being (respectively) 27% and 57% of the way to AGI.
There are, however, at least a few reasons to believe these authors have chosen to mount a head that doesn’t quite fit the neck. The first issue is to do with the foundational premise that the CHC model is a veridical representation of the dimensions of AI capabilities. The CHC model has emerged from the factor analysis of interindividual differences in human performance on mental tests. Consequently, if all humans perform at the same level on some potential dimension (i.e., there is no variability), that potential dimension did not make the final cut for the CHC model. Another issue is the relatively sparse engagement with prior work on mapping machine capability structures or with the possibility that AI might require its own native taxonomy. It can also be said that many of the tests used to operationalize CHC abilities in this paper are not validated psychometric instruments for AI but illustrative examples with arbitrary human-performance thresholds. There is also reason to question whether the model behaviors that the authors refer to as “capability contortions” might more appropriately be considered genuine contributions to ability.
We definitely need better and non-moving goalposts than informal definitions of AGI for the modern Prometheus, but this paper is another trick-or-treat: take this definition or chaos. There are many other options available, which do not have many of the methodological issues of this paper and allow for better characterisations of AI profiles that may be highly transformative and disruptive without having all capabilities humans have. For a more detailed analysis, we refer the reader to this specific blog post.
We don’t know how long Frankenstein’s creature would have lived. The novel doesn’t say, but Mary Shelley warned us: “how dangerous is the acquirement of knowledge”. Scared of this, our editorial team went trick-or-treating in the neighborhood. At the end of the evening, we took a sample of the AGI definition questions as a pub quiz: we scored around 50%, slightly worse than GPT5. We were left with a mixture of disengagement with humanity (“All these years I knew I was a robot”) and wild excitement (“I’m only halfway to become a citizen”). We ordered more bottles of wine to celebrate that AGI is… only a few bottlenecks away. Hurray!
Was Frankenstein’s creature’s girlfriend more gracile and well-conceived? This NeurIPS workshop paper proposes a method to operationalise the definition of “general-purpose AI (GPAI) models” from the EU AI Act, defined as those “competently performing a wide range of distinct tasks”. In particular, the method requires evaluating a set of capabilities, instantiated in benchmarks suitable to the considered model’s modality (text, vision, ...) and annotated according to the difficulty level that each benchmark element poses on the considered capability, using scales normalised based on a human population. The paper evaluates different aggregation procedures for setting thresholds for GPAI.
Benchmark profiling can be done through gradient-based methods. This ends up decomposing LLM benchmark performance into ten cognitive abilities, revealing that most benchmarks test multiple skills rather than single capabilities. Surprise! The method computes Ability Impact Scores to show which cognitive abilities actually drive performance on different benchmarks. This is a more complicated method (but complementary) to some prior work introducing ‘benchmark profiles’, although no comparison is included.
Failure is horrific! That’s why one of the main goals of AI Evaluation is to minimise AI failure! This paper is a survey on AI Failure, a fast growing field as AI is used increasingly more. There are three categories (or reasons) for failure: technical, interactional, and ethical. AI Evaluation could focus on these three separately and jointly too!
What’s more horrific than failure? Becoming redundant. Can LLMs automate AI auditing, as a step to automate AI evaluation? This preprint uses a prompted Sonnet 4 with affordances (benchmarks, python, etc.) to perform these audits. Although specialised to finding issues after finetuning, it shows 56.2% detection rate of adversarial fine-tuning at a 1% false positive rate.
Has Dan Hendrycks taken over the world of AI benchmarks? This preprint maps 2019–2025 data and finds a split: LLM production is decentralizing, but benchmark influence is clustering into a few hubs (authority Gini ≈ 0.89), which helps comparability yet risks path dependence, selective visibility, and saturated leaderboards. A simulation suggests the best way to reduce concentration is to increase the rate of new benchmark creation and keep shared reference suites, by funding a broader, auditable portfolio of benchmarks. Answer: Yes, Dan tops the leaderboard of the number of built benchmarks.
Is performance affected when models have memory of previous interactions and personalisation? This preprint recruited 400 ChatGPT users and 400 Gemini users on Prolific and found out that the results of several “field evaluation” scenarios are significantly different from the offline scenarios. This is not totally surprising but calls for more analysis of these personalised scenarios complementing offline evaluations.
Can language models predict the performance of other language models on a new benchmark, by only looking at some textual description of what the benchmark is about, as described in papers? This preprint finds that GPT5+search can predict LLM performance reasonably well, and better than humans. What’s new? The textual description; the rest is just a Draculaesque take on the classical metalearning problem, studied for decades, but only citing one metalearning paper.
Are LLMs good couch potatoes? From the title we can tell that this preprint is really “meta”: they present “a framework for evaluating models on their capacity to—not just play—but evaluate games”, in particular whether they will find the game fun to play, or watch to play! There are many gems in this paper, but require careful reading!
Horror! Yet another version of the ARC-AGI test (ConceptARC)! This one tests specific concepts and asks models and humans to verbalise how they solve the tasks. They find that AI reasoning models do sometimes abstract and reason in humanlike ways, but they often rely on unintended, nonhumanlike shortcuts. So while they can reach human-level accuracy, they lack consistent object-based abstractions and generalizable reasoning that humans use.
This preprint conducts a rigorous economics-like analysis of the impact of GenAI on scientific productivity in social and behavioural sciences, finding positive effects. In particular, they match authors who are likely to be using GenAI (as determined by the text used in their papers, e.g., “delve”) with similar ones, and compare the productivity change pre- and after-ChatGPT, finding a significant increase in publications output and a small but significant improvements in impact factor. Cause or effect?
What has US NIST been doing amidst the evolution of the NIST consortium, the halt of the US administration and several internal re-organisations? Evaluating a Chinese model: DeepSeek. It would be natural and right to evaluate DeepSeek models and compare them with leading models from OpenAI and Anthropic, especially as DeepSeek models are open-weights, as Nature did recently with its technology. However, the report setting is very scary: the comparison between “capabilities of U.S. and adversary AI systems” is a mandate of President Trump, revealing an AI race narrative. The conclusions are well-known: DeepSeek is less capable and more unsafe than the American counterparts, with a treat: censorship on CCP-sensitive narratives.
Another Draculian preprint conducts a large-scale empirical study assessing when and how well LLMs perform zero-shot probabilistic prediction across diverse tabular-data tasks. They identify key predictors of success, such as task difficulty, label distribution, prompt design, (surprise-surprise!) and propose methods to forecast an LLM performance. They find that success is highly task-related, regardless of whether tasks are from the same dataset or different ones. Most of this was known in previous work on instance-level prediction by LLMs.
“Video models are zero-shot learners and reasoners” looks like yet another Frankenstein’s creature merging the titles of other famous papers: “Language Models are Few-Shot Learners” and “Large language models are zero-shot reasoners”. This preprint looks at video models as another kind of generalist problem solver. In their experiments, they give the model (Veo 3) an initial frame and text description to define the task, and then run the model to generate a video that solves it, such as “segmenting objects, detecting edges, editing images, understanding physical properties, recognizing object affordances, simulating tool use, and more”. We would be interested to see where the evaluation of learning and reasoning in video models goes next.
“I am Dracula; and I bid you welcome, Mr. Harker, to my house. Come in, the night air is chill, and you must need to eat and rest”. Are LLMs sycophants like Dracula? This preprint explores this in everyday conversations and several evaluative settings where users provide feedback in distinct ways. It offers valuable insights into which types of rebuttals are most persuasive in prompting LLMs to change their stance (e.g., providing reasoning), and which ones lead to higher accuracy. Avoiding a strict association between a particular type of feedback and a specific context would make it clearer that the findings have broader applicability across numerous real-world interaction settings.
Trick-or-treat? Reward hacking has long been a problem in machine learning: can a model trained with a reward function genuinely solve the problem, or is it using shortcuts to achieve high scores? This preprint produces an efficient measure of reward hacking in Chain-of-Thought LLMs. The key insight is that a model using a shortcut will solve a difficult problem much earlier in its reasoning trace than if it is genuinely thinking through the problem. Truncated Reasoning AUC Evaluation (TRACE) cuts the reasoning trace and forces the model to answer, with expected reward accessible at each step of reasoning. TRACE outperforms and is much more efficient than competing LLM-based assessor models.
Is this a shadow of the werewolf? This preprint argues that model evaluation should be thought of as inference, rather than simple measurement. In particular, they advocate making the statistical model explicit connecting the measured value (e.g., accuracy) with the latent concept to be inferred (e.g., capability). By rediscovering statistical methods in psychology, they show that aggregate performance is an unbiased measure of capability, leading to biased estimates, which they study empirically for the case of model sensitivity to prompt perturbations. Familiar?
Another ghost in the night? This preprint reframes AI evaluation as an evolving “cognitive examination” that moves from recognition to multi-modal reasoning across tasks. The authors’ claim is that static benchmarks saturate and mask true generalization, so progress now depends on living, multi-metric benchmarks and adversarial/interactive tests that probe reasoning, grounding, and higher-level cognition (including abstract and social intelligence).
This blogpost overviews evaluation of models in long-context tasks, arguing that it takes three forms that are progressively more interesting: 1) can the model “see” the whole context? 2) can it understand it? 3) can it use it to do real-world tasks? The author reports a sizable decrease in performance in long-form tasks of level 1 and 2 as the input length increases, even if models with larger context size are released.
More on this… This NeurIPS LLM evaluation workshop paper explores whether simple, procedurally generated “proxy” evaluations targeting precursor skills (persistence, dexterity, and adaptability) can predict AI agents’ performance on complex, long-horizon tasks such as SWE-bench. The study is heavily constrained by its very small sample size, uncertain construct validity, and limited statistical power, drawing intriguing but early conclusions. And R-horizon is yet another exploration about long horizons.
This preprint is also about long-horizon tasks, but it pokes LLMs where it hurts: “you say you can read long papers? fine, read a whole pile of them and tell me who cited whom.” Funniest bit: Even the fancy models sometimes answer like going blank at a public seminar... they rush, mix up author counts, and blurt out “NULL” as if no one would notice.
What we all know and we all dare to say: “AI benchmarking is broken”. Data leakage, cherry-picking, and opaque metrics inflate headline scores and erode trust in progress. The authors propose PeerBench: a (yet another) live, community-governed evaluation platform with secret tests, rolling task renewal, and reputation-based scoring to deliver contamination-resistant, auditable (weighted averaged) measures of AI progress.
Explanation or trick? The authors of this preprint provide a philosophical and mathematical foundation for mechanistic interpretability in AI based on “Explanatory Faithfulness”, an assessment of how well an explanation fits a model. They argue that scientific interpretability practices can be framed within formal explanatory models rooted in mechanism-based thinking.
An AIES paper introduces an AI Occupational Capability Index (AI-OCI) to quantify how closely AI model capabilities align with the tasks that define human occupations. Using descriptions of 19,000 occupational tasks and 338 AI capabilities, the authors find that occupations with higher AI-OCI values show stronger AI-task alignment (potentially show higher exposure to AI), while lower scores indicate areas where current AI systems remain less compatible with human work.
Moby dick Halloween costume? If whales could talk! A team has proposed a method (ShufflEval) that involves translating vocalizations turn by turn, then checking if shuffling makes the English less coherent; the idea being that real translations preserve conversational structure while hallucinations don’t. They report that tests on low-resource languages and invented alien tongues. Featuring species that communicate via neutrino beams or time-reversed speech they show the method works, offering a practical way to evaluate animal translators without disturbing wildlife. We officially nominate them for the Ig Nobel prize!
Was your secondary school maths teacher terrifying? You now know why. Mathematics requires no less than 433 capabilities! This preprint presents ACE (Active learning for Capability Evaluation), a novel adaptive framework that uses LLMs to decompose domains (like mathematics) into semantically meaningful capabilities and generate reference tasks, embedding them in a latent space to approximate overall performance. Applied to mathematics, it identifies 433 capabilities and 11,800 tasks, using active learning with Gaussian Process regression. We understand your fears.
In recent years, a cottage industry has emerged that copy-and-pastes tests designed to measure human capabilities and directly applies them to Large Language Models. Of course, the validity and reliability of these tests is dubious as soon as the subject population changes. In this preprint, the authors find that psychometric tests of sexism, racism, and morality are not ecologically valid, showing little correlation with downstream tasks. The authors caution the blind use of human-designed psychometric tests for LLM evaluation. Familiar again and again and again.
The authors of this preprint develop a framework to test whether LLM-based agents retain internal consistency - i.e., whether their revealed “latent profiles” (preferences, biases) align with how they behave in conversational settings. Agents often produce responses that superficially mimic human participants, but fail to display coherent behaviour tied to their own internal states. Synthetic model personality plays a crucial role in that. Take-away: your LLM might turn into a wolf at midnight!
CLAIRE is a novel evaluation framework for clustering models that uses pair-wise agreement among different clustering algorithms (i.e., “do the models agree whether two instances should be clustered together or not?”) as the basis for constructing a response matrix. They then apply concepts from Item Response Theory (IRT) to clustering evaluation, thereby estimating each model’s “ability” and each instance-pair’s “difficulty” in the clustering agreement task.
Epoch applies IRT at the benchmark level and creates an index out of it. They admit the limitations of a populational approach openly, where the ECI of a model can change in the future: “The model used to produce ECI scores is fit jointly across all data. As new models and new benchmark evaluations are obtained, values may shift slightly even for models whose data have not changed”. An elusive bird, but at least not a Frankenstein’s monster.
The AI productivity index (APEX) introduces a new benchmark including 200 test cases from four domains where human work has high economic value (investment banking, management consulting, law, and primary medical care). The tasks were sourced from human experts, who were asked to outline the most common tasks in their work and define test cases for them — interesting as it goes towards real-world representativeness.
The Center for AI Safety (again Dan Hendrycks) and Scale AI present the Remote Labor Index (RLI), a collection of projects that can be performed remotely with a computer, usually described with a few sentences and some input files (audio, video, text, etc.). AI agents perform quite poorly, with Manus being the state-of-the-art model with 2.5% of the tasks successfully completed.
Multimodal LLMs are grounded to complete real-world, composite household tasks (e.g., setting a table or cleaning up) in simulated 3D home environments, integrating perception, reasoning, and sequential decision-making. The preprint highlights major gaps in spatial grounding, temporal planning, and robustness to visual ambiguity. Their deployment of Unreal Engine 5 and choice of jigsaw and building block tasks might indicate the chaotic nature of typical student households, but even AI has to start somewhere, hasn’t it?
Two existing benchmarks (ManyIFEval and StyleMBPP) are integrated to measure models’ capability of following multiple instructions at once in the domains of free text and basic programming.Surprise, surprise, reasoning models are the best at following multiple instructions at once. The preprint also trains simple assessors (they don’t call it like that) to predict the performance of models in unseen combinations of instructions, and some of them work pretty well, even with small training sample sizes.
Oh, spiders all over the place! Spider 2.0 is the latest evolution of the Spider text-to-SQL benchmark—now built from 632 real-world enterprise workflows rather than curated question–schema pairs. Evaluation goes beyond static accuracy to measure completion rate, execution accuracy, and coherence, rewarding flexible, goal-directed reasoning rather than rote query matching. Still, it would benefit from more granular metrics, clearer sampling representativeness, and robustness checks on evaluation scripts to ensure stability and generalisability.
Is prompt engineering a walking dead? A team from UC Berkeley and Virginia Tech introduced StyleBench, a benchmark evaluating whether different prompting strategies work better for different tasks and model sizes. Across 15 open-source language models tested on math, logic, and puzzle challenges, they found no universal winner; straightforward step-by-step prompting excels at math, shorter methods save time on simpler questions, and complex search strategies only work well with large models.
FaithCoT-Bench is not a psychotic killer but a new benchmark for detecting unfaithful Chain-of-Thought (CoT) reasoning in large language models at the instance level. The authors create FINE-CoT, an expert-annotated dataset of over 1,000 CoT trajectories labelled for faithfulness across four domains, and systematically evaluate 11 detection methods, finding that LLM-as-judge approaches perform best while logit-based methods struggle significantly.
Can language models learn new words from just a few examples as easily as children do? The BabyLM Challenge seeks to evaluate this capability. Experiments show that language models generally struggle with rare words. Moreover, while model performance is similar for frequent words, differences between models become much more pronounced when dealing with rare ones. Baby-level AI is not here yet.
ConsistencyAI is a new benchmark designed to test whether LLMs deliver the same factual responses when asked the same questions by different personas. The upshot: while many models are highly consistent overall, performance varies significantly by topic (for example the job-market topic was notably low) and by model, underscoring that factual consistency depends on both the subject matter and the provider.
PodEval seeks to evaluate AI-generated podcast outputs. By evaluating textual content, speech generation, and audio quality, the authors aim to give a holistic analysis of podcast generation. They deploy a range of both objective and subjective metrics to assess each, and open-source their work. They construct a dataset of real world podcasts against which to compare other AI generated podcasts from a variety of systems.
METR routinely investigated the degree to which agents demonstrate problematic (or at least annoying) behaviours such as sandbagging or reward hacking when being evaluated on METR’s own HCAST and RE-Bench. They now released MALT, a dataset of 10K cases to benchmark automatic monitors for detecting such behaviour. Any sufficiently complex evaluation needs its own evaluations! And to test whether they know they are being tested.
Great events commemorating 75 years of Alan Turing’s paper ‘Computing Machinery and Intelligence’. We were at two of them. One at the Royal Society, which briefly covered Turing’s legacy and future, to give way for more discussions about the ethics of AI and how doomed we are in the hands of the Big Tech. Another one was the Next Turing Tests conference, organised by E-Lab at King’s College Cambridge, where Turing was both student and fellow, and the Leverhulme Centre of the Future of Intelligence (CFI) at the University of Cambridge, with a more forward-looking focus on what new tests should be measuring in the future.
QUIZ: Can you tell which of these people participated in both, one or none of them? Dermot Turing, Peter Gabriel, will.i.am, Gary Marcus, Kraftwerk, Marilyn Manson and Geoffrey Hinton. If you got this right, there’s a treat for you!
“To put Americans first”. Scary? A bipartisan US senators team proposes US federal legislation to evaluate AI systems and collect data on adverse AI incidents. The program would ensure real transparency by requiring developers of advanced AI systems to submit product information to the DOE before deploying their new technology.
SB 53 passed in California, requiring developers to publish model cards summarising the risk assessment from the model and the role of 3rd party evaluators.
There is a session on AI evaluation in any conference these days. The PyTorch conference, October 2025) had one, with provocative topics about whether benchmarks measure something to do with intelligence, evaluation in the wild, and other cool topics: https://misummit25.sched.com/
Approaching deadlines for the International Programme on AI Evaluation: Capabilities and Safety, looking for 40 exceptional students and professionals from around the world for a 150-hour hybrid course that blends lectures, hands-on labs, and a capstone project week in Valencia, EU.
And we close our Halloween special celebrating that our AI Evaluation newsletter passed 1,000 subscribers and more than 2,000 followers! We’re halfway to fame. Hurray!
News to share? Feel free to reach out to ai.evaluation.newsletter@gmail.com
Getting the digest: Once a month if you join.
Contributors to this month’s digest: Zack Tidler, Behzad Mehrbakhsh, Jose H. Orallo, Kozzy Voudouris, Lorenzo Pacchiardi, Peter Romero, Fernando Martínez-Plumed, Wout Schellaert, Daniel Romero, Marko Tesic, Irene Testini, Manuel Cebrián, Victoria Carro, Ben Slater, Joseph Castellano.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.