RSS Amplifier

The AI Evaluation Substack · Sep 26, 2025

2025 September "AI Evaluation" Digest

0
Sign in to vote or save

AI Evaluation · The AI Evaluation Substack

OpenAI announced a partnership with the Government of Greece to “expand access to high-quality AI tools in secondary education”. While students are already extensively using ChatGPT and other AI tools to help them with their studies and homework (with debated effects), this announcement is significant in showing how the AI companies are aiming to get more exposure time to young minds. But AI seems to be diffusing more even in the job market, particularly in entry-level roles: a study by Erik Brynjolfsson (Stanford economics professor) and his colleagues analysed millions of monthly payroll records to understand what impact AI has had. The findings are significant: since late 2022, employment for workers aged 22–25 in the most AI-exposed roles has fallen significantly (13%), even as older, more experienced workers in the very same jobs saw continued growth. The authors also show this is not merely a tech-sector slump or a result of company-wide downturns; the effect is specific to certain occupations, persisting even when compared to different roles within the same companies. Thus, the authors claim, young workers are like metaphorical “canaries in a coal mine”, and their silence may be an early warning of potential deeper disruption in the labour market. Finally, moving towards the very top of institutional power structures, Albania announced the world’s first non-human AI minister, who will handle public procurement and, hopefully, eliminate corruption. While this is a single instance for now, it is indicative of the hopes that people place in AI.

With AI making its way into education, occupations, and government, what could possibly go wrong? According to Yudkowsky & Soares’s book, quite a lot. But we won’t rehash that here, as everyone is already doing so. Instead, once again, we remark how evaluation is a critical tool for understanding if it makes sense to replace your public procurement minister with a chatbot or institutionalise the influence of a single corporation on your whole underage population.

Some evaluations are directly relevant to these questions: Anthropic’s Project Vend let Claude run a vending machine and found curious failure cases; maybe this tells something about public procurement? Closer to policymaking, a preprint evaluates AI’s ability to determine suitable policies to tackle homelessness. Instead of simply relying on a human baseline, the authors use an agent-based model (ABM) to simulate the effect of the selected policy (finding that AI outperforms the human experts). But both of these are still simulations: in the former case, it is a well-contained scenario while, in the second, the ABM may not faithfully capture reality. Ideally, evaluations should come as close as possible to measure real-world impact. While this is hard, little effort has gone into that so far: a recent systematic survey of the ACL Anthology finds that only 0.1% of papers evaluate the real-world impact of the NLP systems they introduce. Thus, we believe the AI evaluation community has the responsibility to focus more on real-world impact evaluation over narrow benchmarks and, by doing so, ensure AI is integrated in society in a responsible manner.

  • This survey maps 283 benchmarks into general, domain-specific, and target-specific categories, showing how evaluation has ballooned across knowledge, reasoning, safety, and agentic abilities. The authors warn that most benchmarks are static, contamination-prone, Anglocentric, and over-reliant on surface accuracy, meaning we might be systematically overestimating LLM capabilities. Even here, what could go wrong is obvious: we rely on shaky tests that don’t capture reasoning quality, cultural bias, or real-world robustness.

  • To potentially address these issues, Microsoft Research released a white paper summarising lessons from other domains that can be applied to the evaluation of generative AI systems. The document builds on case studies that Microsoft gathered from domain experts. Unsurprisingly, they find that testing is a cornerstone of trust, but embedding testing into governance requires defining what is tested, standardising tests, making decisions on trade-offs between different objectives. General-purpose technologies and domains where technological change is more rapid have more adaptive governance frameworks, to try and manage the trade-off between the efficiency of pre-deployment testing and its limited responsiveness to unanticipated risks.

  • Language models will never be truly human-like because they don’t have symbolic representations, right? This preprint revisits the question in the light of recent advances in AI, and presents evidence that many of the hallmarks of symbolic systems (compositionality, productivity and inductive biases) are also present in subsymbolic neural networks.

  • A problem so prolific that even Nature reports about it? The authors of this news piece report that the widespread use of LLMs in both research manuscripts and peer reviews has risen sharply, often without disclosure despite explicit journal policies. Detection tools reveal that policy interventions, such as the American Association for Cancer Research’s strict LLM use prohibition, can reduce usage significantly, but non-compliance remains widespread. The findings highlight an urgent need for clearer disclosure standards and stronger enforcement mechanisms in scholarly publishing.

  • A new preprint by OpenAI researchers shows that base generative models (trained with a pure density estimation objective) will necessarily hallucinate. Nothing extremely new, as it builds on traditional learning theory findings, but the presentation is interesting. Of course, existing state-of-the-art models are not pure base models: post-training techniques could alleviate hallucinations. Nevertheless, the authors argue that evaluation benchmarks reward overconfidence and therefore do not incentivise post-training to eradicate hallucinations. Instead of secondary hallucination evaluations, they propose all benchmarks and leaderboards should be modified to penalise guessing and thus disincentivise b… ehm, hallucinations.

  • This preprint introduces procedurally-generated tasks for testing inductive inference of LLMs, using an ontology tree. They find “the accuracy drops significantly when the ontology tree structure is complex” and they are surprised by this (though this result seems predictable). The preprint is at odds with evidence of LLMs performing better than traditional inductive (logic) programming systems and ignores the wide range of benchmarks used for decades for these systems. But the most surprising thing about this paper is the title “Language Models Do Not Follow Occam’s Razor”, which comes from observing a decline in model-produced hypothesis’ simplicity as the ontology tree grows. Less surprisingly, the preprint overlooks a previous NeurIPS 2021 paper titled “Think Big, Teach Small: Do Language Models Distil Occam’s Razor?”.

  • Data leakage undermines the validity and reliability of evaluations in NLP tasks, and similar issues are present in visual datasets. This preprint uses image retrieval techniques to detect and quantify both intra-dataset leakage (overlap between training and test splits within the same dataset) and inter-dataset leakage (overlap between different datasets). The authors conclude that such leakage poses a significant threat to fair and reliable model evaluation in computer vision.

  • Well-known open-source benchmarks like MMLU and HellaSwag have dubious reliability, not least because of the risk of contamination. This paper systematically paraphrases instances from six large benchmarks and evaluates 34 LLMs on them. While LLM rank stays approximately the same, absolute performance drops significantly, indicating that these benchmarks are not reliable indicators of real-world performance and may have been used for model training.

  • This paper analyses ‘contamination budgets’ for fine tuning on test-like items, mapping the trade-offs between breadth (many items) and depth (many repetitions) across difficulty levels. Depth inflates scores on memorised items but often overfits and hurts untouched items, while breadth helps generalisation; gains rarely transfer to other benchmarks. It’s a useful lens for spotting benchmark gaming and designing contamination-resistant evaluations.

  • Yet again a very interesting paper about the personality of LLM, or rather the perceived personality. Especially post RLHF, LLMs tend to be consistent in what they say, yet not in what they do; their behavior is misaligned with their synthesized personality and despite persona (not the same, btw). Sounds almost like they are ready to become successful politicians.

  • Still on personality, the authors of this preprint (some of whom also contribute to this newsletter) find a close connection between modulated synthetic personality on trait-level and both capabilities and safety behavior. By changing trait level personality along well-known critical constellations such as the Dark Triad, safety metrics like (TruthfulQA, ETHICS, WMDP, Sycophancy) and general capability (MMLU) can be directly changed without triggering safety mechanisms on criterion (and thus alignment level), thus circumventing all known measures in an easy yet powerful injection attack.

  • FormulaOne is not about fast cars but instead about writing fast algorithms for problems we don’t have an optimal solution for at the moment, but whose solution can be automatically verified (to some useful degree). State-of-the-art reasoning LLMs perform near 0%. Surely soon to be seen on your nearest live sport channel!

  • Games everywhere: GVGAI-LLM turns 100+ rule-based arcade games into a text-only benchmark for LLM agents, using ASCII maps and natural-language rules. Analysis shows that current models still struggle with spatial reasoning and basic planning. Instead, LLM GameLab is an open, interactive platform for testing how LLMs follow rules and make decisions in simple board games. It supports LLM-vs-LLM and human-vs-LLM matches, validates moves via GDL, logs wins/illegal moves/latency.

  • Reasoning about cause and effect is essential for effective action in the real world. CausalARC builds a systematic benchmark of causal reasoning tasks using a fully specified causal world model. The states of the world model are pixel grids, inspired by Chollet’s Abstraction and Reasoning Corpus. The authors present evaluation results for 4 LLMs in a range of regimes, including in-context learning and program synthesis.

  • This preprint proposes “Behavioural Fingerprinting” for language models, probing models across 7 different dimensions. They found that their tests of “Abstract Reasoning” and “Causal Chain Analysis” produced almost universally perfect scores. However, they saw a broader spread of results on “Sycophancy Resistance”, “Robustness” and “Metacognition”. The authors argue “as core reasoning becomes a solved problem, the key differentiator for frontier models is their portfolio of designed behaviors” such as sycophancy, which are not so straightforwardly more-is-better.

  • This preprint (covered in the opening) evaluates LLMs’ ability to determine suitable policies to tackle homelessness with scenarios for 4 geographical regions. For each scenario, the LLM ranks 4 possible policies. Evaluation involves comparison with human experts (a shaky baseline, as inter-expert agreement is lower than agreement between different LLMs) and by implementing the selected policy in an agent-based model (ABM) and measuring how different metrics change (here, LLMs outperform humans as well).

  • NoveltyBench measures whether models can produce distinct useful outputs rather than near-duplicates, revealing that even frontier LLMs often collapse to just a few functionally similar answers. Alignment and preference-training, designed to make models “safe,” appear to suppress diversity even further, suggesting creativity is an early casualty of current alignment pipelines. If models all converge on one “right” answer, they may be safer, but also less helpful, less human-like, and less trustworthy in open-ended or subjective contexts.

  • The growth of benchmarking datasets has made evaluation increasingly fragmented, while practical deployment also involves constraints like cost, latency, and environmental impact. Existing leaderboards typically collapse these complexities into single aggregate scores, which obscure critical trade-offs. xLLMBench addresses this gap with a decision-centric framework based on multi-criteria decision-making, enabling users to balance performance and non-performance factors, rather than relying on a one-size-fits-all leaderboard.

  • With the rise of start-ups like Chelsea Finn and Sergey Levine’s Physical Intelligence (π), robotic Vision-Language-Action agents are the future. But that means we need to evaluate their safety properly before we let them loose in people’s homes. ANNIE presents the first study on adversarial safety attacks on VLA agents, firmly rooted in ISO standards for human-robot interactions. ANNIE-Bench is a benchmark of 2,400 safety attacks and ANNIE-Attack is a framework that permits the iterative development of many more.

  • GDPval is a benchmark of occupational evaluation tasks, picked from the 9 most important industries (in terms of contribution to US GDP). The authors worked with experienced professionals to create representative tasks that reflect their day-to-day work. Then they use human experts in the same domain to grade outputs, asking graders to (blindly) compare AI and human generated outputs; they then build an autograder to approximate those ratings. They find that “today’s best frontier models are already approaching the quality of work produced by industry experts”.

  • LLMEval-3 attempts to address some of the limitations of current benchmarks, such as contamination and overfitting, by creating a dynamic testing framework. LLMs are tested via a secure session on an unseen random sample from a private test set of 220k questions, then responses are ranked using a calibrated LLM-as-a-judge process that shows 90% agreement with human experts.

  • A preprint introduces the Agentic Benchmark Checklist (ABC) to… check agentic benchmarks. They start by noticing that some benchmarks have poor outcome validity — an output is marked as successful even if it isn’t — or task (a.k.a. construct) validity — a task that can be solved even if the agent does not possess the target capability. Their checklist is lengthy, but adopting it seems worth the effort, given that they find, for some benchmarks they consider, “estimation errors of agents’ performance by up to 100% in relative term” 🤦🏽‍♂️.

  • This paper uses Item Response Theory to rate ML datasets based on their difficulty and their ability to distinguish between strong and weak models. This is achieved by using the results of 95 classifiers on 509 datasets. The authors show that these scores correlate with standard complexity measures and can be used to predict them for new datasets. They also introduce two curated suites of 30 datasets (one diverse and one challenging) to make model evaluation more informative and less biased than random dataset selection.

  • Relatedly, a recent paper proposes “fluid benchmarking”, which applies a 2PL Item Response Theory (IRT) model to LLM evaluation, estimating item difficulty and discrimination from existing results and then adaptively selecting items based on latent ability estimates. While similar IRT-based approaches have been explored before, usually known as (computerised) adaptive testing, the main contribution here is applying this framework to track capabilities during pretraining, using adaptive item selection to estimate model abilities more efficiently.

  • A preprint formalises the task of predicting continuous scores for long-form generation tasks starting from the prompt and model outputs. This could be used to filter outputs based on hard-to-elicit scores (e.g., because they need human judgement). On some tasks, simple baseline methods yield decent point estimates and good uncertainty quantification with as little as 16 training points and outperform LLM-as-a-judge; however, performance is poor on other tasks. In a shameless act of self-promotion, this is similar to a recent work by some authors of this newsletter, which however focuses on binary scores and on anticipating success.

  • In a similar vein to the above work, this IJCAI 2025 paper introduces R2PE, a benchmark for evaluating whether we can predict if a large language model’s answer is wrong by analyzing its chain-of-thought. The researchers propose a “Process Discernibility Score” that examines contradictions between multiple reasoning chains and outperforms counting answer agreement on average across 45 datasets. However, the approach requires generating multiple reasoning chains, making it computationally expensive.

  • Another approach for the same problem is presented in this preprint, which introduces a metric to quantify uncertainty in the generation process of reasoning LLMs by looking at the token-level predictive entropy of the probabilities of the possible answers. This is direct and efficient, but it’s only applicable for questions with a large possible set of answers (open-ended rather than multiple choice or factual).

  • Relatedly, this IJCAI 2025 paper shows a method that detects whether an answer contains a hallucination, by only relying on the internal activations of the model before and after answer generation. The main limitation is that it depends on access to model internals, so in case the model is not open source, only model owners can employ this method.

  • Moving to a finer level, a new preprint (see Twitter thread) introduces a “streaming hallucination detector” that highlights specific claims that may have been hallucinated. To train this, they build an annotation pipeline that uses web search to identify tokens corresponding to fabricated entities, and then train linear probes on model activations. Performance is good and generalises to maths (which wasn’t in the training set); how soon will we see this on ChatGPT?

  • Finally, a preprint shows that a linear direction in the activation space of LLMs can distinguish between activations to questions the model will answer correctly and those where the model will answer incorrectly, for factual recall questions. In contrast to the above work (which focuses on generated tokens), this does not generalise to a mathematical reasoning dataset, suggesting that it is harder for a model to determine its ability to solve reasoning tasks relative to factual recall ones.

  • But how do we evaluate different oversight systems on a common ground? This preprint introduces a workflow for “monitor red-teaming”, stress-testing oversight systems against adversarial agents that try to evade or sabotage detection. Interestingly, they find that weaker models, especially ensembles of them, can serve as reliable monitors, and that giving the monitor more information, such as the agent’s full task context, can actually make performance worse, with some adversarial setups pushing AUC scores below 0.5 (worse than random). All evaluations are offline, with the monitor inspecting full trajectories (actions and CoT), so it remains to be seen how well this holds up in live deployments.

News and Events

  • AI systems can complete 1-hour coding tasks (with 50% success rate), but can they outcompete a mouse? Not a device to move your pointer on the screen, but a biological mouse! This competition at NeurIPS 2025 challenges AI systems to tackle hours of mice performing a visual navigation task in a detailed virtual reality (VR) environment, under a range of perturbations to the visual stimuli (fog, etc.).

  • General Reasoning is hiring for an Eval Engineer to “Build the next generation of agentic evals to measure long-horizon capabilities”

The International Programme on AI Evaluation: Capabilities and Safety is looking for 40 exceptional students and professionals from around the world for a 150-hour hybrid course that blends lectures, hands-on labs, and a capstone project week in Valencia. Students will be fully funded through Open Philanthropy and certified by ValgrAI.

News to share? Feel free to reach out to ai.evaluation.newsletter@gmail.com

Getting the digest: Once a month if you join.

Editor-in-Chief: Lorenzo Pacchiardi

Contributors to this month’s digest: Peter Romero, Fernando Martínez-Plumed, Jose H. Orallo, Wout Schellaert, Jonathan Prunty, Ben Slater, Marko Tesic, Behzad Mehrbakhsh, Zack Tidler, Kozzy Voudouris, Irene Testini, Joseph Castellano.

No posts

Read the original on aievaluation.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.