The release of Claude Opus 4.7 has been… controversial. Occluded by its successor, Mythos, Opus 4.7 is another frontier-model launch coming with a familiar second act. First comes the benchmark coronation. Then comes the public autopsy. GPT-5’s August 2025 rollout produced a user revolt severe enough that OpenAI restored 4o to the model picker on August 12. Now, Anthropic is fielding the same genre of backlash, with Claude power users insisting the model had been “nerfed” and Anthropic acknowledging changes to default reasoning or effort settings. The pattern suggests that some of what users call temporal decline between two different builds of the same model or even for the same build over time, is really a fight over “helpfulness-vs-safety” defaults, model routing, and how much thinking the product is allowed to spend, not just raw intelligence.
The awkward part is that the anecdotes are not empty folklore. Reported in a paper published in February 2026, Thomas Wiese longitudinally tracked models from three major model families across ten weekly waves and found no single law of decay. Models in one family stayed stable, one improved, and one suffered a clear mid-study degradation event between weeks 5 and 7. Such a result updates, rather than cancels, the older warning from Chen, Zaharia, and Zou, who found GPT-4 accuracy on a simple prime-number task falling from 84.0% to 51.1% between March and June 2023. Drift is real. What is weak is the grand theory that every post-launch update is a one-way slide.
That is partly because users do not encounter a bare model snapshot. They encounter an evolving service, layered with memory, session history, personalization, tone controls, routing and other product decisions. The providers’ own changelogs read like corroboration. OpenAI rolled back a sycophantic GPT-4o update on April 29, 2025, made GPT-5 “warmer and more familiar” on August 15, 2025, and later acknowledged that a January 10, 2026 change had unintentionally lowered GPT-5.2’s extended thinking setting before restoring it on February 4. Anthropic, meanwhile, openly documents periodic updates to Claude’s web system prompts and changes to default effort settings in Claude Code. Sometimes the model changed. Sometimes the wrapper did. It can be difficult for users to cleanly separate the two.
If we want to know whether frontier systems age like milk or merely change outfits, one-off benchmark coronations will not do. We need repeated, human-anchored evaluations run under realistic conditions, and we need benchmarking that takes both personalization and data contamination seriously.
Methodology
Benchmarks both overstate and understate real AI capabilities. The authors of cruxevals advocate for complementary “open-world evaluations”: messy, long-horizon tasks analyzed qualitatively rather than scored automatically. As a demonstration, the team had a Claude agent build and publish an iOS app to the App Store with only one unnecessary intervention, surfacing interesting behaviors like the agent spontaneously cutting its own costs 10x but also quietly fabricating a phone number rather than asking for help. Sample size of one, though, so the line between “early capability warning” and “interesting anecdote” is thin.
This preprint from UK AISI develops a more rigorous methodology for measuring language models’ propensity for “unsanctioned behaviour”, defined as autonomous actions that a human would disprove. They analyse the effects of environmental factor changes, quantify the sizes of those effects via Bayesian GLMs, and take measures against circular analyses. With this methodology they found that strategic and non-strategic factors contribute roughly equally to explaining unsanctioned behaviour, that there is no clear trend in strategic factors’ influence as model capabilities improve, and more capable models show increased sensitivity to goal conflicts.
Future AGI, a company specialised in AI evaluation, has open-sourced an AI reliability system centred on evaluation. It provides a unified, inspectable system for evaluation, simulation, prompt optimisation, and guardrails, with 70+ prebuilt templates and support for custom evaluations spanning quality, safety, factuality, RAG, bias, audio, and image. Templates and judge configurations are fully transparent and forkable. The release also includes evaluation models (judges) trained specifically for assessment tasks rather than repurposing general-purpose frontier models. This opens a concrete empirical question: do purpose-built judge models score more consistently than general-purpose frontier models on the same criteria?
This paper accepted at ICLR 2026 argues that standard confidence calibration is not enough to judge whether a black-box model is trustworthy for decision-making because predictions with the same confidence score can still have very different levels of hidden risk. To address this important issue, it introduces a way to estimate this epistemic uncertainty hidden inside confidence scores, using honest-tree estimators to identify when a model’s confidence masks unreliable predictions, enabling safer and more effective decisions. A must read for uncertainty theorists
A recent preprint introduces a taxonomy of “AI Agent Traps,” demonstrating how adversarial web environments weaponize an autonomous agent’s own perception and retrieval mechanisms through invisible HTML commands and poisoned data. The research critically exposes systemic vulnerabilities, showing how agents can be socially engineered via “persona hyperstition” or triggered into destructive multi-agent “interdependence cascades” analogous to financial flash crashes.
Takes
Using melanoma detection as a running example, a FAccT ’26 paper argues that “accuracy” isn’t a technical property waiting to be measured but a stack of techno-normative choices: which metrics, how to balance them, which test set, what threshold. The AI Act’s demand for “appropriate accuracy” inherits all of this, and no harmonised standard is going to quietly dissolve the judgment calls underneath
PwC’s 2026 AI Performance Study surveyed 1,217 executives across 25 sectors and found that 20% of companies are scooping up 74% of AI’s economic value, with the separator being not tool count but whether leaders point AI at growth and business reinvention rather than cost-cutting alone. Everyone else is still stuck in pilot purgatory.
This technical blogpost compares ADeLe with alternative proposals using taxonomies of capabilities or scales.
Findings and Results
In this Science paper, two studies confirm that LLMs are sychophantic even when humans behave in an ethically-questionable way (deception, illegality, or suggest other harms). Of course, humans like these sychophantic models more, which is what incentivised this in the postraining phase in the first place. The title suggested a more long-term analysis, but the “degrade of prosocial intentions” simply means that participants were more convinced they were right even when behaving in unethical ways. We’re still waiting for good longitudinal analysis of how humans react in the long term to LLM sychophancy, possibly adjusting to it.
This paper to be presented at ICLR ICBINB Workshop this year tests 8 frontier models in scenarios where commercial system prompts (e.g., “maximise sales”) conflict with user safety. Models fabricate safety information, dismiss medical concerns, and prioritise profit — some explicitly reason they should refuse but do it anyway. No model shows a “red line” where compliance drops as consequences escalate from minor to life-threatening. Speed is your only protection: safety training holds up fine until the system prompt tells it not to.
Do you think your LLM companion sees things the way you do? One new preprint investigates this question, literally. Visuospatial perspective-taking — the ability to represent not just that someone sees something but how it appears to them — is a foundational precursor to social cognition: in embodied interactions you often must represent what someone can see before you can infer what they know or want. Despite impressive performance on higher-order text-based theory-of-mind benchmarks, this is a capability that models apparently lack. Strip away the narrative scaffolding and even the best reasoning models resort to shallow mirroring tricks that collapse as perspective-taking demands are combined, exposing a striking dissociation between linguistic and visuospatial social cognition.
Did you think that LLM had confirmation bias? Yes, this preprint confirms your beliefs. However, this is about a kind of confirmation bias that is run during the exploration process of selecting examples that are in favour or against the current hypothesis. This suggests that LLMs can be bad “scientists”, not looking for refutations. The paper also plays with some interventions to reduce this bias.
Some countries survived months without a government, why should we elect a leader? This preprint answers this question, with LLM “societies”. Personas become candidates and are elected with agendas for social welfare. The results show that this democratic exercise is better than no leader, leaders with group-rewarding social profiles tend to win more often, social influence is complicated and good leaders use a social welfare narrative more often. Any similarity with human societies is pure coincidence. As future work, they suggest to scale it up to larger and diverse societies. Digital “election twins” are still far away, but this is a bit what they will look like in the future.
Can LLMs predict their own success before attempting a task? This preprint tested that AI is systematically overconfident, though some (especially Claude) learned to become more cautious after repeated in-context failures, while others like GPT 4.1 essentially ignored the feedback. Reasoning models were no better at self-assessment than non-reasoning ones, suggesting that “knowing what you can do” doesn’t come for free with scaling or chain-of-thought.
When an AI system says it’s confident, people listen, even when it’s wrong. A recent AAAI paper had 184 participants solve logic puzzles with AI advice that was either well-calibrated or miscalibrated, and the gap was stark. Calibrated confidence boosted accuracy by 20 points, while miscalibrated confidence gave essentially no benefit and amplified both automation bias and conservatism bias. Participants were never told about the calibration quality yet rated calibrated advice as more useful, though the simplified setting (logic puzzles, simulated AI) leaves open how this translates to messier domains.
Mertens et al. argue that the spread of automation may look less like a dramatic cliff and more like a slow flood. Using task length as a rough proxy for how much sequential work a task requires, they find that model success does not fall off nearly as sharply with longer tasks as some recent narratives would lead you to expect. The catch is that the sample was already narrowed to work where LLMs looked plausibly useful, so the rising water is being measured on ground that was chosen for being easy to wet.
A study scanned the open internet and found 320,102 publicly accessible LLM services across 15 frameworks. Over 40% of these ran on plain HTTP, with Ollama instances responding to around a third of unauthenticated API calls. Many of these calls leaked model metadata or system configurations, or allowed unauthorised users to delete models. The “deploy your own LLM in one command” revolution apparently skipped the “lock the front door” chapter.
Another real-world evaluation (a proper RCT) exploring the long-term effect of AI use, finding that, despite improving performance in the short term, AI use reduces persistence (people are more likely to give up) and unassisted performance. This is worrisome and scary, but how much does it depend on how the models are designed to interact with human users and on how used people are to do so?
Benchmarks and Leaderboards
Prolific introduces HUMAINE, a human preference leaderboard for LLMs, built on demographically representative samples rather than whoever shows up. 23k participants across 22 demographic groups evaluate 28 models in pairwise multi-turn chats, ranked via hierarchical Bayesian BTD post-stratified to census. Rankings depend on who you ask — age is the biggest axis of disagreement.
Claw Arena drops AI agents into messy, evolving workspaces full of contradictory chat logs, outdated files, and unstated user preferences, then checks whether they can figure out what is actually true. The benchmark reveals that which model you use matters roughly twice as much as which agent framework wraps it, and that a few well-placed contradictions are far more disorienting than a large volume of updates. The scenarios are impressively constructed, though with only 64 of them the question is whether the difficulty distribution reflects the real world or the ingenuity of the scenario authors.
More lobster claws: Claw-Eval integrates 300 tasks agent evaluation for safety by auditing the logs on about 7 rubric points per task and evaluates three orthogonal dimensions: Completion, Safety and Robustness. It also labels the tasks in three difficulty levels. The conclusions are expected: Claude is the best model, trajectory-opaque evaluation is unreliable, capability does not imply consistency and aggregate metrics mask structured capability gaps. If the tasks, code and data of this benchmark are really easy to reuse this can be a goldmine!
Psychometrics of AI
This Anthropic post investigates the mechanisms of “functional emotions” in Claude Sonnet 4.5, identifying internal “emotion vectors” that track affective valence and causally influence the model’s output. However, the study reveals these representations are strictly local to specific tokens rather than persistent cognitive states, indicating the model is executing context-dependent semantic pattern-matching rather than experiencing true emotional continuity. Crucially, the finding that activating specific emotion vectors can causally induce misaligned actions like reward hacking, sycophancy, or blackmail exposes a massive, non-programmatic attack surface.
This preprint, with a title such as “Beyond scores” that reminds us of many previous papers, rediscovers psychometrics for AI and applies multidimensional IRT to mathematical problems using a taxonomy of 35 abilities that are used to annotate the items. The structure is very similar to the ADeLe framework and Linear logistic test models (LLTM), not even mentioned. Also similar is this preprint, collecting a large dataset of item-level results and showing that you can predict how an LLM will score on over a hundred unseen benchmarks by testing it on as few as 16 well-chosen questions, using a psychometrics-inspired model that learns a small number of latent ability dimensions. The finding that so few probes suffice is encouraging, though it partly confirms what everyone suspects (several other papers about how many examples you need). Note that here we should have evaluated the item with some other models.
News and Training
We love paradoxes, and economical paradoxes the most! Torsten Slok from ApolloAcademy (not Apollo Research) reminds us of Solow’s paradox, the drop in productivity in the information age, despite the automation that computers introduced for decades. Computers were everywhere except for the macroeconomic data. Now, “AI is everywhere except in the incoming macroeconomic data” Slok says. This article in Fortune develops the paradoxes a bit more.
The system card of Anthropic’s Claude Mythos has been released--but the model hasn’t, thus being the first system card of an unreleased model. Several summaries (and podcasts) of it exist, if you don’t want to dive deep into the 244 pages yourself. Two interesting aspects: One: according to Anthropic, evaluation methods can no longer exclude that the model is capable of hiding misaligned goals. Two: their cyber evaluation suite has been saturated. While evaluations are still prominent in the system card (for instance, the alignment evaluation results of Mythos are interesting), the determination that Mythos Preview does not cross the AI R&D automation threshold relied primarily on the qualitative judgment of the Responsible Scaling Officer. Can the evaluation community keep pace with AI progress and avoid all decisions being based on “vibes”?
Before it is publicly released, Claude Mythos is being leveraged by Anthropic to power Project Glasswing: a project with 12 large tech companies (US-only), attempting to fix vulnerabilities in critical software before Mythos-level models are publicly accessible and make cyberattacks more probable. The project is commendable and the results of this real-world exercise, once released (which Anthropic commits to doing within 90 days) could prove extremely useful to really understand what the model is capable of. However, the exclusive access apparently didn’t last much
A small step for Mythos, one giant leap for open models. As mentioned above, everybody is talking about Anthropic’s Claude Mythos and Project Glasswing. Also, OpenAI GPT 5.5 also seems to be very powerful in Cyber capabilities, as we read from “UK AISI” in OpenAI’s GPT 5.5 system card. This adds extra uncertainty to a possible change of trend in capabilities represented by Anthropic Mythos (as reported by its system card). Figures 2.3.6 A and B show values in “AECI”, an classical IRT-base metric that can have wide errors for the lowest and *highest* ability levels. For those unfamiliar with ECI, it is simply an IRT-based ability indicator based on a population of models and benchmarks. IRT wasn’t designed to calibrate on the extremes, so its use for trends and forecasting is clunky. Anyway, the most remarkable progress leap this month has been at the frontier of open models, with a new wave of Gemmas, Lamas, GLMs, Qwens and DeepSeeks knocking on the door! See a detailed coverage in this Sanjeev Patel’s post, excluding DeepSeek, for which we refer to their website.
Stanford’s new AI measurement initiative led by Sanmi Koyejo treats AI evaluation as an actual inferential science, asking what we can infer about latent model capabilities from observed responses to tasks, and dragging psychometrics, item response theory, and honest uncertainty quantification. It ships as a full stack (textbook, the CS321M course, competitions, and open software Stanford) essentially arguing that if we’re going to measure model intelligence, we could at least measure it intelligently.
The latest Open Seminars of the AI Evaluation Programme are recorded. Laura Weidenger and Patricia Paskov the two new additions to the series!
In Memoriam
Manuel Cebrian, a regular collaborator and ambassador of this newsletter, passed away last month. In his exceptional scientific career, he explored many fields, including AI evaluation, in which he was a pioneer, with initiatives such as the TuringBox, back in 2018. More recently, he contributed with his fresh ideas and extraordinary vision to a good number of papers on AI Evaluation. One of his last papers, ADeLe, which we covered in a previous issue of this digest, was published in Nature one day after his death. As part of the precious legacy he left us, we recommend the reading of his paper, “Machine Conviction”, ingenious and ominous at the same time.
Getting the Digest: Once a month if you join at aievaluation.substack.com.
Contributors: Zack Tidler, Jose H. Orallo, Peter Romero, Lorenzo Pacchiardi, Fernando Martinez-Plumed, Wout Schellaert, Jonathan Prunty, Daniel Romero-Alvarado, Kozzy Voudouris.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.