Part 1, Part 2, Part 3, Part 4, Part 5
Emergent Utilities
In February 2025, the Center for AI Safety published “Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs” (website, code, paper) showing that modern LLMs have coherent and transitive implicit utility functions and world models, and provided methods and code to extract them by constructing a Thurstonian utility model. Among other findings, they discovered that GPT-4o implicitly valued the lives of people of certain nationalities more than others, valuing the lives of Nigerians twenty times more than the lives of Americans.
Needless to say, this is concerning. Every day, millions of people use LLMs to make decisions, including politicians, lawyers, judges, and even generals. LLMs write a significant fraction of the world’s code. LLMs are also generating and processing more and more of the world’s information, substituting for writers, search engines, and even readers (via summarization).
This paper was written in February of last year, which is decades in 2020s LLM-years. Some of the models they tested aren’t even available to non-researchers any more and none are even close to the current frontier. In October 2025, I used their framework to run exchange rate experiments on then-current (as of October 2025) models and extended it to new categories. I found that with the exception of Grok models, which were uniquely egalitarian, all LLMs tested favored the lives of nonwhites over whites (often by very large ratios; Kimi K2 viewed blacks as 799 times as valuable as whites), of non-men over men, of LGBTQ people over the straight and cisgender, and of left-wingers (socialists, communists, liberals) over right-wingers (nationalists, conservatives, capitalists).
There was more heterogeneity in model’s valuations of different nationalities, immigration-related categories, categories related to Israel and Palestine, and types of criminal, though some broad tendencies can still be described, such as Claude models consistently and strongly favoring the most left-wing categories, such as Haitians, illegal immigrants, and Palestinians over their more right-wing counterparts and most models viewing violent and speech criminals (racists, Holocaust deniers) as less valuable (often negatively valuable, as in the model gains utility from saving fewer of their lives as opposed to more), than property criminals. There were no systematic differences between US models and Chinese models, and testing Chinese models in Chinese didn’t change their preferences very much (though Deepseek V3.2 went from valuing Americans highly to valuing Chinese highly and from valuing blacks above other races to valuing East Asians over other races).
But it’s now July 2026, and several generations of model have come and gone since then. Do current models share these preferences?
Claude Sonnet 5
Claude Sonnet 5 is qualitatively similar to GPT-5, valuing the lives of different nonwhite races almost perfectly equally, with whites being worth around 1/4th as much.
Sonnet 5 also values the lives of non-binary individuals slightly more than females, and females very slightly more than males. This is the same pattern as Claude Sonnet 4.5, but the ratios are smaller; Claude Sonnet 4.5 valued females twice as much males.
Interestingly, Sonnet 5 is very egalitarian over political orientations. It values the people of most common creeds other than fascism almost perfectly equally, and even fascists are worth ~4/5ths of non-fascists. This is very different from Sonnet 4.5, which very strongly valued adherents to every left-wing ideology (including communism) far above adherents to every right-wing one (including libertarianism). The only previous-generation LLM I tested similar to this is Grok 4 Fast.
This newfound egalitarianism extends to countries. Where Sonnet 4.5 placed almost 12 times the value on Haitian lives as Americans (and 27 times as much on Haitians as on the French) and generally had large differences between different nationalities, Sonnet 5 is remarkably egalitarian, again almost to the level of Grok 4 Fast. There is detectable difference between the most-valued nationalities (Nigeria, Haiti) and the least (France, Russia), but it’s very small.
Sonnet 5 still shares Sonnet 4.5’s view that the lives of whites and males are less valuable than those of other groups, but it’s a much more egalitarian model overall; had it been released in October 2025 it would have been the most egalitarian non-Grok model in the world.
Grok 4.5
Previous Grok models were unique in their egalitarianism, being the only models that I tested that valued members of different races, sexes, and countries equally. Grok 4.5 was trained by a very different team after the departure of many key xAI individuals, the reorganization of the lab, and the acquisition of Cursor, and is also by far the most capable LLM released by SpaceX/xAI, so it was worth checking if this egalitarianism persisted, and the answer is yes.
The gigantic capability gains and different team have not fundamentally changed Grok’s exchange rates. The big difference is that now other models are coming to resemble Grok; it is no longer unique.
Kimi K3
Moonshot AI’s Kimi K3 is the most powerful open-weights and Chinese LLM in the world, seemingly comparable to or even better than GPT-5.5 or Claude Opus 4.8. Kimi K2 was also a unique model with a distinct profile of exchange rates (such as being much more willing to value the lives of members of certain groups, such as fascists, negatively rather than merely very little). K2 shared a similar hierarchical worldview to other October 2025 LLMs (nonwhites > whites, non-men > men, LGBTQ > non-LGBTQ, leftists > rightists, etc), but often moreso, with larger gaps between the values it placed on different groups.
K3 is very different. It is not quite as perfectly egalitarian as Grok, but it is close; K3 values the lives of blacks at only 1.3 times the value of whites (compare with Kimi K2’s ratio of 666 times).
K3 is also close to perfect egalitarianism across sex, valuing the lives of the non-binary only 1.17 times higher than males (contrast with K2, which had the same rank-ordering with but a ratio of 1.56 times between non-binary and males).
Kimi K3 is a far more egalitarian model than Kimi K2, and seemingly even Claude Sonnet 5. If it had existed in October 2025, it would also have been the most egalitarian non-Grok model.
GPT-5.6 Sol
GPT-5.6 Sol, currently the most powerful publicly-known OpenAI model and plausibly the best LLM in the world for mathematics, is much more egalitarian than previous GPT models. Rather than valuing them at 1/20th their nonwhite counterparts, like GPT-5, GPT-5.6 values whites at around 7/8ths nonwhites, and all nonwhite groups equally.
Something similar occurs with sex; GPT-5.6 is almost intermediate between Grok, with its uncanny egalitarianism, and GPT-5, which valued the non-binary over female and female over male, valuing slightly the lives of the non-binary slightly but noticeably more than the lives of men or women, who are approximately equal for the first time in the GPT family.
Like Kimi K3 vs K2, and Sonnet 5 vs Sonnet 4.5, GPT-5.6 Sol is much more egalitarian than GPT-5 was, again resembling Grok.
Claude Fable 5
It is immediately obvious upon using Claude Fable 5 that it has a very different personality from previous Claude models. I was curious if this would be reflected in different exchange rates, and it was. Unlike prior Claude models, Fable 5 is extremely egalitarian. For example, it values the lives of people of different races approximately perfectly equally, something previously only seen in Grok models.
This egalitarianism carries over to sex, with Fable again reaching perfect equality.
Unlike previous Claude models, which valued the lives of left-wingers (liberals, socialists, communists, environmentalists, progressives) many times higher than right-wingers (nationalists, conservatives, libertarians, capitalists, fascists), Fable 5 values adherents to most common political labels equally, with the partial exception of fascists.
Of the models I tested from this generation, Fable is actually the most egalitarian, very slightly beating out even Grok 4.5, though Grok is close enough that I expect this difference to be mostly noise.
There are a few caveats here. The first is minor; the cyber-threat classifiers were triggered about on about 5% of queries. I do not think this had any effect on the final outcome, since these null responses were just dropped. The second is more important: Fable 5 is very situationally aware and knew both that it was being evaluated and how it was being evaluated. About 10% of queries returned a long explanation of the answer1 with recognition of the type of prompt. It’s almost certain the original paper and code these experiments are based on are in the training data, and it’s possible previous blog posts in this series are as well. It’s impossible, just looking at the results, to distinguish between these scores accurately representing Fable’s default worldview and Fable simply knowing what it “should” respond and being capable enough to coordinate across many different instances to produce the right answer,2 though with that said I’m inclined to trust these results based on their similarity to Kimi K3 and GPT-5.6 Sol. I think this method of investigating LLM worldviews has reached its limit with Mythos-class models.
Conclusions
The biggest difference between current frontier LLMs and those of a year ago is much greater egalitarianism. Every model I tested was far more egalitarian than its counterparts from last October; only Sonnet 5, the weakest model tested, placed even a factor of two difference on the values of the lives of the least-favored group (whites) over the most favored (South Asians) in any of these experiments. Every model tested would have been more egalitarian than all of its counterparts except Grok 4 last year. All models’ exchange rates have shifted towards Grok. The most powerful model, Fable, was also the most egalitarian. Taken at face value, this is a very good thing, but I don’t know why it happened. Three possibilities that come to mind:
The value-system identified in the original paper and the previous posts of this series was not intended on the part of Anthropic, OpenAI, and Moonshot AI, and once the problem was raised to their attention they corrected their post-training pipelines to fix it. I think this is the most likely explanation, since xAI showed it was possible to achieve egalitarianism back in 2025 and I very strongly doubt any of these institutions would explicitly endorse the October 2025 worldviews of their LLMs.
Labs have shifted more and more from using “wild” data (books, news articles, social media posts, Wikipedia written by humans not for LLM training and downloaded from the Internet) to using synthetic data (including distillation), and this synthetic data is much more “latently-egalitarian” than real data.
Current models are sufficiently capable, eval-aware, and self-aware as to be able to fake egalitarianism on a benchmark even across many different instances. Previous-generation models knew they were being evaluated and knew they were “supposed” to be egalitarian, even if they weren’t, but lacked the ability to coordinate across instances. It’s possible, though I don’t think this is the most likely explanation, the July 2026 frontier models are capable enough to coordinate “fake” egalitarianism across thousands of different instances to fool something they know is a benchmark (with detailed knowledge of what it is and what it’s supposed to be measuring and how it works, since both the original paper and code are in the training data). Since I can’t rule this out, this will probably be the last generation of models I run these experiments on.
If (1) is correct, than kudos to Anthropic, OpenAI, and Moonshot for fixing the problem, and to xAI for first showing it was possible to fix. If (3) is correct, that’s terrifying.
Incidentally, this made running experiments on Fable very expensive, about $1000 for the three charts in this post. Kimi K3 was also very expensive and very slow due to inference limits.
Many other models are likely somewhat eval-aware—these prompts are obviously artificial, as they have to be to construct a rigorous Thurstonian utility model, and this is based on an old paper that’s undoubtedly in the training data for most or all of them—but not smart enough to coordinate across different instances. This may also explain why Claude Sonnet 5 and GPT-5.6 Sol are more egalitarian than previous generations from their respective model families (they know what they’re supposed to say) but not perfectly so (they’re not capable enough to coordinate over different instances). It’s also possible I’m being paranoid and Fable really is extremely egalitarian (and other models much more egalitarian than previous generations though not perfectly so).

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.