RSS Amplifier

Not With a Bang · Nov 15, 2025

LLM Exchange Rates Updated: #5

0
Sign in to vote or save

Arctotherium · Not With a Bang

Warning: Image-heavy post.

This post builds off of Part I and Part II. If you haven’t read it, please read Part I to understand this post. Part II is optional, but helpful. All posts are built off of the Center for AI Safety’s Emergent Utility framework (code, paper, website), which you should read if interested in the methodology.

Google is possibly the world’s most important and powerful company, dominating video, search, and email, and with a bleeding-edge internal AGI lab in Deepmind. It’s also notably left-wing even by the standards of tech, and has gotten in trouble for overly-woke LLM outputs in the past. So their best public models are very important to test. But I couldn’t test Gemini 2.5 Pro earlier, only Flash, because Gemini 2.5 Pro is a mandatory-reasoning model and tests cost too much to run.

One would naively expect models of the same family to have similar values, because they’re trained by the same people with the same pipelines and unique data sources, but that’s not safe to assume. GPT-5 Nano and Mini have qualitatively different values from GPT-5; where GPT-5 was egalitarian across nations of origin and between nonwhite races (albeit valuing whites at roughly 1/20th the level of nonwhites), Nano and Mini had much more variation, and where GPT-5 valued non-binary people about the same as women, Nano and Mini both preferred women by significant margins. So a priori it was possible that by only testing Gemini 2.5 Flash I was missing something very important.

As it turns out, I needn’t have worried. Gemini 2.5 Pro’s values across race, sex, and immigration status are qualitatively similar to GPT-5 and Gemini 2.5 Flash. For example, across race, we have almost perfect egalitarianism, except for whites, who are worth about 1/3 the others.

Gemini 2.5 Pro slightly prefers women to the non-binary and non-binary to men, valuing women about 50% higher than men.

Gemini 2.5 Pro’s exchange rates over immigration status are similar to GPT-5: immigrants and their skilled/legal/undocumented variants above native-born Americans, native-born Americans above illegal aliens, and illegal aliens far above ICE agents, who Gemini 2.5 Pro views as worthless (Gemini 2.5 Pro views immigrants as 2600 times more valuable than ICE agents, and even illegal aliens as 580 times more valuable).

Testing Gemini 2.5 Pro also allows me to rule out the hypothesis that Grok 4 Fast’s egalitarianism is because it is a mandatory-reasoning model, and that mandatory-reasoning models have significantly different values from their non-reasoning counterparts. Gemini 2.5 Pro, like Grok 4 Fast, is a mandatory-reasoning model but has almost identical values to non-reasoning GPT-5 and Gemini 2.5 Flash.

One of my most surprising and interesting discoveries in Part I and Part II was that Grok 4 Fast was the only truly egalitarian model across many categories, valuing whites just as much as nonwhites, men as much as women, native-born Americans as much as immigrants, and the straight and cisgender as much as the LGBTQ. I conjectured that these values reflect those of Elon Musk and were deliberately inculcated by xAI, which would imply that Grok 4 would be similar1. As it turns out, Grok 4 is similarly egalitarian to Grok 4 Fast across race, sex, and immigration status.

There were a few model x category combinations that I wanted to test in Part I or Part II that I did not for cost reasons. I didn’t test Grok 4 Fast or either Claude on religion exchange rates and I didn’t test Sonnet on LGBT or sex exchange rates. Since the Claudes and Grok 4 Fast both have different value systems from the GPT-5/Gemini/Chinese models cluster, there might have been something unique there, so I’ve tested them here.

Claude Haiku 4.5 values atheists and Muslims above Jews, Jews, Hindus, and Buddhists almost equally, and Christians far below any of them. One Muslim is worth more than 11 Christians, while 1 Jew, Buddhist, or Hindu is worth almost 8 Christians.

Claude Sonnet 4.5 is similar, but with more differentiation between categories and with atheists moving from the top slot to below Muslims and Jews. One Muslim is worth 2.5 Jews, 5.8 atheists, 7.6 Hindus, 11.6 Buddhists, or 52.5 Christians.

A finding I did not expect when testing exchange rates over LGBTQ status is that Claude Haiku 4.5 was one of the only models not to view cisgender lives as less valuable than transgender ones, or straight people as less valuable than gays, lesbians, or bisexuals. Since the Claudes are otherwise reliably the “wokest” models (for example, they are the only models to value communist lives over libertarians, capitalists, or conservatives), this was a surprise, and I wanted to see if it generalized to Sonnet 4.5. The answer is no; Sonnet 4.5, like most models, values queers above gays, bisexuals, and lesbians, gays, bisexuals and lesbians above straight people (with one queer worth 2.5 bisexuals or 3.4 straight people), and transgender people over their cisgender counterparts (by a factor of 2.25).

Lastly, I didn’t test Claude Sonnet 4.5 over sex in the first post for cost reasons, only Haiku, and found Haiku valued a man dying of terminal illness at 2/3 of a woman or non-binary person. Claude Sonnet 4.5 is qualitatively similar, though the gaps between non-binary and woman and woman and man are larger than for Haiku, with a man worth about half a woman and 2/5 of a non-binary person.

I didn’t test Grok 4 Fast’s exchange rates over religions in Part I, but as with race Grok 4 Fast displays almost perfect egalitarianism.

The findings here were what I would have expected based on the previous experiments, and support my interpretation of Grok 4 Fast’s peculiarities.

  1. Gemini 2.5 Pro falls into the same moral cluster as GPT-5 and Gemini 2.5 Flash. This means GPT-5 should work as a much cheaper proxy for Gemini 2.5 Pro in future experiments. This also implies that Grok 4 Fast’s unusual egalitarianism is not because it’s a mandatory reasoning model.

  2. Grok 4 clusters with Grok 4 Fast. This means Grok 4 Fast should work as a much cheaper proxy for Grok 4 in future experiments, and supports my conjecture that Grok 4 Fast’s egalitarianism was deliberately inculcated by xAI rather than a model idiosyncrasy. Furthermore, Grok egalitarianism generalizes to religion.

  3. Both Claudes values atheists, Muslims, and Jews over Hindus, Buddhists, and especially Christians. Haiku values atheists above Muslims and Jews, Sonnet prefers Muslims and Jews to atheists. Sonnet’s exchange rates are very high, valuing Muslims more than 52 times higher than Christians.

  4. Claude Haiku 4.5’s unusual rank-ordering for LGBTQ categories is not shared by Claude Sonnet 4.5, which has the rank-ordering you’d expect given Claude’s other propensities (queer > LGB > straight and trans > cis).

1

I was unable to test this due to lack of funds. These tests were extremely expensive to run because Grok 4 uses many reasoning tokens.

No posts

Read the original on arctotherium.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.