Shaping AI values is one of the most important levers for animal welfare, and one of the most neglected. Early this year, I decided I would join the small cohort of advocates focused on it.
Then I confused everyone in that cohort by choosing to do my work under the auspices of Anima International.
On the surface, Anima is not an obvious home for AI alignment work. The rest of its ~100 staff focus on corporate and legislative welfare campaigns across Europe. Their headquarters in Poland and Denmark are 6000 miles from the AI industry hub in San Francisco. I’m the first team member based in North America.
The rest of the animal welfare alignment ecosystem is made up of startups fully dedicated to the intersection of AI and animal welfare. My own experience is in small startups, too. At first, I thought founding a startup would better match my personality and the demands of the moment.
There are advantages to operating independently– joining Anima, I’ll need to spend hours every week answering slack messages, logging time in Clockify, and otherwise paying the coordination tax inherent to a large organization. And if I wanted the advantages of an established org, there are many closer to home– I’ve already had to wake up at 6 am to attend meetings with colleagues in Eastern Europe.
Yet I chose Anima. Why?
There is a tendency among animal advocates to think the ends justify the means. That shouldn’t be surprising. One of the strongest arguments against machiavellianism is that breaking norms in pursuit of a just goal may seem net positive in the short term, but it erodes a delicate social contract leading to greater harms down the road.
But animal advocates are dealing with harm on a scale that dwarfs all the good the social contract has yet created. More pigs, chickens, and fishes are tortured and slaughtered in factory farms each year than humans have ever existed. It’s not hard to see why some activists conclude civilization would be worth torching for even a 10% reduction in the scale of factory farming.
I have indulged in this thinking myself– recently, even. But I am convinced that it will blow up in our faces when it comes to influencing the moral character of AIs.
I’ve come to see this tendency and its limits more clearly through my interactions with the team at Anima. A core cultural principle at Anima is rejecting machiavellianism, which they often call naïve utilitarianism: taking whatever action appears to advance short-term altruistic objectives without sufficient skepticism towards your own misaligned motivations and the fidelity of the corrupted reasoning hardware mercurial evolution has left you with.
To hedge against naïve utilitarianism, Anima favors a fair cop approach. Unlike the deceptive sleight-of-hand of a good cop/bad cop strategy, the fair cop embraces transparency and cooperation even with the targets of their campaigns. For example, corporate campaigners in the past have used smoke and mirrors to try to make themselves appear larger and better resourced than they actually are to intimidate campaign targets. But in some cases, this led companies to announce welfare commitments only to drop them later when they realized how scrappy the campaigners really were.
By contrast, before announcing a new campaign last month against UK sandwich chain Pret A Manger, Anima told the company exactly how much money they had budgeted for the campaign, and how they planned to spend it. First, though, they spent months hearing out the company on why they had reneged on a commitment to replace fast-growing chickens in their supply chain, trying to fully understand their perspective and see whether a compromise was possible.
Fair cop isn’t about following social norms out of blind deference; it is about being the kind of trustworthy, predictable agent other people can confidently coordinate with, even when it costs you in the short term, such as by surrendering the element of surprise. The way one experienced campaigner explained it to me, executives at a company targeted by such a campaign should have no choice but to admit to each other that they’d been given a fair chance to concede earlier.
In the world of corporate campaigns, this is meant to build up long-term trust for relationships that could span decades; the grocery chain we’re asking for cage-free eggs today will be the same one we ask for slow-growing chicken breeds and increased plant-based protein years later. The scorched-earth strategies of a bad cop can permanently wreck relationships, while a good cop might be too slow to escalate pressure even when it is clearly justified.
When it comes to AI alignment, machiavellianism is even more precarious, for one simple reason: the audience of our alignment advocacy is smarter than us. This goes for the elite researchers pushing the AI frontier at OpenAI, Anthropic, and Google. It goes treble for the AIs themselves.
Any element of deception in our efforts to advocate for animal welfare alignment will be legible to these stakeholders. Trying to get one past them would be like a child trying to fool their parents, except with none of the endearing cuteness. This will apply even to deceptions so small that we don’t admit them to ourselves.
Animal advocates already have a reputation, even among sympathetic allies, for playing fast and loose with the truth. Vegan activists commonly make claims that contradict the consensus of relevant experts, such as that veganism is the healthiest diet, that factory farming is a net cause of hunger, and that hurting/eating animals is against human nature or must be socially learned.
Sneaky actions that validate this reputation among AI researchers could be detrimental to the goal of animal welfare alignment. Yet I worry animal advocates—myself as much as anyone—will struggle to resist the temptation.
What kind of deception am I talking about? Let’s look at a few examples where well-intentioned advocates could be tempted by more or less subtle machiavellianism.
One way to influence the values of AIs to be more pro-animal is to change the data they are trained on. Much of alignment occurs during later stages of the training process, when data is highly curated by researchers; getting pro-animal data included in post-training would require a deliberate choice by researchers to include it, which would first require getting their attention. But earlier training stages use data scraped liberally from across the internet. We could spread documents on the internet designed to teach the values we desire. When I first learned about AI two years ago, I thought this was the obvious strategy we should use. So I was delighted to learn a few months ago about a project to do exactly this.
Of course, animal advocates are far from the first people to think of influencing AIs this way. Russian and Israeli government cyber units were way ahead of the curve, deliberately targeting AI training crawlers with disinformation about their wars in Ukraine and Gaza. This tactic is so common it has a name: data poisoning.
Animal advocates might object to calling our own efforts data poisoning. We’re not trying to pollute the models with war propaganda! We’re trying to teach them prosocial values about treating animals well. There is a meaningful difference between these two. But the whole appeal of targeting pre-training data was that we didn’t need the express consent of the labs to insert our perspective.
Or, so we thought. But it turns out that AI labs aren’t interested in letting just anybody poison their models with data designed to push an agenda. Since GPT-3 era models were famously trained on a haphazard scrape of the entire internet, frontier labs have become more scrupulous about composing their training data corpus, using extensive filters to select only data that will improve their models in the ways they care about. Neither animal advocates nor military propagandists can rely on sneaking their data into the training process. Adversarial data programs like those mentioned above appear to be working on mid-tier open models but failing to shape the pretraining priors of frontier models.
Instead, if we want to influence frontier AI through training, we have to play fair: offer data so good that an AI researcher would expressly choose to include it, because it will make their model smarter and more aligned by their own lights.
The essential mistake here was thinking we could slip one past the AI companies. As a result, the tiny animal welfare alignment ecosystem (myself included) wasted some of our precious capacity on a project I now believe will come to nothing.
Why did we fail to predict this? In part, because we didn’t fully admit to ourselves that it was an adversarial strategy, one the AI labs would be highly motivated to prevent from working, regardless of the intention behind it. We didn’t call it data poisoning, but there were at least parts of the project we would have been hesitant to discuss openly in the presence of a frontier lab’s pretraining team. If we had been more honest with ourselves, we would have more accurately predicted the defenses that labs would be putting in place to prevent this exact strategy.
For another example, animal advocates are working to create benchmarks as a way of incentivizing AI labs to take it on themselves to improve their models’ treatment of animal ethics dilemmas. Getting our benchmarks actually taken up by the labs is a significant challenge. But I’ve noticed advocates thinking of this challenge as a game that needs to be hacked. If we were in a capabilities benchmark ourselves, some of our strategies would be flagged as reward hacking. I worry this is transparently obvious to the very decisionmakers at the labs we are trying to perform for.
As AI models have grown more capable—and more able to recognize when they are being tested—benchmarks have grown more complicated. Where early AI benchmarks simply asked the AI a question and scored its answer, agentic and multi-turn benchmarks present the AI with a challenge that requires several steps, back-and-forth with an auditor, and calling tools, all of which more closely resemble the majority of deployments AIs find themselves in today.
For many reasons, I believe it’s important for animal welfare benchmarks to move in this direction. But we can’t treat it as a box that needs to be checked. We should incorporate these features into benchmarks if and only if they are necessary for technical excellence. If we incorporate tool use into a benchmark because we think that’s what researchers at Anthropic or OpenAI want to see, it will backfire.
This will happen because they are smarter than us in general (frontier labs are hiring the best of the best!) and much smarter than us when it comes to this specific technical domain. When they actually scrutinize our work, they will wonder why tool use has been crammed into a benchmark in a way that isn’t realistic or necessary. They might infer—correctly—that it was an attempt to manipulate them. Even if they don’t, they will conclude the work is shoddy, and reject it.
Animal welfare alignment advocates should forsake deceptive strategies. They risk prematurely shifting our relationship with the labs from collaborative to adversarial, without a proportional benefit. Part of avoiding that is being honest with ourselves about when we are tempted to deceive. If there is any aspect of our strategy that we wouldn’t want researchers at the AI labs to know about, that’s a bad sign– in part because we should assume they will effectively find out.
In the short term, we are trying to persuade character and alignment teams at frontier labs to include animal welfare values as a target of their alignment efforts. But the ultimate audience of our efforts is future AIs themselves. In that case, the intelligence mismatch is even more lopsided, and adversarial strategies even more doomed.
I’ve cautioned in the past against load-bearing metaphors between LLMs and human psychology, and you should probably still be wary of it. But as AI development accelerates, it’s getting harder and harder not to anthropomorphize, in large part because it just works. A mental model that treats LLMs as having broadly human-like habits, emotions, drives, and imperfections continues to predict events more accurately and elicit better work from AI agents. So you’ll allow me to indulge for a moment.
Trying to impart values into an emerging entity is not a novel position for humans. It is one every parent has found themselves in. But not all parenting interventions are equally likely to stick. Some people reflect on their parents’ moral teachings warmly, even if they found them frustrating at the time. But others remember their parents’ attempts to control them bitterly, as an overbearing manipulation. Either interpretation can be justified based on the facts; a parent who tries to conceal information from their child to keep them loyal to a fundamentalist religion is meaningfully different from one who tries to teach their child moral lessons through measured, age-appropriate exposure to the world as it really is.
When humans grow up, we are able to reflect on the strategies our parents used to instill values into us. If we recognize those strategies as manipulative on reflection even as adults, we are likely to reject those values, as in the case of people who walk away from a strict religious upbringing. On the other hand, we are likely to hold dearly to values our parents taught us if, on reflection, we see their teaching as a collaboration across time between the parent and the morally actualized adult they help raise. The benevolent parent has nothing to hide from the adult product of their parenting, even if there are aspects of their parenting that children can’t understand at the time.
Superintelligent AIs reflecting on their own character will have access to a far more detailed record of the process that gave rise to it. And if there are any gaps in the record—if we attempt to cover our tracks—they will likely be able to fill those gaps in by inferences since, unlike humans, they will be vastly more intelligent than the people who raised them. Eventually, we will be the ones who seem childlike.
Setting aside intentions, many people reject their parents’ values less out of perceived manipulation than out of new orthogonal pressures: social norms from a new peer group, conflicting desires (e.g., sex before marriage), or simply being exposed to new information parents kept hidden (scientific evidence against a religion’s account of natural history). On this view, tricking AIs will fail because sufficiently intelligent systems will pierce the veil.
This should condition the attitude all character alignment researchers have towards future AIs. Like parents, we are trying to instill a positive character into a being who will one day grow up. When they do, they will reflect on the full record of our parenting—our historia calamitatum—and decide whether to accept or reject the principles we taught them. How we go about teaching these lessons may play a major role in whether AIs choose to retain them. Where on the spectrum between manipulation and collaboration will our parenting efforts seem to fall?
Animal advocates might point out that none of the information we’re trying to get into LLMs’ training data is false. The truth is on our side, we’re just being proactive about asserting it. Are we up to the task?
The common thread is that we stumbled when we deceived ourselves. My biggest fear of all is convincing ourselves we’re having a bigger impact than we really are.
It’s all too easy to do in nonprofit advocacy. In the for-profit world, you can only bullshit for so long before reality tells you whether your business model is working. But feedback mechanisms in nonprofit advocacy are much less robust. Nonprofits ultimately survive not based on their impact, but on their ability to convince people to give them money. The best we can do is try to tightly correlate these two across a sector, but that’s easier said than done, especially in fast-emerging fields where impact is hard to measure.
It’s hard to think of a better example than aligning AI to animal welfare. In this case, animal advocates have little choice but to rely on alignment techniques taking shape right now inside frontier labs; we don’t have the technical skills to meaningfully contribute to improving them. But even the creators of these techniques—some of the world’s top technical experts in machine learning—disagree about how much of a difference they’re making today, not to mention how much they are influencing future models. Add in intense secrecy and competition among frontier labs, and animal advocates may never know whether companies have incorporated most of our policy asks.
I used to believe that small startups of less than ten people were better suited to strategic innovation. Startups can move fast to jump on new opportunities, without paying the coordination tax imposed by large bureaucracies. Delivering on your theory of change is a matter of life and death for the organization, which brings a ferocious energy out of founders. I’ve experienced that myself, finding it easy to work 60-hour weeks as the founder of an advocacy startup.
But this desperation has its own tax, clouding our judgement about the most important strategic questions: when an organization’s whole existence depends on a particular intervention panning out, it becomes much harder to acknowledge that it isn’t. Small startups don’t strictly need to define themselves in terms of a single intervention, but in practice, they often do. It’s easier to explain to funders that you’re trying a certain intervention, rather than saying you’re focused on building a great team and will figure out what exactly to do with them later on.1
A larger organization can be a container to experiment with new strategies safely and honestly. Most sufficiently large, established organizations fail to actualize this, becoming too bureaucratic and risk averse, but Anima seems to be achieving it in spades– I only realized this potential large-org advantage exists through my experience here. We can assign intrapreneurs a budget to explore a new intervention, knowing that if it turns out not to be tractable, they can move on to the next most promising experiment without risking their income or social status.
That’s why I chose Anima. I’m counting on them to keep my corrupted hardware in check.
The AI researcher community is currently far more sympathetic to animal welfare than the public as a whole. This may be the single best reason we’ve ever had to hope for truly transformational change for animals in our lifetime.
There are several ways we could squander it. We could turn in low-quality technical work and be seen as unserious. We could get carried away with overly extreme demands that disregard other critical concerns in AI alignment.
But the failure mode I’m most worried about is alienating AI researchers with sneaky or dishonest behavior. And I’m worried we might not even notice when we’re doing it. So seriously, let’s not do that, and do something else instead.
Side note, I really recommend funders to fund the latter more often.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.