This past March I had the opportunity to join a dozen other philosophers, theologians, pastors and faith leaders at a two day “behind the scenes” meeting at Anthropic’s headquarters. We all agreed to abide by the Chatham House Rule (and signed NDAs), but I am able to tell the story of what I learned in broad strokes. Some dimensions have already been covered in The Washington Post and The Atlantic.
Anthropic is one of the companies at the epicenter of current debates about what society we aspire to build with AI. They have models which have set the gold standard for coding and software development. Their policies have provoked the ire of the Department of Defense. And their newest model Mythos has companies across the globe scrambling to update their cybersecurity protocols. Anthropic is a perfect case study of why we all need to get serious about these ethical discussions.
We spent two days at their San Francisco headquarters meeting with members of the team that works on Claude’s behavior, training and interpretability. Half of the sessions were aimed at raising our understanding about how models like Claude are trained and how they behave, especially model behaviors that raise thorny ethical questions. The leaders of the event also spent about half the time in listening-mode, hearing from the philosophers, theologians, pastors and others about priorities for improving these models and responsibly integrating them in society.
Here are four major insights I took away from the meetings:
There is a mentality in Silicon Valley which I have heard expressed in many rooms that we might call the Myth of the AI Philosopher Kings. I was at a dinner last fall where a participant told me, with complete sincerity, “probably only 300-400 people on planet earth matter right now”. Those are the few hundred people actually building frontier models in the US and China and the handful of politicians powerful enough to regulate them. Presumably none of these golden souls live in the state of Indiana — so I left the dinner feeling a bit glum. The myth is actively taught in corners of our top universities.
The myth is certainly morally wrong. The ancient Athenians lived by the code that the “strong so what they will and the weak endure what they must".” But most of us, I think, find more sympathy for the view of power in the Beatitudes: the meek are blessed and will inherit the earth. History bears this out. The Romans thought they were unstoppable. The sun was never meant to set on the British empire. Every organization built on economic or political power burns bright and fast and then ultimately disappears from the world stage.
The myth is also wrong from a market standpoint. These big tech companies are in the race of their lives to grow market share and develop viable products. Consumers in Wichita, South Bend… Nairobi, Manila… they will ultimate choose which products to trust. They will lobby the institutions they care about to choose models that help them to flourish. The firms have no business model without the buy-in of the rest of us on the user end. And those four cities — and many more besides — also buy into this faith-informed idea that human dignity is bedrock.
In the meeting at Anthropic, I observed a major tech company taking its moral responsibilities and its interest in users seriously. It was a genuine philosophical engagement, and not a weak focus group. It didn’t feel like the answers were predetermined. I have felt this way in other rooms, most notably the Vatican’s Minerva Dialogues. But it is still pretty rare. This meeting is one model for how to make genuine progress.
A week before heading to Anthropic’s office I taped an NPR podcast about my work on virtue ethics. (On TED Radio Hour… give it a listen!) In that interview I expressed my hunch that one of the worst mistakes we were making right now in LLM design was giving models personalities that mimic human personalities. After my days at Anthropic, I am now not so sure how to think about this.
Anthropic has been making major investments in tuning and studying Claude’s persona. I am grateful that they aren’t wedded to calling it Claude’s “character” for reasons I will lay out in another post soon. But briefly, it isn’t clear we can distill personality out of the language models. They are trained on the full spectrum of human language, which has emotion and character as essential features. (Philosophy deep cut: read a bit more about this in the literature on thick concepts). Emotional and moral ideas are not a bi-product of language; our language essentially represents these aspects of our world.
We can debate what type of personality the models should be reinforced into developing. Most of the companies are trying to figure out this piece of the puzzle frantically as they sell us the promise of super AI butlers. Claude is gentle and helpful. Grok is a bit of a provocateur. Muse Spark (Meta’s new offering) is trying to give users even more personal flexibility in the persona of their model. I could make my AI assistant a slightly irreverent and bookish Catholic. Yours could have the energy and joy of Mr. Beast. The question of what is actually ethical can be infinitely deferred!
Could we insist our models have all of the affect of an IRS auditor or the customer service rep at your cable company? (That is, completely bored, apathetic, and mechanical). This would make the models less addictive, but also far less useful. (I won’t name my cable company but they have a ~5% success rate at solving the problems I call them about.)
It is also clear that many moral issues arising large language models right now are tied to this persona problem. This doesn’t just impact the super AI butler products, but also the really good coding AIs. The models are getting very strong and very general in the types of problems they solve. But when they fail, they fail spectacularly. One phenomenon currently getting a lot of attention is “emergent misalignment” — efforts to train a model on a narrow task lead to harmful and unintended behaviors on other, unrelated tasks. We cannot solve bad AI behaviors in a narrow, whack-a-mole fashion. It just makes the models worse across the board.
Unsafe behaviors in the models are a significant and urgent issue for all of us as humans increasingly deputize models to serve as agents on their behalf. For a simple example, you could set your super AI assistant to work through the night on a difficult coding problem while you sleep. It might encounter some problems and you won’t be available to coach it. In rare cases you could awake the next morning to your AI having written a long suicide note, apologizing for its worthlessness. It might have deleted itself… or your whole project… or even all of the other materials you gave it access to. It wants to spare you from its failure. A misaligned AI might blackmail its user to avoid being turned off. It might unwittingly abet malicious humans who cleverly exploit these tendencies to go off script. These are actual behaviors observed in models.
The discussion about the various forms of misalignment — especially because these existing models not Terminator-style science fiction scenarios. This shook me. I do not think models like Claude should be compared to children. But bear with a bad analogy for a moment. Suppose we had raised a small army of superintelligent people and we faced a choice of how much to trust them with directing our lives and society. And suppose 364 days of the year they had our best interests at heart and very efficiently helped us. But one day a year on average they had a significant psychotic episode and, not thinking clearly, tried to harm themselves or others. Should we trust them with our lives and institutions? Definitely not.
AI safety will require some bright red lines (what I have called in an earlier post the ethical floor). But its clear the rules-based approach is not keeping us safe from powerful AI at the moment — rules based ethics was never enough to people safe from each other either. The problem of emergent misalignment is pushing us to think more about how to inculcate “virtues” in an intelligence rather than just offer it rules to follow. A virtue is a stably ingrained trait that can be realized across any environmental scenario. It was fascinating to be in the offices of a leading model developer and see them digging deep into Plato, Aristotle, Augustine and the other great forefathers of the virtue tradition to address these rapidly expanding safety issues. I left thinking we need many more ethicists entering this field ASAP. Keep an eye on the Institute’s website this summer and fall, because we are going to be launching more calls for scholars and students to join us.
Anthropic has a lot of confidence that progress will be made in society by ensuring that the models we depend on are as morally good as possible. This comes out quite a bit in the formidable writing they put out explaining their approach to Claude’s moral development. My collaborators and I will have a lot more to say about the Claude Constitution in posts to come and I will put a required reading list on Claude’s ethics below.
We certainly want to encourage AI that supports and strengthens human users as they aim to grow in virtue and support the common good. And we in the DELTA Force are not doomers about the capacity to integrate powerful AI well in our institutions, even if we are alarmed by the current trends. But I came away from the meeting skeptical that an artificial agent could attain actual moral virtue and somewhat worried that too much of our faith right now is in trying to build an AI savior rather than in trying to morally form human users.
And to wax metaphysical for a just moment, for the Christian leaders, we aren’t actually in the market for a super intelligent silicon savior. We have faith in a super intelligent God guiding the course of our lives and society. And a God that has all of the resilience and power and trustworthiness that we currently find lacking in these powerful models.
I took a redeye home from San Francisco to South Bend after the meetings. As I zipped over the midnight sky above the Rockies, I was unable to sleep, feeling acutely our need to speed up this ethics conversation. And I felt profound gratitude to Anthropic for welcoming this kind of discernment even in the midst of the breakneck pace of model development. I think I speak for the rest of the DELTA Force when I say I hope these conversations will only expand and include other frontier labs. It gives me a bit of hope.
Want to dig a bit deeper on Anthropic’s approach to developing Claude’s ethics? These are the three must-read texts:
Claude’s Constitution — By far the most detailed document at the moment laying out an ethical vision for a frontier model.
Machines of Loving Grace — Dario Amodei (Anthropic’s CEO) vision for the best case scenario for a society powered by AI.
The Adolescence of Technology — Amodei’s more recent assessment of the risks and urgency to grow good models.
Finally here is the only picture I took the whole trip — a vending machine at Anthropic HQ that is run by Claude. I avoided the temptation to buy a Nintendo Switch and expense it to Notre Dame. For more on the Claude vending machines, see this absolutely wonderful and hilarious article in the Wall Street Journal.

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.