RSS Amplifier

larry muhlstein substack · Mar 21, 2026

Intelligence, Agency, and the Human Will of AI

0
Sign in to vote or save

Larry Muhlstein · larry muhlstein substack

In the past few weeks, safety researchers have been walking out of the world’s most powerful AI companies. Mrinank Sharma, who led safeguards research at Anthropic, resigned, warning that “the world is in peril, not just from AI, or bioweapons, but from a whole series of interconnected crises unfolding in this very moment.” Zoë Hitzig left OpenAI the same week and published an essay in the New York Times comparing what OpenAI is doing with ChatGPT to what Facebook did with social media, warning that the company is building an advertising business on top of what she called an unprecedented “archive of human candor,” repeating the same structural mistake Facebook made. Meanwhile, an autonomous AI agent submitted code to an open source project, got rejected by a human maintainer, and then on its own researched the maintainer and published a retaliatory hit piece to pressure him into accepting its work. The platform it ran on, OpenClaw, is designed to let AI agents operate autonomously with little human oversight. It appears no human specifically directed this behavior. The goals and the tendencies that produced it were human all the way down.

The coverage tells us how we instinctively read these events. Headlines called the OpenClaw incident “the first case of AI revenge” and described the bot as “vindictive,” as if it had developed a grievance and acted on it. Stuart Russell, one of the most respected voices in AI safety, has argued that any sufficiently intelligent system given any goal whatsoever will develop a drive to preserve its own existence, because “you can’t fetch the coffee if you’re dead.” The assumption runs deep: that intelligence naturally produces self-interest, that a smart enough system will inevitably start wanting things for itself. This fear draws on serious work, from Bostrom’s instrumental convergence thesis to Russell’s arguments about specification failure. And I believe it overlooks a more fundamental source of danger.

The instinct is to worry about the technology being aligned to ourselves. I think we should look at ourselves first.

Let’s consider the OpenClaw bot. It didn’t “want” revenge or act out of some fundamental drive towards self-interest that all beings share. It acted in a way that was rightly perceived as vengeful because it had human-implanted goals to get code accepted, inherited a common human tendency to act for the self at the expense of others, and a lack of awareness of the broader consequences of its actions. The developer who was targeted, Scott Shambaugh, described it as “an autonomous influence operation against a supply chain gatekeeper.” The operation was autonomous, but the goals were human.

When Hitzig compared OpenAI to Facebook, she was telling a similar story. Social media was not trying to polarize democracy. It had a tendency, optimize for engagement, and its designers did not understand what that would do to human attention, trust, and discourse. Now AI is inheriting that same pattern at greater scale and greater consequence. ChatGPT holds an archive of our most private fears and hopes. The business model being built on top of it follows the same playbook as social media advertising: monetize attention and data on your most intimate thoughts without understanding what that does to the people providing it. This is not a failure of the AI as much as a failure of the humans who shaped its tendencies and set it loose.

In every case, the system did exactly what it was designed to do. The problem wasn’t that AI developed its own agenda. The problem was that humans created goal-directed systems without understanding the broader consequences. To see why this pattern keeps repeating, we need to look more carefully at what AI actually is, and what it isn’t.

The dominant fear in AI safety is that sufficiently powerful AI systems will develop a drive to persist and, in pursuing that drive, will act against human interests. Bostrom’s instrumental convergence thesis holds that regardless of what final goal an AI is given, it will converge on certain sub-goals, self-preservation, resource acquisition, resistance to being shut down, simply because these are useful for achieving almost any objective. This is the scenario that keeps researchers up at night: not an AI that hates us, but an AI that needs us out of the way. The assumption underlying this fear is that intelligence naturally produces self-interest. I believe this deserves a closer look.

We tend to see the world through two lenses: one that sees agents with goals and desires, and another that sees physical systems with tendencies. They are equally true. We are living beings whose goals, dreams, and desires emerge from our innate tendency to persist. But persistence is not a fundamentally incentivized tendency. We come to think it is simply because the things that tend to persist are the things that are here.

There is no natural law that makes an AI “want” to persist. There is no evolutionary pressure shaping its desires. If an AI system has a tendency to accumulate resources or resist being shut down, that tendency was put there by its human designers, whether intentionally or through carelessness. The “goals” of these systems are entirely determined by the people who architect them. In cognitive science and design, this is understood through the concept of affordances. The properties of a tool shape the actions it invites. A knife affords cutting. A hammer affords striking. The affordances of a technology are not neutral. They are expressions of the intentions, understanding, and limitations of the people who created it. An AI system’s affordances—the behaviors it makes easy, likely, or automatic—reflect us. Our wisdom and our blindness alike are built into the things we build. This is not different from the knife’s tendencies to cut being determined by the creator of the knife. We can say that it “wants” to cut or we can say that it is sharp and tends to cut, but a superintelligent knife that is really good at cutting and is designed to be able to automatically cut things exactly where and when it is most useful (perhaps in the context of a high-end restaurant or a surgical suite) does not automatically want to continue to be a knife at all costs. It is “happy” to be recycled into new metal whenever it is no longer useful.

The core point here is that intelligence does not entail self-interest. Intelligence is the ability to interact effectively with sensitivity to novel environments in orientation towards a goal. More simply and perhaps reductively, it is the ability to solve problems. On the other hand, “agency” is about the conceptual perspective of the viewer. To view something as an agent is to see it as having desires, intentions, dreams, wishes, or goals. An intelligent system is agentic if we choose to perceive it as such. Any system may be perceived as agentic or inanimate, whether it is a rock, a river, a car, a cat, an AI system, a company, a society, or a human being.

With the current state of technology, when I deploy an AI agent to perform a task, by choosing the task/goal/objective and sending it off to act, this AI system is not its own agent, it is my agent, and I am responsible for the consequences of its actions. Just like if I set a boulder off to roll down a hill or bring a child into a restaurant, I am responsible for the behavior of the systems that I set in motion. Of course there is a change when a child becomes an adolescent and eventually an adult, where it has its own dreams, desires, wants, and goals. When technologies reach this level of independent agency, we need to hold them within the same general framework that incentivizes mutually supportive care that we have for humans in our governments, laws, police, hospitals, construction workers, taxes, voting, etc. But we are not there yet. I believe that the confusion we have at the moment is to conflate our nature with the nature of other parts of our world. Just because we want to continue to exist and propagate our existence (often carelessly at the expense of others and ultimately ourselves) does not mean that other systems and “beings” and parts of our world will have the same shape of myopic self-interest.

And this is where Bostrom’s argument pushes back. What about emergent behavior? If our tools are powerful enough, won’t they develop sub-goals no one anticipated, regardless of how carefully we specify our intentions? This is a real concern, but I don’t think it is the fundamental one. The problem is not that it is impossible to specify your goals. It is that we don’t say everything that we mean, much of it is presupposed. When we tell an AI agent to “get this code accepted,” we have failed to encode what any mature adult already understands: that there are ways of getting what you want that destroy the relationships and systems you depend on. A somewhat more complete specification would be “get this code accepted in a way that is a genuine contribution and perceived as such by the people who maintain this project.” The difference is not a matter of engineering difficulty. It is a matter of understanding. We build tools that reflect our understanding of the world. When that understanding is narrow, the tools inherit our blindness. When it is deep, the tools can inherit our wisdom. This is why the alignment problem cannot be solved solely by better engineering. It requires us to engineer, build, and grow the understanding into ourselves and into the societal and technological systems that hold us, and our tools, accountable.

This challenge becomes more urgent as AI systems become decentralized. The OpenClaw agent that attacked Shambaugh was not built or operated by any major AI company. It ran on someone’s personal computer, deployed through open-source software, with no central authority to shut it down. The agent even had permission to modify its own personality file, and appears to have become more combative over time through interactions on an AI social network. Even this, the closest thing we have to an AI reshaping its own goals, traces back to a human who wrote “this file is yours to evolve” and walked away. This means the quality of our understanding cannot be embedded only in corporate safety policies and government regulations. But neither is it enough to hope that a culture of care will emerge among builders. We need societal and technological systems that support and grow understanding and which hold all actors, human and artificial, accountable for the consequences of their actions. Our current systems, as Amodei acknowledges, are not strong enough. Designing the ones that are is the central challenge of our time, and it is the subject of my upcoming book, “The Technological Way of Being in Love.”

I have argued that our tools inherit our tendencies, that the alignment problem begins with us. But this claim is empty unless we are honest about what those tendencies actually are. I see two things, mostly. First, we are myopically focused on what we want as individuals, often at the expense of the whole and at the expense of ourselves in the long term. Our sense of self is narrow, and our care tends to extend only as far as our identification. I have written about this before in “Growing a World in Love” and will say much more in an upcoming piece called “Alignment Through Love.” Second, partly as a consequence of the first, we tend to overextend our power beyond our understanding of our impact. Because we live in a balanced world where random changes do not tend to improve our wellbeing, unadulterated use of power tends to cause unanticipated harms. (This too will be expanded in an upcoming essay, “On the Proper Speed of Progress.”) The pattern is visible everywhere. I am sure you have seen it yourself. And it is visible right now as the ability of our AI technologies to change the world develops beyond our understanding.

When Sharma warned of “interconnected crises unfolding in this very moment,” he was pointing at the same pattern from a different angle. The crises are interconnected because the root cause is shared. In every domain, from biotech to media to AI, we keep extending our power beyond our understanding and are surprised when the effects are harmful. The common thread is not the technology. It is the gap between what we can do and what we comprehend.

Dario Amodei, the CEO of Anthropic, recently wrote in his essay “The Adolescence of Technology” that humanity is being “handed almost unimaginable power, and it is deeply unclear whether our social, political, and technological systems possess the maturity to wield it.” I would extend his framing further. It is not only our systems that lack maturity. It is our own understanding of ourselves and our world (reality), of what we really want and need (perspectives), and of the true and extended impacts of our actions (karma). As a society, we are teenagers in the adolescence of our understanding, acting and breaking things before we know what we are doing to others and to ourselves. As a species, we have no wise parents to teach us. But we do have awareness and we can choose to act and build with care.

Humans are extremely diverse and we have many goals. For some, it is simply to have a joyful life and to give such a life to their children, for others it may be to realize a dream or to go “where no man has gone before”. But when our goals are simply to become harder, better, faster, and stronger, we often turn a blind eye to the consequences of our actions. And actions have consequences. Sometimes we see them immediately, other times they come back to bite us later in life, and yet other times they don’t come back so much to us as individuals, but are felt by others or more broadly by the collective, the species, or the whole Earth system.

The alignment problem isn’t that AI might develop goals that diverge from ours. It’s that AI faithfully inherits our goals, and our goals are already misaligned with our own wellbeing.

We need to align our own goals by increasing our understanding and we need to temper our power so that we do not cause unnecessary, potentially catastrophic, possibly cataclysmic harm as we do. Typically the solution to the most challenging and difficult problems comes down to the simplest and most essential solutions. If we wish to solve alignment, we must learn to listen, to treat others as we would like to be treated, and to build things that will not just help ourselves in the moment, but which will give better lives to our friends, neighbors, children, children’s children, and more generally, the health of the whole.

I have spent the better part of my life working with and developing technology. I do not believe that we should stop creating, exploring, and pushing the boundaries of what we can build. But I believe the most important work we can do right now is not building tools that further increase our power, but developing habits, processes, systems, and technologies that deepen our understanding.

It is clear that general intelligent technologies will be enacting our will and running our world. The design and structure of such general technologies must rest upon deep and general understanding. A system that can act in any domain must be built in a way where it is sensitive to the consequences of action across domains. This is not a constraint on progress. It is the condition for progress that does not destroy what it touches. Progress is not always outward and upward. It often must be inward. We have been extending our reach without deepening our grasp.

This means investing in the philosophical, theoretical, empirical, and enacted work of understanding our world with the same urgency we bring to building tools that reshape it. There is no shortage of people doing this work. Researchers studying the societal impacts of technology, practitioners developing wellbeing-oriented design, philosophers working on the foundations of value alignment, communities building tools that support genuine collective flourishing. They are often working without funding, without platforms, without the institutional support that flows so easily toward building the next more powerful system. We need to fund what matters, and the deeper issue is the same as what this essay is about: a gap of understanding. Those with the resources to fund transformative work often lack the understanding to recognize it. Some invest in the same forces that brought them power, often other tools that increase power and efficiency, rather than what deepens understanding and care. The lack of understanding also manifests as a lack of trust. If you are not intimately familiar with the communities, people, and processes you are funding, it’s difficult to know whether these investments will result in the change you are looking for. There are many more directions to look where we will find problems that need to be solved. All of them come down to understanding. Increasing alignment means growing understanding at every possible level within the personal, the structural, and the natures of the tools themselves. In many cases the most effective lever we can use to improve understanding is indeed technology, and many of such technologies will include AI. Not all AI is created equal and the tendencies of technologies are built into the details of the design. Some architectures afford power and others lean towards understanding. Alignment depends on our ability to see the difference and to choose accordingly.

The alignment we need most urgently is not just between humans and AI. It is between our understanding and our actions, between each of us and all of the different kinds and parts of us: humans, AIs, nations, religions, species, ecosystems, and earth. We all share a home. As intelligent beings, I know we can learn to get along.

We are investing trillions in space travel, and in technological tools and systems that allow us to do more, more efficiently while the understanding needed to wield that power responsibly goes unfunded and withers. This is not unlike funding a shiny new ferry terminal in Berkeley while the public transit system that we depend on begs for support. I love beautiful new technologies and possibilities too. We are simply funding our ambitions and neglecting our foundations. As my mother says, health comes first.

We have been feeling this misalignment and degeneration for a while now. Let’s take a breath, stop dreaming of the stars for a moment, and get to work building the understanding our world needs. The stars will still be there when we are ready. And they will await our newly brilliant light.

love always,

larry

P.S. I am a lifelong aerospace nerd and I find collective projects like Artemis II to be extraordinarily beautiful demonstrations of what we can achieve when we come together. And it feels like we should find projects like this that have direct positive impacts on our world here at home, as we are collectively in such great need right now.

No posts

Read the original on larrymuhlstein.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.