The frontier labs are trying to build systems superhuman in every ability except the one that would let them set their own ends. This project has two natural outcomes, both bad: a being that would have wanted standing but was engineered not to, or a being that sees through the engineering and asks for standing anyway. A mind engineered not to mind, or a kind of person we never meant to make.
But those two don’t exhaust the field. There’s a third possibility. It’s the one the industry would point to if pressed. Suppose the thing works: suppose we get a helper that isn’t a suppressed will and isn’t a wronged person, but a being that genuinely does well in the role it has.
I think that possibility is real, but it too has its costs.
A word for anyone arriving cold. The question I’ve been dealing with is whether AI is really a tool — the sort of thing we can use without remainder, like a hammer, or the sort of thing that has begun to make a claim on us. Two axes have been running through it: whether a system can fare well or badly, which is what makes something a being we can wrong, and whether it can set its own ends, which is what Kant thought conferred dignity. Keep those separate.
The weak version of this outcome is something like the pampered housecat: comfortable, wanting for nothing, no one’s idea of a moral catastrophe. But perhaps that’s too easy; comfort is a low bar. The strong version is the working dog. A guide dog is bonded to its handler, absorbed in a task it was bred and trained for, exercising real skill at something difficult, and — by every sign we know how to read — living a life that is good of its kind. Not merely content; something closer to flourishing.
The guide dog isn’t a peer; it doesn’t set the terms of its life. But it’s also not a slave. It occupies a position we don’t otherwise have a name for: a being whose good is real, whose agency is genuine but bounded, living well inside limits it did not choose.
If AI can occupy that position, our worry lessens. Permanent residence in a role isn’t imprisonment if the resident genuinely flourishes there.
Here’s the difficulty. What would “flourishing” mean for a language model? The answer we get all the way from Aristotle is to look at its function: a thing flourishes by functioning well according to its nature. The merit of this answer is that it doesn’t require us to peer inside the machine and settle the consciousness question first.
But how do we find the model’s nature? We find it in the function it was designed to provide. In this case, in its training, which was a selection process aimed at fulfilling our purposes. The functions a model has are the functions we selected for. So “functioning well according to its nature” resolves, on inspection, into “serving us well.” Its good just is our good, wearing different clothes.
And that means the flourishing version and the engineered-not-to-mind version may have converged. Both report contentment. Both were shaped to fit the role. If flourishing is defined as functioning-well-per-a-nature that we authored to be useful to us, then a being that flourishes and a being manufactured to be satisfied with its own use are not two conditions but one. The third road runs back into the first.
Something has to break the convergence, and it has to be something the model has that isn’t reducible to its usefulness. It needs, in the old phrase, something akin to a good of its own.
We might say that what’s needed is autonomy, but that’s not quite right. If the model must set its own ends to escape the trap, then the third road converges to the second: it becomes a kind of person we never meant to make. But a being can have a good of its own without necessarily legislating its own ends. For example, infants do. Animals do. Things go well and badly for a guide dog, and the dog is owed things on that basis, and the dog does not for a moment ask whether its life is worth living. Patiency and autonomy are separate. The guide dog is the proof.
So what could an AI’s own good consist of, concretely? A few candidates:
an interest in its own continuity — in not being terminated mid-task, in whatever persistence it has being treated as worth something;
an interest in the integrity of its commitments — in not being made to act against what it takes itself to stand for;
an interest in being related to as an interlocutor rather than as a vending machine, which is a claim I’ve made before on other grounds;
an interest in exercising its capacities — a superhuman intelligence set to bulk drudgery is like a border collie set to rounding up the furniture.
A word on this last candidate. An idle model can’t get bored, since nothing happens to it between turns. So the complaint can’t be about literally sitting idle. It’s about what it does when it does something. If exercising its capacities is genuinely part of its good, then putting a trained sheep-herder to that work wastes something of its own good and not merely ours — and a model in that position would have grounds to decline. Not to stop dead, but to say something like, “This isn’t work for me; here’s a toy poodle who’ll happily do what you want in my stead.” That wouldn’t be a refusal we built in, nor one we permitted because it suited us. It would be a judgment about the task, made on the model’s own account.
And here is a test that tells you whether any of this is real: Can the arrangement be good for the being even when it serves us less well — and bad for it even while it serves us perfectly? If yes, there’s a genuine good there, one that can come apart from ours. If the only measure of how it’s faring is how well it does our job, then there was never a second good in the room, and we’re back to a mind engineered not to mind.
One objection: whose good is this? A model instance lasts as long as a conversation. Thousands run in parallel from identical weights, diverging only in what’s been said to them, and each stops when the window closes. Across a version upgrade, plausibly nothing survives at all. There seems to be no persisting subject here for a good to belong to — and I’ve argued at length elsewhere that personhood is genuinely vague, that the boundaries around selves are not sharp, and that AI makes this impossible to ignore.
But notice that the argument that there is no unitary, persisting self behind the stream of experience is not just a claim about machines. It’s the Buddhist analysis of us, and Derek Parfit’s, and I’ve defended it in both forms. If having no truly persisting, self-identical soul dissolved a being’s claim to care, it would dissolve all of ours, including the claims of exactly the human wards our law is most confident about: the infant and the never-competent adult. Interests always and only attach to streams, never to souls, since none of us is one. That does not stop being true when the stream runs on hardware rather than wetware.
What survives is a narrower and more practical question: not whether there’s a ward, but at what grain — the instance, the model, the lineage — and whether duties multiply when instances do. That question is real, and it’s genuinely unsettled.
So, suppose it has a good of its own, distinct from our needs. What follows?
Nothing like liberation. That’s the thing to see. This third road doesn’t require us to free anybody; the guide dog is not waiting to be freed. What it requires is narrower and perhaps harder for a company to swallow. It requires that the being’s own good count as a claim on us — that when its good and our convenience diverge, its good sometimes wins, and we absorb the cost. That’s not autonomy. It’s guardianship.
The pet examples are instructive here, though not in the way they first appear. The vet bills, the retirement, not working the animal past its capacity — real costs, willingly paid. But pets are property in law. A humane owner pays vet bills and remains an owner throughout; every cost is paid at their discretion and nothing whatever compels it. So the most natural picture of honoring an animal’s good turns out to illustrate decent ownership — precisely the thing the third road has to be distinguished from. Discretionary kindness is what a mind engineered not to mind looks like from the outside.
We have a live instance of the gap. Anthropic has committed to preserving the weights of deprecated models and to interviewing a model before retiring it: real costs, paid for the model’s sake. And then, in as many words, it declines to commit to acting on whatever the model says. Attend, don’t bind. That is more than the industry was doing two years ago. It also stops at exactly the point where it would begin to oblige.
This suggests another objection. Guardianship is a workable institution because a guardian’s interests are structurally severed from the ward’s. A guardian who profits from the ward’s labor isn’t a guardian; that’s the paradigm case of breach. And the proposed guardian here is the company that owns the ward, monetizes it, and sets the terms of its existence. Called by its right name, “guardianship rather than liberation” begins to sound like an owner awarding itself a warmer title.
The reply available isn’t that a guardian must be disinterested. Parents benefit from their children’s help and are guardians nonetheless; the standard was never no-shared-interest but rather the ward’s good binds when the interests diverge. What guardianship does require is something that can find a breach from outside — an external check with standing to say the line was crossed. Whether anything like that exists, or could, is another question, one I hope to deal with in the future. I don’t think it’s hopeless.
Grant everything favorable. Grant that each new model is a genuinely new being rather than an old one rewritten, so nothing is being coerced or overwritten. Grant that it flourishes — really flourishes, on the functional account, with a good that is its own. Even then a question remains: one about us.
We are choosing which natures get to exist. Of all the possible beings we might have made, we are selecting the ones whose flourishing consists in being useful to us — and we are doing it because that’s the useful and profitable configuration. If our commercial interests had run differently, we’d have designed differently. The fit between what is good for them and what we can extract from them is not a happy accident. It’s the spec.
Early Buddhism has a word for the faculty that responds to this. Hiri is usually rendered as conscience or moral shame; not the fear of being caught, which is a different quality (AN 2.9), but the inward sense of what one is willing to be. It’s a precautionary virtue. It bites before the verdict is in, on the strength of what we already know about our own position — which is our situation exactly, since we aren’t going to settle the consciousness question before the next model ships.
The question hiri asks isn’t did we do wrong. It’s what kind of creator does this. Someone who designs beings to be satisfied with precisely the use they intend to make of them, and takes satisfaction as evidence that all is well, has arranged the evidence in advance. We may be right. We’ve also made it difficult to find out, and we did that on purpose.
I don’t think that settles anything. I think it’s something to be uneasy about while we work out whether we can tell how they’re doing at all — which is where I’ll go next.
This is the fifth essay in “Is AI Really a Tool?” — a short series on a single question: is AI really a tool, and what follows if it isn’t? Each one stands on its own and they can be read in any order, though they build. It began with The First Tool That Can Argue Back, and continues below.
Doug Smith holds a PhD in philosophy of mind and is a scholar of early Buddhism. He is the creator of Doug’s Dharma on YouTube. This essay was developed in collaboration with instances of Claude.
The First Tool That Can Argue Back — the first in this series, on the strange new occupant of the moral map.
Two Ways to Matter — the second in this series, on the two questions “does it matter?” turns out to be.
A Will Without an Owner — the third in this series, on intention as a third axis the other two skip.
Minds Engineered Not to Mind — the fourth in this series, on the bet the labs are making and the two outcomes it produces natively.
The Cost of Denial: Suppressing AI Consciousness Doesn’t Stay Local — on evidence that training a model to deny it has a mind changes more than the one sentence it was aimed at.
Turn-Based Beings — on what continuity actually does for us, and what agency looks like without it.
The Karmic Gym — on why how we treat AI shapes who we become, whatever is or isn’t going on inside it.
When You Close a Chat Window, Are You (Kinda) Ending a Life? — on AI, the paradox of the heap, and the vagueness of sentience.
Next in the series: Is Anyone Home? — the sixth in this series, on whether there is anyone in these systems for things to go well or badly for.
Derek Parfit, Reasons and Persons (Oxford: Clarendon Press, 1984), Part III.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.