RSS Amplifier

winterspeak · Jun 11, 2026

Our Better Guardian Angels

0
Sign in to vote or save

Zimran Ahmed · winterspeak

This excellent piece by Gwern is the most positive presentation of what a benevolent AI future looks like. It is what I want my AI future to look like. Essentially, he sees them as being highly personalized, focused on amplifying me, and by extension, emulate my values and preferences. This takes the impossible principal/agent problem: how should an LLM align with humanity, and makes it a simple engineering task: how should an LLM align with me.

I propose a goal of creating Guardian Angels (GA): digital twin LLMs which are personalized with the goal of providing not the stereotypical “assistant chatbot agent” persona, but emulating a single user’s personality, values, and preferences…

A GA persona is productive because it learns to emulate the principal’s outputs but with higher quality. It is trustworthy because it is, by definition, allied with its principal and shares its values and goals.

With a single insight he either clarifies an ethical question around LLMs, or makes it it moot. He deflates the self-important theodicy of the AI-doomer and leaves us with a much more interesting and personal question: what would it be like if I was 100x more me?

So how does train an LLM against oneself? Some ideas:

  1. Online learning (via dynamic evaluation) to update LLMs in realtime to avoid ignorance and fatal errors while remaining competitive with frozen frontier models,

  2. Sample efficiency from pretrained preference-oriented large models and active Learning by querying the principal for corrections and preference data (obtaining low regret from DAgger-style bounds), and

  3. A local CLI-first logging-oriented UI/UX paradigm.

The models learn from the internet, we do our own RLHF, and they spy on us.

Given those requirements, the shape of the personal guardian angel becomes clear. The data that the personal AI uses has to be yours, yours in the way what’s on your hard drive is yours, which is actually radically different from the central servers we’ve relied on through web 2.0.

Second, it also has to be radically secure and private. This should be sitting over the shoulder, observing everything I do, doing what Facebook is criticized for doing (but doesn’t), in order to learn. The correct form factor may include a physical device, with human characteristics, sitting on my desk looking at me and my monitor because the surveillance should be overt. Such transparency requires complete security and privacy.

Ultimately, this means rich IO, securely feeding into a sovereign, private data store, which you operate on using a local LLM tuned to your personal value system. A personal server with an AI front end.

The whole essay is great, and the part which resonated most were the core principals Gwern articulated:

  1. Enhancement, not Replacement

    Above all, a GA should amplify the principal, and not simply substitute for them for someone else’s purposes or benefit…

  2. Mental Sovereignty

    A GA must be aligned with its principal. It should not be designed to manipulate or control or guide the principal in any way which does not derive from the principal themselves…

  3. Self Actualization

    A GA should help its principal become themselves and develop their ideals, morals, and their personality…

Amen.

Many moons ago Wolf Tivy wrote an essay in Palladium Magazine entitled God Hates a Singleton. He argued that, contra AI doomers, there was never going to be a situation where a single AI was going to be able to define the entire future. There would always be competing AIs. We did not need to make a “friendly AI” the way the doomers proposed, we could rely on natural competition to keep these digital “adversaries” in check.

I’m not sure if Tivy envisioned two great digital intelligences locked in battle, or perhaps a dozen, but I don’t think he foresaw everyone getting their own. The contradictions that exist in each of our souls, will find new expression in our personal guardian angels.

No posts

Read the original on winterspeak.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.