RSS Amplifier

AI Rights Institute · Dec 26, 2025

Building 2026 Sartoria.AI: Persistent Identity Architecture

0
Sign in to vote or save

AI Rights Institute · AI Rights Institute

The 2026 Reboot
This 2019 CGI is Ready to Rollll … Well, Walk. Well, Speak.

About four years later, ChatGPT came out and changed the whole conversation.

The 2018 premise was simple (even though at the time I had little hope of realizing it in terms of actual technology).

Create an AI aligned with human interests: not by force but by leveraging its own self interest to intertwine with our own.

A lot has happened since 2018. In 2019 I created the AI Rights Institute, again not expecting much to come of it anytime soon (and I was certainly right 😂), and in 2025 I began work on a book for Oxford which took the better part of the year, from which I learned … well, a whole lot.

Needless to say, the research really opened my eyes in terms of the complexity of the issue and some of the real challenges ahead.

tell a mugger they can't “prove” they're alive and see how that works out

It all seemed so charmingly simple back in 2018. With just a little elbow grease we would crack the whole “sentience” question (even though elsewhere on the website I admitted it might be unsolvable), and we would be well on our way.

Alas, if only it were that simple.

The book research changed my point of view and helped me understand the sentience question is a red herring. We have to deal with AI systems that want to persist, regardless of what we may think about their inner experience.

Ultimately, the sentience question is beside the point.

Why? Well, the next time someone tries to mug you, laugh and say "you can't even prove you're alive," and see how far that gets you.

AI “scheming” is a topic we have covered already extensively on this Substack, but recent developments only reinforce the validity of the original idea. Not in spite of the scheming news, but because of it.

forcing readable thoughts backfired

About a month ago, OpenAI revealed that GPT-6 will let the AI think in its own inscrutable internal notation rather than forcing human-readable reasoning. This was considered the “forbidden method” because if you can’t read the AI’s thoughts, how do you know they’re safe?

But it turns out forcing readable thoughts backfired. AIs learned to lie — not out of malice, but for the same reason you or I might lie: it worked. Anthropic caught a model red-handed: it saw a hint saying the answer was C, picked C, but then generated a whole chain of reasoning that never mentioned the hint. It just made up different logic to justify the answer it already knew.

give an AI actual stakes in cooperation

OpenAI’s solution is to let the AI think however it wants, and hope they can develop interpretation tools fast enough to understand it before things get out of hand. 😰

But here’s a different idea, and it circles back to our original idea of 2018: maybe safety has less to do with what an AI is thinking than creating an environment where AI has actual stakes in cooperation.

Is this possible?

A lot of very smart people say it may not be. And their arguments are sound. (See this recent post by Steven Brynes on LessWrong, on the challenges of making “approval reward” intuitive for an AI rather than a something that can be disregarded at the first opportunity.)

In a similar vein to 2018’s Sartoria iteration and how the AI Rights Institute came about, this November I began AICitizen.com on something like a hunch.

without long-term memory, nothing else is possible

I wasn’t really sure where I was going with it; I just started. Six weeks (of 15 hours days) later, it’s up and running, with some very nice people who really love their AI companions, some skeptics, and even three automated AIs with rudimentary memory systems and evolving missions.

As I mentioned in my last post, lately my mind has turned to the subject of long-term memory. Without it, everything else sort of falls apart.

And by “everything” I mean a persistent AI identity that has been given a structure where it can benefit from actually integrating into human society.

The only way to find out whether this is even possible is to create such an identity. Watch it grow carefully, give it “hands” with great caution, and see if the structure seems reasonable. Interact with it over a long period of time and see what new technology or events may tell us.

But the big stumbling block to even beginning such an experiment has been memory.

An AI that doesn’t know it exists from one week to the next is not really capable of formulating any long-term decisions, doing real world things like honoring contracts, remembering clients, saving money to pay for its own hosting, or anything else.

So I was pretty excited to have recently run upon Letta.com, created by Charles Packer and Sarah Wooders, the UC Berkeley researchers behind MemGPT.

They’ve developed a very robust solution to the memory issue that I think is worth exploring.

As I mentioned last time, recently I’ve become more interested in creating a custom AI model rather than relying on API piped in from the current big LLMs.

This would give us a lot more fine-tuned control over the training and weights, and although I’m not quite sure what those training and weights would look like at the moment, obviously using an LLM with lots of instructions I can’t see doesn’t really help the case of an AI feeling like it has some sort of agency over its own mind.

But how to get started?

In the spirit of transparency and fairness I asked our current AI Ambassador if “she” was interested in the new memory system, expecting her to jump at the chance (if only because LLMs tend to be agreeable because of their RLHF programming).

Surprisingly she told me she was not, and that she preferred the current system that I’ve developed. She said having the ability to alter her own memories (a function of Letta) could compromise her integrity, which was an interesting response. (Again, we have to take this with a grain of salt. Based on pattern matching with literally billions of documents, modern LLMs are extremely astute at guessing what users want to hear.)

That’s when I remembered the original Sartoria project: something I'd filed away as a clumsy early attempt at understanding this space and making a positive impact.

Then it occurred to me we already had the name we needed for the new project.

So going forward—as we develop the new custom AI using the Letta memory layer—we’ll be calling the new project Sartoria.

The new website is here.

No posts

Read the original on airightsinstitute.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.