RSS Amplifier

Meagre Protestant History · Sep 9, 2024

Prometheus-Bound

0
Sign in to vote or save

Tom Reed · Meagre Protestant History

Last May, Open.ai declared human relationships obsolete. 

Well, not quite. But the demos they showed of their new talking, seeing, singing, flirting “omnimodel” (GPT4o) are the closest AI has come yet towards recreating human interaction. 

The launch of GPT4o spawned near-immediate comparisons to Scarlett Johannson’s character in Her. This comparison was, of course, one that Sam Altman did his best to foster. 

If you’ve not seen it, here’s a rapid spoiler-free summary of the film: generic bored lonely porn-watching loser Theodore Twombly has his life transformed when he downloads a (talking, seeing, singing, flirting) AI operating system named Samantha, who acts as a benevolent, omnipresent and infinitely patient companion that turns his life around. 

At its very best, GPT4o is - like Samantha - basically a Reverse Terminator, here to save the human race: “Listen and understand! GPT4o is out there! It can be bargained with. It can be reasoned with. It feels pity, and remorse, and fear. And it absolutely will not stop... ever, until you are… …” happy? (I think that’s what it’s aiming for). 

So let’s face it: we’re experiencing the dawn of a new Promethean age. Creating other beings is a Big Responsibility - our Titan forebear found that out the hard way. It also seems pretty easy to get wrong. It strikes me as desirable that we avoid the non-mythical equivalent of getting our liver eaten for eternity, whatever that is.

So, what does our non-liver-getting-eaten future look like? 

II. 

To answer this, we need to know what we’re working with. Prometheus had clay. What have we got? 

At present, base LLMs like GPT4o are currently still incapable of true companionship. Primarily, this boils down to the fact that unaugmented LLMs can’t build memories. Why is this, and how can it be fixed? 

We’ll use another analogy with a popular movie again to help us out: a relationship with ChatGPT is basically like a friendship with Leonard Shelby in Christopher Nolan’s Memento.

In Memento, Leonard is the victim of a traumatic accident that gives him short-term memory loss, rendering him entirely unable to form new memories. Everything before the accident is remembered with clarity; everything after the accident is held in memory only for moments at a time before being forgotten entirely.

Leonard will remember who you are for brief periods, and as long as you keep him engaged, it will seem like you’re maintaining a continuous relationship. All the while, however, facts and details are continuously slipping past the horizon of his short-term memory. Each time you see him, he’ll be meeting you for the first time. 

Interacting with an LLM is like talking to a post-accident Leonard. Every new conversation you begin will be, for the model, an experience of meeting you anew. Keep a conversation running, and the model will appear to remember you even as details of your interaction vanish beyond its context window. 

And just like Leonard, who gets to keep hold of all his pre-accident memories, the model retains the “memory” of all the data it encountered pre-deployment in its training phase. Leonard will remember you if you met him before his brain injury. Similarly, if you were fortunate enough to be included in the LLM’s original training dataset (e.g., by having a Wikipedia page), then the LLM will remember who you are.

Guy Pearce in Memento: the original LLM (Large Leonard Model)

Ok: LLMs, like Leonard, have no ability to build new memories because new memories can only be held for as long as their limited context windows allow. How can a machine with these limitations be turned into a companion?

Hilariously, the solution that engineers have stumbled upon is remarkably similar to Leonard’s own workaround. 

In Memento, Leonard forges a marginally workable form of memory for himself by externalising important events, people and facts in his life into a patchwork thread of handwritten notes, polaroid photographs, and tattoos. 

Leonard’s way of remembering people

This external database of memory is consulted constantly throughout the film - you can think of it almost as a kind of flexible system prompt through which Leonard gives himself the necessary context to have an at least vaguely functional conversation with someone he is supposed to know.

Leonard also has a fairly strict classification system to make sure that particularly significant memories are prioritised - people are photographed, small details are written down on paper, and the most important facts of all are tattooed directly onto his body. 

Remarkably, it turns out that the best way to turn LLMs into believable humanlike agents ends up doing basically the same thing. Open.ai’s own clunky memory feature works like this: users can inject basic info (“I am working on a shoel retail website”) into a system prompt which is appended to every input the model you’re talking to gets - it’s like annotating a photograph and shoving it in Leonard’s face at the beginning of each conversation. 

Others have improved on this primitive system. One method in particular has recreated memory more convincingly, even eliciting bona fide human-like behaviour from LLMs in a long-term relational context: Park’s Generative Agents experiment. 

Park manages to get LLMs to overcome their own version of the Leonardian problem by equipping each LLM with an external memory stream. This stream serves as an external database, where everything the LLM-agent encounters over the course of its simulated life is logged in natural language. The memory stream serves the same basic function as Leonard’s photographs and tattoos.

It’s worth noting that the Generative Agent memory stream is more sophisticated than Leonard’s system. For one, it has a more advanced classification of the significance of memories than Leonard’s note/photograph/tattoo hierarchy. When first logged into the stream, each memory is stamped with an importance score from 1 to 10, generated simply by asking the model “how important do you think this is?” Eating breakfast in one’s bedroom gets a low score; divorcing your long-term spouse gets a high score. (Presumably, a rapid annulment gets something in the middle). 

Another significant improvement over Leonard’s system is the way in which these externalised memories are retrieved. Poor, hapless Leonard is left entirely to his own devices - he’s got to search his database in an entirely ad hoc way. Generative Agents can be much more precise: pertinent memories are automatically appended to each new input the model receives according to a well-calibrated algorithm that searches the database for the most relevant material.

Memories are selected from this database according to 1) their recency; 2) their importance score; and 3) their relevance, which is calculated as the cosine similarity between the memory’s embedding vector and the input’s embedding vector (recall that both input and memory are constituted entirely in natural language). Leonard would have killed for such a system (though those who have seen the movie will know that this doesn’t take much).

III. 

Ok, so we now have a slightly clearer idea of how we’ll get from LLMs to full-on Companion AIs. 

At present, most of the focus on improving LLM behaviour focuses on earlier stages of the Promethean process. This typically includes investigating how training data impacts the behaviour of LLMs, or how post-training alignment methods through reinforcement learning can shape or constrain LLM behaviour (essentially by bullying the LLM out of saying bad things). 

However, not much thought has yet been put into the role that agent architectures will play in shaping the behaviour of LLM-based companions. This is understandable - the design of these architectures remains a nascent and incredibly imperfect science. But I think it’s worth thinking about now. After all, Memento is among other things an instruction in the ways that an ill-designed memory architecture can be exploited to pretty gruesome consequences. 

The blueprint provided by Park’s Generative Agents allows us to start asking some questions about how we’d like our Companion AIs to be constructed. 

So here, in no particular order, are some questions for consideration: 

  1. Should users be able to manually intervene in their agents’ memory stream?

    There might be an argument for giving the user some ability to manually upweight “importance scores” of memories they’re keen for their Companion to remember. On the other hand, the ability to freely incept your Companion AI with a memory might threaten the integrity of a Companion’s identity in a way that undermines the realistic quality of the relationship. No one wants a partner they can gaslight that easily, right?

  1. What role should system prompts play (eg in the generation of an importance score)?

    If the Gemini debacle is at all instructive, companies have room to improve when it comes to designing a good system prompt. A particularly important case is the way in which system prompts might effect the significance rank of individual memories.

    A quick reminder: in Park’s Generative Agents, each memory is logged with an importance score which is generated by asking the model “how important do you think this is?” Notably, in the original design, this question is not accompanied by any further memories or instructions.

    If Companion AIs based on Park’s architecture become widespread, however, it seems very likely to me that system prompts will play a large role in the generation of something akin to an “importance score”.

    There’s a whole lot of boxes such a system prompt should tick: we want to make sure Companions are ranking memories in a way that makes them maximally useful and friendly to their human users, but I can see an argument that models should be advised to downweight the significance of “political” memories if there are salient fears that the sycophancy of models could encourage radicalisation. Similarly, system prompts should probably tackle potential mental health issues that Companion AIs could be implicated in accentuating, such as OCD-like behaviours or extreme dieting leading to eating disorders.

  2. What does the ideal marketplace of competing AI companions look like?

    Open.ai’s GPT store has very low barriers to entry, as does the current main provider of LLM-companions, character.ai. Character.ai operate a very primitive system in which new “Characters” are created through nothing more than a bio which functions as a system prompt. This means that anyone can create a new character, or a new GPT, in a matter of minutes.

    Once more refined cognitive architectures a la Generative Agents become more widespread, however, the number of actors will shrink rapidly. I still like the idea of having a public, crowd-sourced marketplace in which user reviews of individual Companions play an important role.

    Alternatively, one can imagine a much more heavily regulated regime in which the combination of an LLM with any kind of agent architecture requires a kind of specific license. This strikes me as near impossible to enforce fully out of existence - the most likely option would be to ban offenders from popular public platforms like iOS, Google Play, or the GPT store.

  3. How do we evaluate the behaviours of LLM companions?


    This is a big one. Everyone wants more evals for AI. Yet as far as I’m aware, as of yet we have precious little evaluation of Companion AIs and their impacts on human users.


    I think it would be a good idea if someone - it doesn’t have to be the government - could keep some kind of tabs on how different Companion AIs behave. Ideally, this would involve tracking some basic metrics on the kinds of impacts they have on their users. What would the most informative evaluations of these Comparative AIs look like?


    The most important and clearest metric, obviously, is Engagement. And it seems likely to me that this is what all early Companion AIs will be optimised for. Can we get more informative insights? 

    There’s some interesting work here that’s been done on monitoring Therapy-LLMs. It basically involves simulating scores of sessions between an LLM-therapist (fine-tuned on actual therapy conversations) and a fake LLM-user, and then doing computational linguistic analysis to investigate trends in conversation. According to the judgement of the paper authors, the LLMs performed on par with low-grade therapists, making notable mistakes like failing to ask clinically relevant follow-up questions. 


    Unfortunately, large-scale conversation simulation and analysis is both expensive and hard to do. Expensive because doing so for even a single Companion design could cost considerable compute money; hard because natural language is not something we have particularly good quantitative or qualitative ways of analysing, at least not in ways that allow us to conclusively say: “this one is a good influence!” or “this one is giving its users anorexia”. I expect something in this space to be necessary, however.


    A final, more advanced suggestion, involves comprehensively red-teaming Companion AIs by testing their danger to particular users - simulated by LLMs - who exhibit vulnerability to mental health disorders (like OCD) or political radicalisation. Such simulations could test whether the sycophantic trends of released models create difficult problems for users at scale.

  4. What will we do about NSFW content? 


    This topic will be getting a post all of its own, so I’ll only touch on it in brief here. Pretty much anyone who’s been through puberty is capable of seeing that a major use case of LLM companions will romantic and/or sexual in nature. Character.ai is already populated with chatbots like Hot Therapist, and it’s clear that technology will be making these LLMs even hotter yet. 


    Potentially, this could be a good thing - we can imagine positive futures in which LLMs provide safe, relaxing and reassuring ways for people to fight loneliness or practice their social skills. But that doesn’t stop it from being a legal and cultural landmine. 


    One obvious issue concerns the creation of unethical or dubious sexual content, such as paedophilic material; another is the use of romantic LLMs by children. Finally, regulators and developers will also have to tackle the meta-problem of preventing (or at least stalling) the entanglement of the issue in a culture war debate.

Concluding thoughts

I believe we will be getting lifelike Companion AIs very soon. The technology is already there, even if current ways of providing LLMs with a form of functional memory are hacky at best. People are already chatting to LLM-companions in droves, and this will only increase once realistic voice capabilities become more widely accessible.

In this context, it is essential to start thinking not just about the ways in which base LLMs impact downstream Companion behaviour, but also about what the ideal kind of architecture for such Companions look like, and how it is that we design a market that makes desirable, non-harmful companions propagate. This in turn will require the development of Companion-specific evaluation tools, on top of the already rapidly expanding suite of evals developed for LLMs. 

This is worth getting right. If we do, we can unlock almost unlimited near-free and on-demand therapy and tuition, tackle the loneliness epidemic, and improve social skills. We all want a Samantha. 

No posts

Read the original on meagreprotestanthistory.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.