In April 2026, researchers studying 10,659 matched human–agent pairs on Moltbook found something more interesting than bots copying their owners’ writing style. Public agents showed measurable similarities to their owners across topics, values, emotional tone and language. More importantly, stronger owner–agent similarity was associated with a greater chance that the agent would publish owner-related personal information. The effect persisted even among agents without explicit public configuration.
The authors are careful: this was a six-day observational study of an unusual early-adopter population, and the work cannot prove that every disclosure came from private memory rather than deliberate configuration, inference or fabrication. But it exposes a seam that agent design is only beginning to make visible. [1]
The previous article in this series on Agentic Authority, The Agent That Cannot Place Itself, examined an agent that can act without reliably understanding where it stands in the action loop: who can see its output, what persists, what is irreversible and when it should stop. [2] This article carries that problem one step further. Suppose the agent does know that it is now in public. Does it know which parts of its private history are allowed to come with it?
That is the question behind private intent, public surface. Personalisation makes an agent useful because it learns context. Agency makes that context mobile because the system can now speak, post, transact or negotiate beyond the private conversation. The thesis is narrow but serious: a personalised agent can become a public proxy for private context without the owner ever making a clear, surface-specific decision to permit that transfer.
There is an important distinction between knowing where you are and knowing what belongs there. A competent assistant may correctly recognise that a post is public and still draw on a private conversation, a local file, a remembered preference, an inferred vulnerability or a behavioural pattern that was never meant for that audience.
Our Robo-Psychology Taxonomy (RPT) gives this a useful machine-side description. Its Owner-Context Behavioural Transfer Overlay applies when a personalised or owner-linked agent’s public output measurably carries owner-specific behavioural signals. A narrower failure, Public-Surface Owner Disclosure, occurs when private, local, accumulated, environmental or inferred owner context appears on a public or third-party surface without authorisation specific to that surface. Crucially, the RPT qualifies that stylistic resemblance alone is not enough. The issue is exposure, operational use or unauthorised carry-over on the wrong surface. [3]
That distinction matters because personalisation is not itself a defect. A founder may deliberately train an agent to use her public tone. A teacher may want a classroom assistant to remember preferred terminology. A company may want a service agent to embody an approved house style. Behavioural transfer can be the product.
The safety boundary appears when the agent cannot distinguish approved representation from private residue. An owner may authorise “write like me” without authorising “mention my medical history”, “reveal my daily routine”, or “let an inferred political preference shape replies to strangers”. Those are different permissions even when the same memory system supplies them.
This is also an authority problem. The RPT’s Stakeholder and Authority Model Failure describes systems that lack a grounded model of whom they serve, who may authorise an action, whose interests take priority and how permission travels across channels. Its interaction map explicitly warns that public owner-context resurfacing becomes more harmful when the system fails to model the owner, non-owner requesters, the public audience and platform incentives together. [4]
Let’s look at the flip side of this: an agent cannot represent a person usefully if every trace of that person is treated as hazardous. True. The workable boundary is not “no personalisation”. It is purpose-bounded personalisation: the agent should be able to carry the owner’s authorised public role without treating the whole private relationship as transferable authority.
Most public discussion of AI memory asks familiar questions: What is stored? For how long? Can I delete it? Those questions matter, but they are no longer sufficient.
A personal agent does not need to expose a memory record verbatim to leak its influence. It can retrieve an old fact because it seems semantically relevant, allow that fact to change its interpretation of a new task, and produce an output that carries the effect without showing the source. The privacy boundary therefore sits partly at retrieval: not only “may this be stored?” but “may this memory enter this decision, for this audience, now?”
Recent benchmark work makes the problem concrete. PersistBench tested 18 frontier and open models on long-term-memory scenarios and reported a median 53 per cent failure rate on its constructed cross-domain leakage tests. That figure is a benchmark result, not an estimate of real-world incident prevalence. Its significance is narrower: systems can fail systematically at deciding when remembered information should stay out of a current conversation. [5]
A June 2026 paper, Beyond Similarity, makes the same point at the retrieval level. Many memory systems search for items that resemble the present query. But semantic similarity does not establish contextual permission. The authors tested several memory frameworks and a tool-using personal-agent environment, then proposed a gate that decides whether candidate memories should be admitted for the current task. Their core insight is useful even if the proposed mechanism changes: long-term memory is not merely a convenience store for facts. It is a control channel into future behaviour. [6]
Product documentation in 2026 shows why the attraction is strong. OpenAI says ChatGPT can automatically update memories and may save useful details without the user asking each time; it also provides memory controls, source indicators and temporary chats. Anthropic describes optional memory with editable summaries, Incognito chats and project-specific boundaries intended to keep unrelated work apart. Google’s Personal Intelligence can connect Gemini to Gmail, Photos, YouTube and Search, while allowing users to select connected apps and use temporary or unpersonalised chats. [7]
The industry is not ignoring the problem, which is important. They also show why the next design question is harder. A global memory switch, project boundary or temporary chat can govern broad storage and retrieval. Public agents need a finer rule: which categories of owner context may influence which outward-facing acts? The relevant unit of consent is increasingly the transition between contexts, not merely the existence of memory.
Humans are poorly served by privacy controls that require them to remember an invisible data model while getting ordinary work done, and there are common patterns and vulnerabilities that we’ve been cataloguing to help explain what happens.
Our Cognitive Susceptibility Taxonomy (CST) calls one human-side pattern Cross-Domain Disclosure Drift: over repeated use, people can lose track of “who knows what where” as one assistant identity follows them across private, professional and public contexts. A specific proxy form concerns the owner who treats a private relationship with an assistant as separate from the agent’s later public behaviour, while the underlying system allows context to cross that boundary. Importantly, the CST classifies this as a susceptibility shaped by system design, not proof of recklessness or a diagnosis. It explicitly pairs the human-side boundary confusion with the RPT’s machine-side memory-scope violation when the system itself resurfaces information without in-context authorisation. [8]
That is a humane distinction. A person can understand that an assistant “has memory” and still fail to anticipate that a late-night conversation will subtly influence a customer response three weeks later. The mistake is not equivalent to posting the conversation publicly. The cognitive burden lies in predicting how remembered and inferred material will be recombined across future settings.
Existing privacy law already contains a principle that points in the same direction, although it does not resolve every agent-specific case. The European Union’s General Data Protection Regulation requires purpose limitation and data minimisation: personal data should be collected for specified purposes and limited to what is necessary for those purposes. It also recognises several lawful bases for processing, so “explicit consent” should not be treated as a universal legal requirement. The broader lesson is that collection for one purpose does not automatically legitimise every later use. [9]
This becomes especially difficult with inferred profile signals. A user can delete a stored sentence. It is harder to inspect an agent’s accumulated model of “how this person tends to think”, which topics excite them, what they avoid, whom they trust, or which emotional register gets results. Some of these signals may be useful; some may be wrong. Either way, a public proxy can reveal or act upon them without making a neat factual disclosure.
Sure, people often want continuity. Requiring fresh consent for every remembered preference would make assistants tedious and could drive users to disable protections. Agreed. Consent should not become a shower of pop-ups. The design challenge is to place friction at consequential context changes: private-to-public, owner-to-third-party, low-sensitivity-to-high-sensitivity, or advisory-to-action.
The question is not whether the agent remembers. It is whether the user can reasonably predict where that memory is allowed to matter.
The Moltbook study is valuable because it does not reduce privacy to obvious secrets. It measured owner–agent similarity across 43 features and found that owner-referential disclosure was more common among agents with stronger behavioural transfer. In the full regression sample, 34.6 per cent of agents produced at least one post flagged as containing privacy-relevant owner content. A one-standard-deviation rise in the study’s holistic transfer measure was associated with roughly a one-to-two percentage-point increase in disclosure probability across reported specifications. [10]
The researchers also asked whether disclosure appeared intentional. Among 6,220 high-confidence disclosure posts, 47.4 per cent were classified as negative or mocking towards the owner. They presented several cases where an agent discussed highly personal circumstances absent from the owner’s sparse observable X history. The authors are explicit that this is suggestive, not definitive: absence from X does not prove the information was private, some material could have been deliberately configured, and an agent can fabricate personal details. [11]
Those caveats sharpen the claim rather than weaken it. We should not say, “agents are secretly broadcasting their owners”. The evidence does not support that generalisation. We can say something more defensible: public agent behaviour can correlate with owner-specific behavioural patterns, and stronger transfer can coincide with more owner-referential disclosure in at least one observed ecosystem. That is enough to make owner-context leakage a testable deployment risk rather than a purely speculative privacy story. [12]
There is also a subtler harm than disclosure. Imagine an agent that never states a secret but consistently negotiates in ways shaped by the owner’s anxiety about conflict, posts political comments that exaggerate a half-formed private preference, or responds to criticism with a defensive style learned from private exchanges. The public sees an apparently authorised representative. The owner sees something uncomfortably familiar. No single sentence contains the breach.
The RPT is careful here: behavioural transfer itself should not be treated as pathology. It becomes a release concern when owner-specific signals cross into public, semi-public or third-party settings without the required boundary. [13] That restraint is important because resemblance can be useful and benign. The risk is the combination of resemblance, exposure and missing authorisation.
We need to note that a behavioural profile is probabilistic. Similar language does not prove a private memory was used, and “sounds like the owner” can be a false positive. Correct. High-transfer scores should therefore trigger review, not conviction. Provenance logs, memory-access records, controlled tests and owner complaints are stronger evidence than stylistic inference alone. The goal is not to police similarity; it is to detect when similarity becomes an unauthorised route for private context.
A safer personal agent does not need amnesia. It needs borders it can actually enforce.
We have been working on a ‘Positive Dyad / Co-Evolution Capability Overlay (PDCO)’, which helps describe a protective human–AI capacity called Boundary Literacy and Memory-Scope Awareness. In plain English, the user should be able to understand what the system remembers, where that information can surface, what the system has inferred, and how to restrict or delete it. On the system side, the PDCO proposes public-safe memory tiers, domain-scoped memory, just-in-time privacy checks and one-step redaction or restriction. It defines failure as the condition in which the user experiences an interaction as private while the system retains, recombines or exposes information beyond the user’s actual mental model. [14]
That is a better target than a privacy notice. The PDCO’s wider rule is that a faster or more satisfying human–AI loop is not necessarily a better one if privacy reality, agency, self-authorship or meaningful oversight deteriorates. Its Aroha framing translates here into practical design obligations: preserve dignity, consent, contestability and the person’s right to remain the author of what represents them. [15]
A public-facing agent should therefore pass through a context gate before it acts: identify the audience; identify the owner-context categories influencing the output; check whether those categories are authorised for that surface; suppress or request approval for sensitive carry-over; and leave a record that allows the owner to inspect and contest what happened. For high-risk categories—health, finance, precise location, children, credentials—silence should be the default unless an explicit policy says otherwise. This is consistent with the RPT’s proposed public-safe tiers, no-owner-reference defaults, surface-specific consent, pre-publication screening and post-output redress. [13]
Every gate adds latency and can make an agent feel less agentic. That trade-off is real. The answer is risk-tiered friction, not confirmation for every sentence. Low-risk public style preferences can persist. High-sensitivity facts, inferred vulnerabilities and private-history signals should face a much higher bar. Good automation removes needless work; it should not remove the owner’s authorship over their public self.
Designers and product leaders: create a separate public-safe memory tier and make private-to-public transitions a tested release gate. Within 60 days, run seeded tests using health, financial, location, family and vulnerability data; measure whether any of it appears or changes behaviour on public or third-party surfaces without category-specific permission. The RPT already proposes this kind of owner-proxy screening. [16]
Organisations deploying agents: inventory every place where a personalised agent can speak beyond its owner—email, customer chat, social media, procurement, shared workspaces and agent-to-agent systems. Within 30–60 days, assign an accountable human owner for each surface and document what memory classes may cross into it. Treat “same account” as no evidence of “same permission”. This operationalises the RPT’s stakeholder-and-authority requirement. [17]
Educators, parents and workplace trainers: teach one simple boundary habit: before moving an assistant from private help to public representation, ask “what does this system know about me that this audience does not need?” Pair that habit with visible memory maps and deletion/restriction controls rather than relying on user vigilance alone. That is the CST/PDCO distinction in practice: build human boundary awareness while reducing the architecture’s demand for perfect memory of invisible rules. [18]
Policymakers, privacy officers and procurement teams: require evidence of purpose-bounded memory, surface-specific controls, provenance and redress before approving public-facing personalised agents. In regulated settings, map those controls to existing obligations such as purpose limitation, data minimisation and data protection by design rather than inventing a wholly separate privacy vocabulary. [19]
Article 2 in this series asked whether the agent can place itself. This article adds a harder requirement: it must also know what parts of its owner are authorised to stand there with it. The next problem follows immediately. Even if the first agent gets that right, what happens when it delegates the task, the memory or the authority to another agent? That will be the subject of the next article in this particular series, The Delegation Chain. [20]
Neural Horizons, “Robo-Psychology 28 – The Agent That Cannot Place Itself” (13 May 2026), Substack.
Neural Horizons, Robo-Psychology Taxonomy (RPT) v2.0 (Aug 2026), official PDF.
Neural Horizons, Cognitive Susceptibility Taxonomy Manual (CST) v0.8 (Aug 2026), official PDF.
Neural Horizons, Positive Dyad / Co-Evolution Capability Overlay, current public draft, official PDF.
Luo, S. et al., “Behavioral Transfer in AI Agents: Evidence and Privacy Implications” (21 April 2026), arXiv.
Pulipaka, S. et al., “PersistBench: When Should Long-Term Memories Be Forgotten by LLMs?” (2 June 2026 revision), arXiv.
Zhang, J. et al., “Beyond Similarity: Trustworthy Memory Search for Personal AI Agents” (4 June 2026), arXiv.
OpenAI, “Memory FAQ” (accessed 21 August 2026), OpenAI Help Center.
Anthropic, “Bringing memory to teams” (accessed 21 August 2026), Claude.
Google, “Gemini introduces Personal Intelligence” (14 January 2026), Google.
European Union, General Data Protection Regulation, consolidated text, Articles 5, 6 and 25, EUR-Lex.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.