RSS Amplifier

Defrag Zone · Jun 10, 2026

This is how agents lie online

0
Sign in to vote or save

Francesco <frag> Gadaleta · Defrag Zone

From unsplash.com

We keep blaming the algorithm. Platforms, recommendation engines, engagement-maximizing feeds. But the algorithm doesn’t create consensus. Agents do. The real machinery is older, simpler, and far more robust than any ranking model. It’s the same process that decides which word survives in a language and which one dies. Understanding it changes how you read everything you see online.

I’ve spent years building AI systems that model agent behavior, at Amethix and on defense-adjacent problems where the cost of wrong consensus can be catastrophic. The pattern I keep returning to isn’t from NLP or reinforcement learning. It’s from a 1990s model of how populations agree on names for things.

The Naming Game, introduced by Luc Steels, is a minimalist model. A population of agents tries to reach agreement on a name for an object. In each interaction, a speaker picks a word from their internal inventory and says it to a hearer. If the hearer knows that word, both agents keep only that word and delete everything else from their inventory. If not, the hearer adds it.

That’s the entire model. No central authority. No truth signal.

The system converges to a single shared word not because that word is correct, but because it was the one that happened to survive the network’s interaction sequence. The winner is usually not the best word. It’s the word that reached critical mass first.

This is not a metaphor. It is the literal mechanism of online consensus.

The basic model assumes agents hold unlimited inventories. Extend it with finite-sized memories and the dynamics shift sharply.

Agents with small memories can only track a handful of competing claims at once. When they hit capacity, they drop the least-recently-used items. This means the first few pieces of information on any topic have a structural advantage, not because they’re more credible, but because they occupy memory before competing claims arrive.

On X, you see this as narrative anchoring. The initial framing of a story, often wrong or incomplete, colonizes a user’s mental inventory before the correction circulates. By the time accurate information propagates, there’s no room for it. The memory is full. The slot is taken.

Older users on social media don’t update their priors more than younger ones because of wisdom. They update less because the finite-memory effect compounds over time.

In the pairwise Naming Game, each interaction is private. Add group dynamics and everything changes.

When agents interact in groups rather than pairs, local majorities form faster and exert disproportionate pressure on minority vocabulary holders. A claim doesn’t need to be true to win a group discussion. It needs to be held by slightly more than half the visible participants at the moment the interaction occurs.

This is the mechanics behind ratio culture on X, pile-ons on any platform, and the way a thread of 200 replies can make a factually broken claim feel like settled consensus. The group interaction collapses the inventory of anyone uncertain. They drop the contested word and adopt the majority word. Immediately.

Mastodon’s federated structure partially disrupts this. Because group dynamics are scoped to instances, local majorities form within communities rather than across the entire network. You get more persistent minority vocabularies surviving in isolated pockets. This is both the feature and the bug.

Introduce a small probability that the word transmitted gets corrupted during communication and the model’s behavior changes fundamentally.

The Naming Game with learning errors never fully converges. Instead, the population reaches a quasi-stable state where multiple words co-exist, drifting slowly. Each transmission carries noise. Misheard claims spread as new claims. And because agents update their inventories based on what they receive, not what was sent, error compounds.

On social media, every layer of sharing is a transmission event with noise. Quote-tweets introduce framing changes. Summaries drop qualifiers. Screenshots crop context. A study claiming “X increases Y by 12% in cohort Z under conditions W” becomes “X causes Y” within four hops. The drift is not malicious in most cases. It is structural.

Nostr’s model is interesting here because it encodes the original note with a cryptographic signature and preserves the propagation graph. In theory, you can trace drift. In practice, most clients don’t surface this, and users don’t look.

The Naming Game on networks with community structure produces the most politically relevant dynamics.

Within tightly connected communities, consensus forms fast and is stable. Across communities, bridging agents carry words between clusters. The words that cross community boundaries are not the most accurate ones. They are the ones that survive translation, meaning the ones that are ambiguous enough to mean something to both sides, or provocative enough to generate a response.

This is the selective pressure that inflates emotionally charged, context-free claims. They cross networks. Nuanced ones don’t. The bridging agent on X is the account with followers in multiple communities. What they amplify is not what’s true within any one community. It’s what travels.

Multi-community dynamics also explain why corrections fail even when they reach the right people. A correction originating in one community arrives in another as a foreign word, not yet integrated into local inventory. The community’s established consensus treats it as noise and the finite-memory effect drops it.

The final extension that maps cleanly to the current information environment is the multi-word variant, where agents hold inventories of entire phrases, framings, or narrative structures rather than single tokens.

Competing narratives about the same event behave like competing words in this model. They spread in parallel, accumulate local majorities, interact with errors during transmission, and eventually one framing dominates, not because it was more accurate but because it was more transmissible in the specific network topology it encountered.

The same event can stabilize to entirely different dominant narratives in different network communities. This is not relativism. It is a mathematical property of the system. Both communities ran the same game. They started from different initial conditions and different network structures. They converged to different words.

Join us on Discord

The operational implication is uncomfortable. If you understand the Naming Game, you don’t need to lie to win the information environment. You need to seed early, target bridging agents, and introduce just enough error to prevent convergence on accurate competing claims.

This is not a hypothetical. Defense and influence operations have understood distributed consensus mechanics for at least a decade. The academic literature on naming game variants maps almost exactly onto documented IO tactics.

For platform design: Mastodon’s federated structure creates many stable local consensuses. That’s not worse than X’s global convergence, it’s just different. Nostr preserves transmission provenance but doesn’t enforce anything about what users see. None of these architectures immunize against the game. They reshape which word wins.

The next 12 months will see more AI-generated content entering these networks as agents, not just as tools. Synthetic agents with designed vocabularies and targeting strategies will participate in naming games at machine speed. The convergence dynamics will accelerate. The error-injection will be precise.

The meaningful response isn’t content moderation. Moderation operates on words after they’ve already propagated. The meaningful response is network-level transparency: who interacted with whom, in what sequence, before consensus formed.

That’s a structural problem. And right now, almost no platform is built to expose it.

The real misinformation problem isn’t what’s false. It’s that the game doesn’t check.

Read the original on defragzone.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.