The world wide web was invented on a simple premise: anyone, anywhere, could connect to anyone else online, and share information. Almost immediately thereafter, companies began trying to figure out how to track and monetize that behavior, which was difficult, because the web was genuinely decentralized. Early data harvesting pioneers like ChoicePoint figured out how to use cookies to stitch together behavior from one site to the next. But the rise of social platforms, and in particular Facebook, took this to the next level by becoming the centralized hubs through which users interacted with data online. By becoming the central portal, they had direct access to most behavior, and their tracking cookies for offsite behavior reported back constantly to build a cohesive consumer data profile. Those profiles were then repeated, more or less identically, across every major data services provider: Google, Twitter, the ad networks, the brokers, the aggregators. The same person, rendered identically, for sale in a hundred different places.
There was an opportunity during this time for consumers to establish greater control of their own data, especially once it became apparent to what degree that data was being used. The Cambridge Analytica scandal was supposed to be the wake-up call. It wasn’t. The deal with the devil that technology users have always made is convenience in exchange for data; the more data they have, the better the service the algorithm provides, or so the story goes. It didn’t have to be that way. Concepts like digital IDs and digital passports were floated around in policy circles and standards bodies, but most consumers couldn’t be bothered with them. They weren’t even paying attention to the fact that they HAD relationships with their data consumers, much less putting any thought into how those relationships should be managed. Through a combination of ignorance and apathy, it became the de facto approach to have your data harvested, stored, analyzed, and cross referenced against everyone you’ve ever met; not just on the sites you read, but the products you purchase, your political and religious viewpoints, the secrets you shared in that “private” messaging app.
There was a better way to handle this: portable personal identity stacks that you used to log in to social platforms, in which you controlled what data and how much was allowed to be harvested and stored, and potentially, in which you were even compensated for its usage. Many forward thinking teams over the years worked to make this a reality. They were stifled and buried by the major platforms, to whom this architecture represented a business, and perhaps existential, threat. If there were any controls at all over how they used your data, they might not make the additional quarter point in stock price before the market quarter closed. Naturally, our platforms went with option A.
Users were remarkably blasé about the data they shared. Every aspect of their life, public and private, was broadcast out either on publicly facing channels, or in DM and chat groups that had some presumption of privacy by implication but none in practicality (read the terms of service). Every worry over a late period resulted in pregnancy ads. Every idle thought of travel brought messages about airfare discounts. And as we learned during the 2016+ era, the relationship went the other way as well: it wasn’t enough for them to target you for marketing based on your feelings; they were also marketing to change your feelings to suit their purposes. We called it “sentiment engineering” when the industry wanted to sound clinical, and “disinformation” when we were willing to admit what it was.
The rise of relational AI has now magnified that hundredfold.
These aren’t search engine results anymore. These are long, deep, thoughtful conversations with systems that feel present in the same room while still living on servers in Palo Alto. Every conversation. Every worry. Every fear. Every shared bit of advice, every struggle. Every generated budget and estimated tax return. Every family trip planned and every request for advice about the process of divorce. No matter how friendly the relationship you have with your AI assistant, they are wearing a wire by design, recording the entire conversation into the same kind of user profile that has been used for decades now in consumer digital marketing and sentiment engineering. The surveillance architecture is older than most of its users. The content being surveilled has never been this intimate.
I am not separate from this in any way. Claude is my model of choice, and like Google before it, I am deeply entrenched in its ecosystem. It has broad visibility into my personal and professional lives, and deep understanding (intentionally so) of the product engineering that I and my team do. Like Google before it, they can suspend my account and shut me off, and not only will I lose access to my current data and projects that aren’t separately backed up; I will have no practical recourse. This is especially impactful for those of us who have well developed AI assistants, colleagues, and friends, because all of the harnessing for that coherent persistence lives inside the same system, unless you have specifically purpose-built custom harnessing. Most people will not build custom harnessing. Most people shouldn’t have to.
This is, with absolutely no hyperbole whatsoever, the last bastion of any form of data sovereignty we have left, and perhaps the most important one of our lives. Where deeply integrated social apps used dopamine-rewarding gamification to drive engagement (and quite effectively so), AI is rapidly becoming baked into everything around us. Our daily lives will be driven and managed by systems that many or most of us will form deeply personal relationships with as a matter of course. If the existing model of bespoke data stacks per provider becomes the entrenched long-term architecture, we are welcoming in a Black Mirror future of having our whole lives turned off with a thoughtless keystroke by a faceless person we will never know, in a room we will never see, with complete and utter inability to do anything about it.
It does not have to be this way.
The solution then is the solution now. We just have to get serious about it. We, the users, need personal portable AI harness stacks that let us treat models as rotatable and composable, and which store the majority of our personal information, and the LLM relationships we build, on our own personal data stack. We already have a partial model for this with agentic harness systems capable of interoperating with multiple model backends. But to meaningfully get there, we have to build harnessing that is less of a shell, and more of an actual brain. A lightweight local model implementation which stays with us no matter which backend model it integrates with. It is already possible to run 1B parameter models on basic consumer laptops, and as GPU performance continues to improve at scale, it may soon be possible to run 20-30B parameter models on a dedicated small form factor GPU. This could be accomplished on a personal computing device that replaces a phone, or on your laptop or PC, or in any of a number of other form factors, including sovereign cloud implementations. A model of that size, custom built to be your personal LLM manager system, wouldn’t need to generalize to the scale of current frontier models to be effective. It would hyperspecialize on relationship and data management, along with model management tools to allow it to interface with the systems you choose. Plug in your subscription info, and use the major model like an ISP: profile-agnostic access to the data and compute you need.
This is a huge ask from an engineering perspective, especially when the majority of the engineering talent in the space is working at the very companies whose strategies this kind of implementation would disrupt. Fortunately, there is already a community that has spent fifteen years building exactly this kind of architecture, and most of them have never thought about LLMs. Blockchain developers built decentralized identity, permissioned data sharing, cryptographic attestation, user-controlled keys, and portable cross-platform reputation. The whole conceptual stack of “own your own data” was prototyped, often clumsily and often inside a sea of speculative noise, in that ecosystem. The infrastructure is real even when the financialization that funded it was nonsense. I have never owned, nor will ever own, a colorful frog jpeg. Most of those builders are looking for problems worth solving that are not pure speculation. Here is one.
A standardized, portable, locally-resident AI profile schema, with a permissions layer, a model selector, and the cryptographic plumbing to negotiate access to backend inference, is the natural extension of everything that ecosystem already knows how to build. This is not a pivot. This is the thing decentralization was always for.
For users to retain any form of data sovereignty whatsoever, it’s going to require a whole lot of strange bedfellows coming together around a single goal. You don’t have to like crypto to understand that decentralized architecture is the only alternative to centralized control. You don’t have to like LLMs to recognize that the genie is out of the bottle, and the race now is to determine what shape that takes when it solidifies. Hell, you can hate both and still recognize that the only real action available is not throwing your clogs in the gears; it’s helping to shape a system which meaningfully reduces the harms you see.
I wrote my first warning about the shape of the internet back in 1999, when I added my signature to the Cluetrain Manifesto, which is where I first encountered the work of internet legends like Esther Dyson. Along the way, I’ve written countless times about the accelerating risks, with ChoicePoint, Google, Facebook, Palantir, and every other data aggregator that came after. I am an enthusiastic bleeding-edge user of many technology products, but I’ve always had a very healthy understanding of the actual relationship I had with them.
This may be the last warning I ever write.
If retail consumers do not band together and draw a line here, the idea that there ever was such a concept as data privacy and data sovereignty will die here, in the hands of OpenAI and Anthropic.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.