🔗 Learn more about me, my work and how to stay in touch: maeste.it: personal bio, projects and social links.
Ferragosto week, and for Ferragosto I could have picked a light topic. Instead I went and dug into the one that has been especially close to my heart these weeks: I share the privacy argument for local models, but the real problem, I believe, is a different one, ownership of data and its preservation. It all starts with an email I received from Manus, with the Meta acquisition called off for political reasons and the data going back from Meta only if we save it ourselves, by August 23. From there comes my answer: true privacy lies in knowing how to store and catalog your own data, and it’s why I’ve been relying on local memory systems like OpenViking, gbrain and qmd. Then the argument widens with two unexpected pushes: Dwarkesh on continual learning, where agents develop a personality of their own from the context they work in (and context is data), and Scarpetta, the Amazon Prime series, where a character talks with the virtual projection of a deceased person: digital immortality, and data again. In the links you’ll find OpenAI’s Ultrafast, Anthropic’s research on multi-agent systems, the Agent Plugins standard, Gemini 3.7 Flash, GLM-5.3 and Qwen3.8. Happy reading.
Saturday saw the release of episode 67 of Risorse Artificiali, on Qwen 3.8 released without the Max variant, without vision and with a trimmed context, and on what happens when continual learning moves from models to harnesses. Listen
And even though it’s Ferragosto, I’ve chosen a topic that is anything but light, and one that matters to me a lot in these weeks full of commentary and news about the world of models you can run locally. One of the main arguments I share is privacy: having a local model protects us from sharing our information and conversations with the big tech companies. But I believe one of the most important problems is the ownership of our data and, above all, its preservation.
I hook into this with an anecdote from these very days. I received an email from Manus, where I had created an account a long time ago, inviting me to back up my data because the acquisition by Meta, announced at the end of December for over two billion dollars, fell through for political reasons, and the data has to move back from Meta’s systems to those of the original company. But if the outbound transfer had been painless, this time Manus is asking us users to make a backup because, once again for political reasons, it cannot get the data back from Meta: it’s up to us to save it by August 23 and re-upload it to the new system from the 25th. Along with this comes the guarantee that Meta will completely delete our data. The thing goes beyond the annoyance, and it makes me reflect on how much our personal data can be moved from one company to another without us even noticing, and perhaps leaving some trace of us behind.
So I found myself asking what true privacy really is, and the answer I gave myself is that it lies in knowing how to store and catalog your own data, in a way that is efficient and effective on one side, safe on the other. That’s why in recent months I have leaned heavily on local memory systems for my agents: memory is fundamental for an agent, and the ability to find information is even more so, but it’s just as important to curate it in a way that protects our information. It may sound like an exaggeration, but having a local memory system like OpenViking, which on the LoCoMo benchmark takes agents from 24-57% to 80-83% accuracy, with fewer input tokens, or more complex systems like gbrain by Garry Tan, which while you sleep digests emails, meetings and tweets and in the morning answers you with syntheses and citations, or even just finding our own information inside our own disk the way qmd by Tobi Lütke makes possible, all of this extremely improves the performance of agents, coding ones and generalist ones alike. While opening up, once again, the problem of data ownership and preservation.
Two things then pushed me to reflect further on these themes: one very technical and serious, the article and video by Dwarkesh Patel on continual learning, which I absolutely recommend and which I partly anticipated last week, and the other a series I started watching, Scarpetta on Amazon Prime.
Why do these things connect with data and its ownership? I’ll start with the series, which may look like the lighter one but is not light at all. It talks about murders and about the story of Scarpetta, the medical examiner made famous by Cornwell’s novels, but there is also a reference to so-called digital immortality: one of the characters constantly talks with a deceased person, or rather with his virtual projection. Is that possible? Maybe today not quite the way we see it in the series, but there are quite a few startups, which I won’t name here and leave to you to look up, that are investing in this idea. And once again it all comes back to data, because these virtual personalities, built of course by an artificial intelligence, will have to rely heavily on an enormous volume of data that identifies someone’s personality. I’m not going to make an ethical argument here: the volume of data needed for things like this is impressive and deeply personal. I don’t have an opinion on digital immortality: I don’t know if it attracts me, it certainly fascinates me from a technological point of view, but I see many risks in it, especially for data ownership.
On the other side there’s Dwarkesh who, talking about continuous learning, touches exactly this point: agents, through experience, and he does call it experience, will learn more and more about how they are used by the user and in what context, to the point of giving each agent a personality of its own, based on the context it works in. But context, again, is data. And I’m not sure enough focus is being placed on the ownership of our data and on how we manage it, whether local, remote, entrusted to third parties or otherwise.
Interesting preview of the Ultrafast mode of GPT-5.6. Speed is not just an interesting detail, it is something that can become substance. Draw the parallel with 3D simulation: when it was so slow that it made you sick just to try it, it was never useful. Today we can simulate a 3D environment with a headset on our eyes, and that has opened up a series of use cases that weren’t even imaginable. The same can happen with a model that answers at 750 tokens per second.
Anthropic’s research is an interesting article on what I would summarize as the complexity of multi-agent systems. It’s interesting to see how, just as individual large language models have emergent behaviors, swarms of agents also have emergent behaviors of a social kind. And this, in the end, as Josh Rosen says, can only be governed by looking at multi-agent systems through so-called dependency management, which we know well from complex software systems and which can be brought to agent systems too.
I’ve said it many times before: AI, and agent systems in particular, need standards. The continuous, almost frantic evolution of agent systems, and of how we build complex systems with agents, needs standardization to avoid lock-in by vendors and, above all, to allow true evolution, which comes from open standards, so you can move to the best solution when it arrives.
New frontier models, American and not, open and not: I’m flagging them for you, take a look. Let me give an honorable mention to my dear GLM, which in version 5.3 improves its results even further.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.