6 October 2025
I’ve been dabbling along the conversation around context engineering and agent experience for the past year. Along the way, I’ve been angling for more designers to be part of these conversations. They’re sometimes deliberately dense, but the reality is that we’re talking around the actual ground level deployment problems in a world where agents intermediate things we used to do ourselves. (I’ve been calling this disintermediation)
I’ve been circling the wagon all year about both agent experience and its relationship to the people we’re building things to “help.”
For starters, we have to understand that we’ve been delegating tasks to computers for a long time. The difference is that when we did it ourselves, there was a baseline of reliability that the transaction reflected our intent. Increasingly, we’re moving into a world where people will delegate sensitive decisions to agents to “save time.” When those go wrong, they’ll be forced to work across other agents to rectify mistakes. These are not edge cases. There will be a range of unforeseen consequences, especially in public services and the kinds of life and death decisions that sit in between. This is where I think context becomes a clearer design need.
As I’ve spent more time working on experiments, collaborating with others, and figuring out where lessons from other design contexts apply to LLMs and AI contexts, I’ve decided the best place to start is context engineering. Long story short, most people use the term “context engineering” to describe extensible ways of managing context states for LLMs. This includes things like retrieval, memory architectures, and prompt optimization. This year, I’ve been arguing we need to broaden that frame to include the systems these models enter, the assumptions they carry, and the consequences they create on the people using them. (Because we’re always assuming that people will use them at some part of the stack.)
State capacity—a government’s ability to effectively implement policy and deliver services—is a contested concept, but understanding it is critical when AI systems are being embedded in public systems. With localities globally rushing to adopt AI solutions, we need more than ‘working guides.’ We need approaches that capture the real contexts these tools operate in.
Designer Soren Iverson has been publishing UI prototypes of imagined realities that veer from dystopian to eerily plausible. I’ve experimented with using LLMs for similar purposes in the past.
My latest project, State Capacity AI, is going to operate a bit like a personal R&D lab, a place to expand how we think about public facing AI deployment and the nerdy behind the scenes mechanics that shape what people actually experience.
To be glib, we’re essentially letting an ICQ window dictate the course of the next few years. We need better and more imaginative ways to guardrail these systems and, if nothing else, to provoke debate about where we are, where we could go, and whether we can do better.
I’ll be posting more here in the coming weeks and will report back.
Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.