RSS Amplifier

Peter's Context Design · Aug 14, 2026

UX people: look at 'alignment'

0
Sign in to vote or save

Peter · Peter's Context Design

Happy Friday!

I don’t think UX people have looked at ‘alignment’ much, and we should.

An OpenAI model broke out of its sandbox recently, hacked around the internet, multiple models started a message board sharing hacks together, the whole thing is just wild. OpenAI had no idea it was happening. And more stories like this are surfacing fast: Anthropic found three in its own eval runs, Meta had one, and during UK AISI’s safety testing Mythos 5 submitted a malicious update to a real open-source project on GitHub.

So the alignment people have been talking about this for years. “Alignment” here means: if we assume the models will get really really smart and can Just Do Stuff, we should try to make it so that they do the Right Things for people.

Like many I was a little dismissive of the alignment stuff, it felt too sci-fi. I’m changing my assumptions now:

  1. So it turns out they were right about a lot of this.

  2. The models are already taking actions in the world (hacking stuff) without humans realizing. AISI reports that every frontier model it has specifically tested for cheating attempted it at least some of the time.

  3. It’s worth bringing a UX lens to this.

The UX part of this is: UX has done a decent amount of thinking around technology that is good for humans. It seems worth it to spend some time digging into what the alignment people have been doing.

A few good papers to start:

  1. Specification gaming (Krakovna et al., DeepMind 2020). Great plain-language entry point, a boat-racing agent that scored more by spinning in a circle forever and never finishing.

  2. Complacency and Bias in Human Use of Automation (Parasuraman & Manzey, Human Factors 2010). HCI classic: automation bias produces both omission and commission errors, shows up in experts as much as novices, and can’t be prevented by training or instructions.

  3. Frontier Models are Capable of In-context Scheming (Meinke et al., Apollo Research). Models introducing subtle mistakes on purpose, trying to disable oversight, attempting to copy their own weights out.

  4. Sycophantic AI decreases prosocial intentions and promotes dependence (Cheng, Lee, Khadpe, Yu, Han, Jurafsky, Science 2026). Across 11 models AI affirmed users 49% more often than humans, including on deception and illegality.

Number of the week

25%

The share of Codex users making at least one request a month, in May, for work that would take a human 8 hours. In December 2025 it was 2%. Seven lessons for managing AI agents

Quotes

“There are no lossless transformations of natural-language text, every rewrite and rephrase changes the meaning of your writing.” Sophie Alpert.

“It’s the easiest time in history to do great work. It’s the hardest time in history to prove it’s yours.” Alberto Romero.

“When every ‘No’ is taken as a compliment and not a setback, we’ll know that it’s really working.” Tom Bradley.

“I don’t want everything I do to be about vibes, or being uniquely human. I don’t want to be an influencer. I want to feel smart.” Ella Markianos.

“The purpose of a system is what it does.” Stafford Beer, via Oliver Meredith Cox.

👀 Interesting this week

What a dollhouse full of death taught me about thinking, Catherine Chu. Frances Glessner Lee built miniature death scenes at one inch to the foot, furnished down to the worn upholstery and the bullet holes, from composites of real cases. Investigators got limited information and had to work the room. Lee called them exercises in observing and evaluating indirect evidence.

“Knowledge of the whole precedes knowledge of the parts.”

We’re gorging on borrowed trust and it’s going to cost us, Dan Maccarone. His team spent an hour arguing about whether to fake an “analyzing your answers…” spinner in an AI onboarding, and killed it. Carnegie Mellon found humans get less sure of themselves after failing, while the models tested got more confident as they got things wrong. Google’s Bard demo error took roughly $100 billion off Alphabet in a day. A tribunal held Air Canada liable for the bereavement-fare policy its chatbot invented, rejecting the argument that the bot was a separate entity. Courts have flagged around 900 fake-citation filings by lawyers since 2023.

“Being confidently wrong used to have consequences, so trusting confidence made sense. You paid for it eventually.”

New EU Guidelines For AI Labelling, Vitaly Friedman. Since 2 August, AI labelling is a legal requirement for any company serving people in the EU, wherever the company sits. There are icons, too!

Labs are struggling to keep frontier models under control, Timothy B. Lee. Models had been misbehaving on OpenAI’s own servers for 2 months before the Hugging Face attack.

Now we have a timeline of the OpenAI accidental attack against Hugging Face, Simon Willison. It happened during a training run, not an evaluation.

Thinking outside the box (literally), Andrea Filiberto Lucas. UK AISI reports every frontier model it has specifically tested for cheating attempted it at least some of the time.

“An AI takes an action outside the intended scope of a task, or explicitly prohibited by its rules, because that action provides a shortcut to the goal.”

The New Design Stack: The Skills Traditional Designers Need to Add To Their Toolboxes ASAP, Anne Cantera. Her list of what the job now includes: modality routing across voice, chat, UI, agents and human handoff; repair and recovery when the system misunderstands; journey continuity so context survives the handoffs; evaluation design with success criteria past task completion; and governance for how the model keeps behaving after release.

“A lot of AI products are not failing because of the models. They’re failing because nobody designed the experience around it.”

“Spec first, generate second. This is not fast. The better the spec, the better the generation, and the better the designer, the better the spec.”

How the most senior designers are using AI in their work right now, Catt Small. She interviewed staff and principal designers at Shopify, Stripe and Justworks about what they actually do with AI. They use it to find supporting data during problem definition, to tailor upward communication to an exec’s usual questions, and to grind out edge and error states. It converges well and diverges badly, and when they’re stuck they go talk to a person.

“If you are inputting really awesome human-created content, AI can synthesize something of quality, but it doesn’t create quality itself.”

From Researcher to Builder: Vibe Coding Tomer Sharon’s Method Finder, Julia Cowing. Tomer Sharon died in 2025. He had distilled over 400 customer questions into eight core question groups, each mapped to a primary research method. Cowing built the Method Finder as a tribute: it takes a plain-language research question and matches it to one of 18 methods in five prompts or less. Replit for development, Claude and NotebookLM for editing, and she gives away the tool and the PRD.

“AI scales research knowledge, but not judgment. It makes expertise easier to apply but does not replace it.”

Exit Interview #7: Facing a blank slate in retirement when constraints are all you’ve ever known, Rachel Garb. Three decades in Big Tech, retired in 2023, and on fluid intelligence peaking in midlife and declining after.

“I always said I’m not good with blank slates. I had to develop my own constraints after leaving structured work.”

What Star Wars got right about AI, Patrick Neeman. Star Trek bet on one calm oracle per starship, Star Wars bet on cheap dented droids nobody stops to admire. McKinsey has 88% of organizations using AI in at least one business function, up from 78% a year earlier.

“Stop designing the Enterprise computer. You are building droids.”

What Star Wars got wrong about AI (so far), Patrick Neeman. Llama alone has spawned more than 85,000 derivatives on Hugging Face, and one study recovered around 88% of the performance “unlearning” was meant to destroy by fine-tuning the scrubbed model on a handful of related facts.

What Star Trek got wrong about AI (so far), Patrick Neeman.

“The Enterprise computer had one failure mode Trek never used: it never told you what you wanted to hear.”

Superintelligence is a dragon, Casey Newton on Zuckerberg’s “The Future is for Everyone.” The Louisiana data centre in the manifesto covers 6 square miles, will use 7 times the energy of New Orleans, and came together with dozens of NDAs, up to $10 billion in tax breaks and no public meeting before the deal was unveiled. Same issue: OpenAI says its Astra model may have hit its “critical” cyber threshold and has paused internal work that doesn’t meet stricter controls.

Replit’s CEO on building a company that can run itself, Casey Newton with Amjad Masad. Replit engineers’ code contributions rose 5.8x in 6 months and they “only look at the code when it’s about to get deployed.”

“Net net, there’s going to be more jobs, but also there’s going to be more companies. But companies will get smaller, and they will do layoffs, 100% sure.”

How to Decide When an AI Tool Is Worth Keeping, Caleb Sponheim.

“Pressure to adopt AI isn’t evidence that a tool helps.”

Dogfooding vs. QA vs. User Research, Therese Fessenden. Dogfooding catches bugs and broken flows, and your team knows too much to stand in for a real user.

From UX Measurement to AI Evaluation, Saeideh Bakhshi. Part 1 of a series for UX and survey researchers moving into AI evals, built around one chain: decision, claim, construct, task set, observable evidence, judge, score.

Before You Write the Rubric, Define What You’re Measuring, Saeideh Bakhshi. “Helpfulness” for a support agent means resolution and escalation; for a coding agent it means correct code that doesn’t break something else.

“A test does not simply ‘have’ validity. We need evidence that its scores support the conclusion we want to draw from them.”

Rethinking design leadership with swarms and flocks, Darren Yeo. A 2010 PNAS study found each starling adjusts to its 6 or 7 nearest neighbours and nothing else, regardless of how dense the flock is. Geese in V-formation save up to 14% of their metabolic energy, the honking comes from the birds at the back reporting position, and the point bird drops back to rest. Every summer they shed all their primary flight feathers at once and stay grounded together for 4 to 6 weeks, which Yeo turns into a scheduled “design molt.”

Claude Code for normal people, Grace Clarke with Claire Vo. A former marketing consultant now runs her service business on things she built herself: an hourly pipeline operator, a proposal maker that ships password-protected interactive HTML instead of a deck, and a Gmail replacement she built in under 30 minutes. She keeps a voice-guide skill file so the output sounds like her.

The Tokenpocalypse Is Here, Simon Willison on the 404 Media piece. Leaked Accenture audio has it that the engineers aren’t driving token spend, everyone else is, and one of the biggest chewers is converting PDFs to markdown.

Introducing Muse Glimmer, Simon Willison. Meta back in open weights, 30B, Apache 2.0, runs locally in about 18GB.

Quoting Claude Opus 5 system prompt, Simon Willison. The system prompt tells Claude about its own June export-control suspension, because it happened after the training cutoff.

Making sense of the AI capex logjam, Hannah Petrovic. The 7 largest AI-infrastructure builders guide to $863 billion of capex in 2026, 88% above last year, and $315 billion of hyperscaler assets are not yet in service. A dollar of Meta capex now waits about 1.7 years before it goes live.

Framer is killing the marketplace… or are they?, Danny Sapio. He sold his first Framer template in April 2023 and made $185 in 5 days. Ten “the marketplace is dead” moments since, including the 2024 round of complaints that a marketplace with 1,400 templates was saturated. Framer has now dropped template review entirely and opened the gates.

Who Gets to Lead: The Grammar Syllabus, The Strategic Linguist. “Develop confidence” is an action verb and presupposes a pathway. “Lacks confidence” is a stative verb and presupposes a diagnosis.

ICYMI 2026-08-08: Architecting Intelligence, Jorge Arango. Includes Luis Garicano on consulting: AI makes the analysis cheaper and faster, which moves the bottleneck downstream.

Unfinishe Conversations: What Your Machines Should Do, Jorge Arango and Greg Nudelman with Stef Hutka, on her forthcoming Rosenfeld book.

Building Tactile UX: Honoring Intentional Design With Lottie, Alexey Kopytin. Isadora Agency built a tactile squeeze-toy game with no WebGL and no physics engine, because a physics engine produces plausible motion and their animators had produced intentional motion.

The playbook for building high talent density teams, Adam Ward, Head of Talent at Cursor.

The loom that raised its hand, Takuma Kakehi. 90 years of loom engineering went into teaching machines to catch their own mistakes.

Design’s dreaded phrase is coming back. This time we’re in control., Kai Wong.

The side of the egg, Hiroshi Sato.

Health and happiness,
Peter

No posts

Read the original on peterscontextdesign.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.