jonplummer.com

Care has to show up in the product

If attention to detail isn’t visible in your product, it’s hard to say you care about the details.

What is in your heart doesn’t count. Values on a slide don’t count. A team that talks about craft while shipping something that looks unattended is not a team that displays care – not in a way that customers and colleagues can verify. You care or you don’t, as an organization, and it shows in what ships. If you think you care and the product doesn’t show it, you need to change your behavior.

Designers: make intentional roughness clear

There’s an uncanny valley between rough-on-purpose (we’re moving fast, expressing concepts, have time to clean up after we make some big decisions) and rough-from-neglect (we’re moving fast, shipping under deadline, maybe this isn’t that big a deal, maybe we’ll fix it later). Is this prototype obviously unfinished, or does someone think it’s done?

Intentional roughness is fine. Unclear roughness isn’t. If a design is meant to be provisional, say so in a way a product manager or engineer can’t miss: what’s settled, what’s placeholder, what you’ll return to. When that signal is missing, people guess. Some will treat draft work as final. Others will treat finished work as optional. Neither helps. And if I can’t tell that your “final” is final, it isn’t.

Part if the job is to make the state of the work clear so people can tell where we are in the project.

Leaders: stamp out “too small to matter”

It’s common in engineering, product, and design to treat some details as too small to bother with. A misaligned control, inconsistent label, or empty state that doesn’t offer action, by itself might seem tiny. Together they create the (correct!) impression of a company that doesn’t care, even when people insist they do.

The cost of getting those details right is often not high. When it seems high, you usually haven’t looked for the simpler fix. “We’ll live with it” should be a conscious trade against an agreed-upon quality standard, not a reflex because you are in a hurry or the ticket felt nit-picky.

Pressure to ship is ever-present, and letting a few small things through can be fine if you have a real commitment and habit to come back and clean those things up. Pressure to ship plus “too small to matter” is how neglect becomes the norm in your organization.

Build the capability and the expectation

If the product doesn’t show attention to detail, ask your designers to name the details – the short list of things that, if fixed, would make care visible.

If they can’t produce that list yet, that’s a capability to build and an expectation to set. Attention to detail is often visible in portfolios, but the story doesn’t end there. Hire for it and expect it to be a habit. People who attend to detail show it in the work and in the list of improvements they’d make if they had the time.

If they can name the details and the product still doesn’t show them, the blockage is elsewhere – process, incentives, partnership, or ownership. I’ve written about diagnosing that elsewhere. The point here is simple: care is a behavior visible in what ships. Talk alone is not evidence.

The plateau of sad gray icons – Yesenia Perez-Cruz
“Without strong, concrete examples, the agent reverted to generic forms and basic tables the moment things got complex. Tokens and prompt rules worked fine for basic screens. Basically any place where a general SaaS software pattern was sufficient. But in areas where I had refined the experience into a domain-specific shape, like treating a paycheck as an editable set of ledgers instead of a standard form, the agent needed real references.”

Your website is boring (but that might just set it free) – Sam Belt
“The rise of Shopify, Squarespace, and Webflow templates didn’t only make it easy to build functional websites; they made it almost impossible to build anything else. It is easy to understand why this happened; templates solved massive engineering headaches, offering speed, reliability, and cost-efficiency that businesses desperately needed. But the result? Mass homogenisation.”

Fire Weather: On why tech’s worst-rated managers are not the problem – Caitlin Steele
“Designers and researchers are, in the report’s words, at the epicentre of AI anxiety across the board: “Among designers, 63% feel ‘overwhelmed by the pace of change’ and 61% feel ‘tired’, the highest of any role.” Sixty-one percent feel the comp squeeze directly, the most of any discipline. Designers and researchers report the lowest willingness of anyone to recommend their field.”

The Lost Discipline of the Alarm: What Notification Design Forgot – Saleh Kayyali
“Notification design has a problem it did not invent and has never named. The problem is alarms: which signals deserve to interrupt a person, and which should be kept quiet. And it was diagnosed, with a whole discipline built around it, long before a single app ever buzzed in your pocket, in oil refineries, nuclear control rooms and hospital wards.”

I handed a UX review over to AI. Here’s what happened. – Kolozsi István
“Both models put together an analysis, but they were full of irrelevant, and even non-existent, problems. There were useful observations too, but even those I could only really use at the level of wording and structure. Claude was noticeably more accurate than ChatGPT, at least in phrasing and structure, though this is my subjective impression, not a controlled measurement.”

Why Human-Led Research Still Matters in the Age of AI – Maria Rosala
“Research produces two things. The first is outputs: the themes, the recommendations, the report. The second is harder to see but just as valuable: the learning that happens to the people doing the research – the observing, the wondering, the working through what the data means. AI is good at producing the first. It cannot produce the second, because the second isn’t a deliverable.”

How do you raise the level of craft?

I flubbed an interview question the other day. Not because I gave a wrong answer, exactly, but because I gave a truthful answer that was too narrow, possibly not suited to the company, and not representative of what I would do were I hired.

The fateful question

How do you raise the level of craft?

It’s simple, full of possible directions to pursue. I should have asked about the live situation the prospective employer faced so I could address that; after all, it’s why they are asking. Instead, I described what I’d done in the past. At two different companies, design craft wasn’t reaching the shipped product because design, product, and engineering weren’t well-aligned. Design work got approved, then thinned out somewhere between specifying and shipping. UX, product, and engineering didn’t agree on what quality factors to focus on, what level of quality we were going for, or how we’d notice.

There was no shared definition of quality – ask three people what “good” meant here and you’d get three different, equally defensible answers. There was no agreed bar for how much was enough – whatever shipped fastest won, no matter what anyone said they wanted. And there was no way to repeatably prioritize – no shared sense of which craft fights were worth having this cycle and which could wait. We either fought over everything or let all of it slide, depending on the scrum team’s habits.

So I worked on the “product trio” partnership – closer collaboration among product, design, and engineering; shared ownership of what “done” meant; engineering at the table from the beginning. It helped a lot, both times.

But that’s not me

It sounds pretty good, but I did myself a disservice with this answer because it didn’t reflect what I would actually do if hired. I explained a fix instead of a process – and a fix that would only apply if the obstacle to craft is the same one I bumped into last time.

Obstacles to craft can appear in several parts of the development lifecycle. Poor partnership between design and engineering is just one potential cause; the job in a new org begins with finding out where craft is stalled before reaching for an intervention.

The better answer – diagnose, intervene, repeat

Diagnosis starts with evidence. Find a place quality went missing: a specific screen, a release, a moment where what shipped doesn’t match what was designed or what the team is capable of. List the ways in which we’re disappointed with the results, be it fit and finish, workflow, error handling, accessibility, etc. Then work backward: which failure modes explain how we arrived at those results? Often more than one failure mode is operating. The job isn’t picking a favorite, it’s figuring out which one is doing the most damage right now, in this org, at this moment – because that ranking is what tells you where to intervene first.

This is the step I skipped in the interview. I didn’t bother to investigate. I leapt past the diagnosis.

Intervention comes after diagnosis, and it should be sized to match your confidence in that diagnosis. A lightweight version of the fix – one critique session, a prototype-first spec, a new metric added to a dashboard – tells you whether you found the right cause before you commit more budget, headcount, or political capital to a bigger structural move. And it needs a real signal attached, something witnessable, not just vibes.

Failure modes

  • People and skill
    • Designer skill level, or skill variance across the team
    • Engineering skill level – it often takes front-end expertise to know what to care about or to feel like it’s sensible
  • No shared definition of quality
    • Nobody disagrees that craft matters, but ask three people what “good” means and you’ll get three different answers
    • The org’s idea of quality is narrower than it should be – burned by poor error handling but nobody’s ever asked for visual polish, for example
    • Specs and tickets convey function and sequence, not feel – craft, or its absence, stays invisible until something’s built
  • Timing and discipline
    • Quality is addressed late, as a pass at the end, where it’s most likely to get skipped
    • The “we’ll fix it later” lie – debt the org tells itself it’ll repay, with no actual plan, piles up and compounds
    • Estimates that never included the detail work in the first place – not cut, never planned
  • Org support
    • Speed (or scope, or date) outranks quality in what actually gets rewarded
    • Leadership says craft matters but doesn’t back it with expectation, evaluation, or investment
    • Someone benefits from the bar staying low and pushes back on raising it
  • Structural
    • Design, product, and engineering don’t share ownership of the outcome
    • Turnover and reorgs – the unspoken parts of the standard lived in people’s heads and left with them
    • Nobody measures craft after shipping, so decline is invisible until something’s obviously broken
    • Technical debt or poor design system makes craft hard to execute
    • Ownership fragmented across teams – each slice looks okay, but the whole becomes incoherent

Interventions

Some of these are useful in multiple failure modes; there’s not a clean 1:1 mapping.

  • Skill-building
    • Critique – borrow the group’s brains, praise good examples, coach the group to raise the floor
    • Coaching, training, pairing – targets individuals to raise their skill level and that of the group
    • Hiring bar and process – slower, reduces your need for remedial training but not maintenance or standards
  • Making quality visible in work processes
    • A written, concrete quality standard, with real examples instead of adjectives
    • Structured design and code review, aimed at specific quality hallmarks
    • Prototypes and working artifacts as the default way expectations are communicated – behavior is better shown than written about
  • Structural fixes
    • Shared ownership rituals for the trio – everyone takes part in delivering quality and fixing problems
    • An accountable owner for the end-to-end experience
    • Estimates that include craft from the start
    • Documentation and onboarding that carries the standard through turnover
  • Incentives and measurement
    • Align incentives – asking design for quality but engineering for speed just creates a tug-of-war
    • Usability testing, accessibility audits, quality-bug tracking, customer-facing quality metrics
  • Direct
    • Take a person aside – if someone’s in the way, they need to hear it
    • Name it plainly if the org doesn’t actually want this – less a fix than a request for honesty

Repeat?

Opening a bottleneck helps, and it reveals the next bottleneck. You might have an overall low level of designer craft, but raising it doesn’t get fully realized in the product; this tells you that there’s another obstacle after the obvious one you went after. This is true for any process you might work on. Speeding up the slowest operation in an assembly line helps, and it reveals the next-slowest operation, the next focus of intervention.

Then repeat – not as a formality, but because as you work on the process, and as business conditions change, the diagnosis doesn’t hold still. People turn over and take the tacit parts of the quality standard with them. Reorgs sever ownership. Growth outpaces whatever onboarding used to establish the bar. The causes recombine. Sometimes the same one comes back, sometimes a new one takes its place. Raising craft once is a project. Keeping it raised is the same discipline I’ve written about before under a different name – goal maintenance.

I should have said that. Instead I told a story about the wrong question.

The Case for Monochrome Data Visualization – Kay
Color is just one of the seven “retinal variables” for differentiating values or categories. It grabs attention but must be used judiciously. And there’s a case to be made that monochrome visualizations can lead to clean and uninterrupted reading when done well.

Designing with web standards: The playbook for this AI moment – Patrik Neeman
“The web got semantic markup because designers agreed on what a heading, a list, and a link meant before the tools decided for them. Do the same for AI interfaces now: agree on what “show your work,” “cite your source,” and “I am not sure” should look and behave like, as reusable components rather than one-off features. The same discipline belongs a layer down, in how a skill is described, discovered, and combined, so capabilities compose instead of colliding. If you keep a component library, this is where shared AI patterns belong.”

A side project is the fastest way to upskill in the age of AI – Phil Morton
“What I’m suggesting goes a step further: use tools like Claude Code, Codex or Cursor to build the real thing and do more of it yourself. When you’re working a little closer to the metal, then you learn more. Lovable is nothing like how products get built commercially, whereas using Claude Code to build an iOS app is. Tools that do everything for you aren’t going to teach you much.”

Creativity is fundamentally not an efficiency problem – Marchin Wichary
“Paul Cantrell: ‘Creative work keeps taking roughly the same amount of human labor/attention/care, even as new technologies accelerate or remove things that used to take time. This is because creativity is fundamentally not an efficiency problem; process is not just the means of producing output, but rather a labor vessel that holds the near-invisible work that is truly important.’”

Make It Work vs. Make It Good – Jim Nielsen
“…sometimes it’s like, good job, you made a bear ride a unicycle. Not really what bears are supposed to do – and they’ll probably never be good at it – but it’s novel and functioning! However, the task of making something good – of arriving at a solution that is obvious – is often met with a kind of ambivalence, like ‘Nice work…I guess? Seems obvious tbh.’ That’s the work of design: to make something so good, it’s obvious. But there’s often little acclaim for the obvious because, well, it’s so obvious (in hindsight).”

The Beginning of Programming as We’ll Know It –Daniel Jalkut
“But for now, real programmers will always win. Why? Because we are uniquely positioned to harness most of the power of AI while augmenting it with human taste, wisdom, and caution, among other qualities that an AI is thus far incapable of possessing.”

Craft is Untouchable – Christopher Butler
“Structure still communicates before content. Visual hierarchy still guides attention. Negative space still creates rhythm. These principles don’t vanish because I’m working through AI rather than directly manipulating pixels.”

When Shipping Software Becomes Too Easy – Julie Belião
“If shipping becomes frictionless, the real scarcity moves elsewhere: clarity of intent, product judgment, and long-term coherence. Product management becomes less about leading and coordinating work and more about protecting direction.”

The complaints that wouldn’t go away

Building is easy, taste is hard. That line’s been going around design and engineering circles for a while now, and it’s true – for a team. A single designer with good instincts can make fast, sound “taste” calls all day. But a team can’t run on one person’s taste, and it can’t run on group vibes. It needs a standard people can actually check their work against. I watched a design team, with an established design system, prove their compliance with it was real and rising, and still be dissatisfied by the results. Fixing that took an investigation.

Fixing the basics

When I arrived at Invoca as the new product design director, the complaints I heard were familiar to anyone who’s led a design org. Accusations that UX was “gold-plating” things, though nobody could point to evidence of it. Developers barely following mockups, sometimes not opening Figma at all. Engineers who didn’t come to designers with questions and chafed at the suggestion that UX should review things that were about to ship. Product managers running shuttle diplomacy, working with their designer and their engineers separately, because the two groups weren’t really talking to each other.

I fixed the relationships, restored the connection between design and engineering, built strong product trios: design, product, and engineering making key decisions together instead of individually or in pairs. I put the detail segments of UX practice into the scrum sprints instead of running apart from them. Designers’ design system compliance rose sharply, and I could show it. Engineering managers started telling me it finally felt like UX and engineering were pulling in the same direction.

And yet! Complaints about people not following the design system didn’t go away. Designers found their very compliant mockups were not realized in code. Questions were coming back to the UX team, but not about the front-end implementation. The relationships were better, but the results weren’t improving much.

Getting real

So we decided to find out what was actually happening instead of continuing to argue about it. I made the call to audit UX deliverables directly – not to defend the compliance numbers, but to make sure they were real. Jake Rowe, our design system product owner, made a gutsier call: watch the engineers work. He built a mockup he knew was 100% compliant with the design system and asked engineers to let him watch them build from it.

What that revealed had nothing to do with designers slacking off or engineers ignoring instructions. Though they sometimes claimed otherwise, engineers weren’t actually familiar with Figma or Dev Mode, where the components and variants were clearly pointed out; they’d never learned how to look there or understand what they were seeing. The difficulty went beyond this skills issue: some primitive values and other magic numbers persisted in the design system in Figma AND the design system in code, and not always the same ones. Some values were captured in tokens on one side, some on the other, some not at all, and there was drift between these two sets. Component naming didn’t always match between the two. Variants didn’t always exist in both places. Teams had built up their own private sense of which deviations were fine and which weren’t, with no shared answer. Some components existed only in Figma, or only in code, never both. And underneath all of it, engineers described feeling such pressure to ship quickly that they were reluctant to stop and ask a question that might slow them down.

None of that was a compliance problem. It was a trust problem – between two systems that were supposed to be one system, and between two groups of people who didn’t really know how the other actually worked.

The intervention

What we did about it took leadership air cover from me and our director of dev enablement. 1) Stop adding new components until the ambiguity underneath them was resolved. 2) Make the design system in code the source of truth instead of Figma, since that’s what the majority (engineers) defaulted to anyway whenever the two disagreed. 3) Reduce the decision-making load by cutting typography from eleven overlapping styles down to five, each named for how it is used. 4) Build the components that only existed on one side into both. And 5) build something we hadn’t had before: an actual way to measure whether a piece of AI-generated output matched the system, instead of arguing about it by feel.

This work would have helped regardless. It mattered twice over because of when we did it. Agentic coding was catching on at the same time, which meant more of our UI was about to be generated by something with no ability to guess what a designer meant or a developer would choose – it would just read whatever the system told it and go. And Invoca had been quietly narrowing its front-end development capacity for years. Fewer front-end specialists meant fewer people left who could bridge design system ambiguities with judgment. More AI-generated code meant more code being written by something with no judgment to bridge with at all. The gap Jake and I closed wasn’t optional cleanup. It was the difference between a system that mostly works because a person catches the edge cases, and a system that has to have far fewer edge cases. Over time we got code compliance up to 89%; not a number meant to impress anyone, a new baseline to beat.

Especially now

That’s the case for doing this work now, whether or not you’ve got our exact numbers to chase. Any company building with agentic coding tools is going to need its design system to do double or triple duty: readable by the people who use it, the coding agents they run, and by the agents composing UI from it. That means resolving the same ambiguities we found: one source of truth, not two that drift apart. Names that mean the same thing everywhere. A real way to measure whether output matches the standard, not a feeling about whether it does.

A design system was never just a component library. It’s the shared language a team uses to turn intent into working software – a system for people, and now for agents too.

This Moment We’re In, Ep. 3 – Jorge Arango
“AI as a thinking tool – i.e., a medium that allows teams to frame problems, explore possibility spaces, and inform organizational strategy. This includes improving research operations and lowering the cost of producing high-fidelity prototypes. This isn’t just the same kind of work, faster. The ability to move faster changes what kind of work can be done. Designers explore solution spaces by making – ‘make to think.’ AI lets that happen at a different level.”

The Website Specification – Joost de Valk
Perhaps the beginning of common (technical) quality standards for websites? “The Website Specification is an attempt to answer one question:regardless of the stack you build on, what should a good website do? Not a framework. Not a guide. A spec – what is required, what is recommended, and what to avoid.”

IDEO IQ 2026 Report
“For our inaugural IDEO Innovation Quotient (IDEO IQ) report, we surveyed leaders from 100 of the world’s largest companies and mapped five core cultural behaviors against real business outcomes. The results? In a world of rapid technological change, the behaviors associated with design and creativity set the conditions that support lasting growth.”

Discovery vs Delivery – Buzz Usborne
“When involved in discovery, AI is the participant in the group that says “yes” to everything with breathless enthusiasm, but rarely produces ideas that change the conversation… that generate smiles. Which is why, during discovery, the relationship needs to flip: not human-in-the-loop, but AI-in-the-loop.”

Punch Yourself in the Face with Reality – Aditya Anand
“I have seen way too many startup founders delude themselves into building more and more for months without a single conversation with a real user. The builders and the technical people really struggle with this. If all you know is how to build, and you just use AI as an excuse to keep building more and more and more, you are just procrastinating and avoiding reality.”

After Forty Years, Still No Silver Bullet – Jorge Arango
“What will the system do? How will it serve strategic objectives? How will it enable better judgment and allow people to derive meaning from data? These aren’t implementation questions, they’re design questions. Somebody must define the “construct of interlocking concepts” that define the system, aiming for good fit between the system and the context it serves. LLMs can help, but they can’t replace human understanding and judgment, at least not yet.”

Danielle asks about “validating components”

Danielle asks:

Does anyone have experience doing research on design systems? Specifically, validating existing components? Would love to learn from anyone who has done something similar, or could point me to resources like articles, etc.

There are three topics here:

  1. testing components and patterns to demonstrate their fitness and find problems to fix
  2. preventing errors when authoring components
  3. preventing errors when building things that use those components

In re the first, I’ve seen usability or accessibility problems surface that could be traced to a component and then fixed there, but I’ve not found a way to validate components out of context. You can put a bunch of components on a plausibly-constructed page and run axe-core over it to get some light accessibility auditing of those components, but that doesn’t cover a ton, it just rules out egregious mistakes.

You might have better luck testing patterns rather than components because a pattern has a natural context for the interaction: a larger assembly such as a file upload widget, with its various behaviors and states, can be put into a testable workflow (or just observed in a real workflow in the wild) and its problems detected and sorted out at the pattern and component level.

Second, if you are worried about catching mistakes when authoring new components, contracts might be your friend – for example, every INPUT element should have label association, error association, focus-visible, touch target size, zoom/reflow, etc. and an agent skill could be written to watch for those things when a PR hits the design system in code.

When Gov.UK publishes design system updates they sometimes have a blurb about what they’ve learned about that component. For example, see the “research on this component” heading on https://design-system.service.gov.uk/components/details/. Another example: this DWP Design System page about filters has a list of theories about a filter component, with some tests. Note that the context is critical.

Third, you can have instructions in your design system in code for the agents that build with it to help them fulfill your quality standards, assuming you’ve documented these in a place the agents and humans will find them. These same instructions can be used to evaluate new or changed work before a PR is merged.

Vision isn’t magic

Ideas are cheap. Every designer I’ve managed has had three ideas by mid-morning. If your org’s problem is “we don’t have enough ideas,” that’s easy to fix: creativity is a muscle, and a muscle needs reps. Generate a lot, including bad ones, and the good ones show up more often. That part is trainable.

The hard parts are choosing which idea to build, and then faithfully building it.

Choosing well requires something most orgs are worse at than they think: staying clear on what success means, and holding that clarity during the twists and turns of the project. I call it goal maintenance. It’s not glamorous. It’s also where I’ve watched more visions die than in the idea stage – not because the idea was wrong, but because a few months in the team had lost their grip on why the idea was chosen in the first place. Every debate about the product turned into a fight about what was easy for engineering or design, and not about what the customer actually needed.

Nobody kills a well-chosen idea with one bad call. They kill it with a hundred small, locally reasonable ones – a scope cut here, a shortcut there – that each look harmless and, together, push the built thing away from what it was supposed to be.

This isn’t only a product problem. Most of what gets labeled “bad engineering” is that same drift, one layer down. A quick band-aid that quietly becomes a pattern, a migration that never gets finished, a shortcut that was fine under this week’s load and isn’t fine under next month’s. No single decision was wrong on its own. Nobody was adding them up. This is the same failure as before, just harder to see at the code level than the roadmap level.

Here’s where I think “vision” gets misunderstood. Executives often want it sharp: a specific target, fully defined, implications worked out in advance. That instinct makes sense. Fuzziness feels like risk. But a vision defined that tightly is fragile and expensive. Pin down every implication before you’ve built anything, and the first real customer conversation or technical constraint breaks it, because you’ve left no room for the plan to be wrong about details you never had a way to check. All that pinning down takes patience that the org probably doesn’t have.

The alternative is a vision fuzzy enough to survive contact with reality, but attractive enough that people will fund exploring it before it’s detailed out. Creating something malleable enough to adapt, but attractive enough to invest in, is a design skill.

Coding agents change the pressures here, but the fundamentals remain. When code is cheap to generate, “we’ll figure it out as we build” starts to look like a real alternative to developing and committing to a vision at all. It’s tempting for the same reason that just pursuing your first idea is tempting: you skip the hard part. You get to start building without defending a claim about where you’re going.

But when “figuring it out as you build” works, it works because there’s implicit direction behind it. It only works if someone is doing the same goal maintenance quietly – holding a stable sense of what success looks like, near-term and long-term, and testing the current direction against it as new information comes in. Take that away, and “we’ll figure it out” is an aimless accumulation of a lot of shipped code that’s unlikely to converge on anything.

It turns out that a vision-led approach and a feel-your-way-through-the-dark approach need the same underlying discipline. Not a plan or a spec, but goal maintenance: a clear sense of what success means, near and far, and an explicit list of the assumptions the current idea depends on. Without both it doesn’t matter whether you started with a bold vision or an empty backlog.

That’s the actual differentiator – not who has the better idea, but who can tell, at any point, whether the idea they’re building is still the right one, and who wrote down, in advance, what would prove them wrong.

Read the original on jonplummer.com ↗