RSS Amplifier

Product Tribe 🔥 · Aug 25, 2026

Dear UX Designer, You Can’t Delegate Attention | Your Roadmap Is Trying to Do Fifteen Things at Once. That’s Why It Does None of Them Well

0
Sign in to vote or save

Destare Foundation, Alex Dziewulska, Sebastian Bukowski, Jakub Sirocki, Łukasz Domagała, Katarzyna Dahlke, Michał Kosecki · Product Tribe 🔥

💜 Dear UX Designer, You Can’t Delegate Attention (guest article by Michał Kosecki)

💜 Your Roadmap Is Trying to Do Fifteen Things at Once. That’s Why It Does None of Them Well (by Łukasz Domagała)

💪 Interesting opportunities to work in product management

🍪 Product Bites - small portions of product knowledge

🔥 MLA week#60

Join Premium to get access to all content.

It will take you almost an hour to read this issue. Lots of content (or meat)! (For vegans - lots of tofu!).

Grab a notebook 📰 and your favorite beverage 🍵☕.

DeStaRe Foundation

Much has been said about who a product manager should be. Almost nothing about the one competency that decides whether any of it matters.

Much has been said about who a product manager should be. Visionary. Mini-CEO. Storyteller. Decathlete. Owner of the why, voice of the customer, CEO of the product. I’ve read the decks, sat through the keynotes, and drawn a competency map of my own — nine categories, sixteen roles.

And there’s one competency that almost never makes the poster.

Humility. Intellectual honesty. The ability to say “I was wrong” in a meeting, in a full sentence, without a support group and without quietly rewriting the goal so that you weren’t.

We don’t hire for it. We don’t train it. We barely have a word for it that doesn’t sound like it belongs in a church. And its absence has killed more products than every skills gap I’ve ever diagnosed, combined.

Your ego is one of the biggest obstacles in product. Not the market. Not tech debt. Not the stakeholders, not the budget, not the org chart.

You.

Where did all this ego come from? We handed it out.

Every product curriculum opens with a keynote about a man in a turtleneck. Customers don’t know what they want. Reality distortion field. Ship the vision. The founder who ignored the focus groups, bet the house and was right. It’s a great story. It’s also a story about a founder — someone who owned the company, paid for every mistake in their own equity, and got fired for it at least once.

Then we hand that story to a salaried specialist with a Jira board, two designers and a quarterly OKR, and act surprised when they start behaving like they mortgaged a house for the feature.

You didn’t. Nobody cares that it’s your baby. You are not a spoiled founder. You are a mature specialist. And a specialist builds on data, not on say-so.

Think about what that word means anywhere outside tech. Nobody hires a surgeon for their vision of your appendix. Nobody wants a structural engineer with a strong personal brand around load-bearing walls. You hire a specialist for exactly one thing: the discipline to change the plan when the scan comes back different, and to not take the scan personally.

A founder’s ego at least has an invoice attached. They pay for it in their own runway, their own equity, their own sleep. A specialist’s ego is financed by the company and paid for by the team. Ego on someone else’s money isn’t conviction. It’s an expense account.

Behavioral economists have a name for why “my baby” is the most dangerous phrase in a product review. The IKEA effect: you overvalue what you assembled, because you assembled it. In the original experiments people paid more for a storage box they’d put together than for an identical one someone else had. Same box. Same wobble. They also rated their own amateur work as roughly on par with an expert’s — and assumed everyone else would agree.

Read that last part again. That’s not a bias about boxes. That’s a product review.

And in product, the box wobbles more often than not.

When Microsoft actually measured — well-designed experiments, on ideas teams believed in enough to build — one third moved the metric they were built to move. One third did nothing. One third made things worse. At Bing the win rate was ten to twenty percent. At Airbnb, the team Ronny Kohavi later ran tested over two hundred and fifty ideas and got about twenty wins. His experimentation team at Microsoft had a running joke about what their job actually was: telling people their new baby is ugly.

Read that again with your roadmap open. Ideas that experienced people believed in enough to fund, build and ship. Two out of three didn’t work. And those are the companies that checked. In the companies that don’t check, the success rate is a hundred percent and you can ask anyone.

It cuts the other way too. The single most profitable idea in Bing’s history — an ad headline tweak worth over a hundred million dollars a year — was prioritized low and sat in the backlog for more than six months, because nobody’s gut recognized it. Nobody’s. Not the person who proposed it, not the people who ranked it. The gut is not a product instrument. It’s a digestive one.

Meanwhile, at the same company, roughly a hundred engineers built a third pane into the search page. The experiments showed no value. It shipped anyway. The word used was “strategic.” A year and a great many more flat experiments later, it was rolled back. Remember that word. “Strategic” is what gets said when the data won the argument and the ego refused to leave the room.

Here is what ego looks like in practice, before anyone’s feelings get involved: someone has one crayon and is painting the whole picture with it.

If you have too few crayons in your pencil case, you add more. You don’t paint everything with the one you’ve got.

Philip Tetlock spent two decades tracking 284 experts and 28,000 forecasts. The ones with one big idea — the hedgehogs, the framework-for-everything people — did worst. Worse than well-read amateurs, and worst precisely on long-range calls inside their own field of expertise. They were also the most confident and had the biggest media profiles, which is its own small tragedy. The people who won carried many small models, held them loosely, and updated when reality disagreed. When Tetlock and Dan Gardner later named the foundation of that style, the word they used was humility.

Every framework-monogamist in product is a hedgehog. The JTBD-or-nothing shop. The team that has “done Shape Up” for three years and reads any other method as a threat. The consultant with one diagnostic who is somehow never surprised that every client has the same disease. That’s not expertise. That’s a sunset painted in brown, with great confidence, by someone who has never opened the rest of the box.

Intellectual honesty is just this: you notice the picture needs blue, and you go and get blue. Even if you’ve built a personal brand on brown.

Ego isn’t one thing. It fits itself to whatever chair it’s sitting in.

In the PM’s chair it’s called “my feature.” You hear it in the language long before you see it in the metrics. The initiative has a codename and a Slack emoji before it has a problem statement. Discovery gets done — three interviews, all with people who already liked the idea, and the word “validated” in 48-point font. The PRD reads like a letter to a future biographer. The roadmap, if you squint, is an autobiography.

Nobody hired you to author a feature. They hired you to find out what works. Those are different jobs, and only one of them has an ego in it.

In the head of product’s chair it’s called “my org.” The feature ego graduates into an empire ego. The hiring bar quietly becomes “agreed with me at the interview.” The process gets a name, and the name has your initials in it. The product review is a ritual in which the team learns what the answer was before they’ve finished presenting. And the head of product — ask around — hasn’t changed their mind in public in two years, which the team reads correctly: disagreement here isn’t feedback. It’s insubordination.

Chris Argyris wrote the definitive piece on this thirty-five years ago and it hasn’t gotten any more comfortable. The most successful professionals are the worst at learning, because they’ve almost never failed and so never learned how to learn from it. So when something finally goes wrong, they get defensive, screen out the criticism, and find someone else to blame. They don’t dig in because they’re stupid. They dig in because being wrong isn’t a data point for them. It’s an identity event.

In the sponsor’s chair it’s called “strategic.” The ego wears a budget now, and a name. The initiative is on a board slide with the sponsor’s face next to it. Killing the initiative means killing the face. So it doesn’t get killed. It gets “a phase two.” It gets “a broader mandate.” It gets a hundred engineers and a third pane. Barry Staw described this fifty years ago as escalation of commitment — the more of yourself you’ve sunk into a bad decision, the more you’ll sink to avoid admitting it — and every quarter, somewhere, a sponsor is proving him right with other people’s headcount.

The sponsor isn’t a founder either. The sponsor is a steward of capital that belongs to shareholders, or taxpayers, or a founder who is not them. Treating the money like a personal bet isn’t boldness. It’s a category error with a budget line.

In the consultant’s chair it’s called “I’ve seen this before.” And I’ll put my own chair in the line-up, because leaving it out would be exactly the thing this piece is about.

I’m paid for pattern recognition. Walk in, minimal data, find the root cause. That is the job — and it’s also precisely the mechanism that turns a diagnostician into a hedgehog. Pattern recognition curdles into pattern imposition the moment you’d rather be right than useful. The consultant with one methodology sells the crayon. The client gets a brown sunset and a very nice deck.

The most valuable sentence I say to clients is “I was wrong about that on Monday — here’s what Tuesday’s data says.” The second most valuable is “I don’t know yet.” Neither is in any consulting playbook I’ve been handed. Both are the reason people come back.

None of this is a personality trait. Nobody is born humble and nobody becomes humble in a workshop. It’s a practice, and a cheap one — which is what makes its absence so damning.

Write down what you think will happen before you run the test. A number. Then read it afterwards. That’s it. That’s the whole training. It’s the cheapest calibration exercise on the market and almost nobody does it, because the results are private, precise and humiliating.

Say “I don’t know” in the meeting, not in the retro.

Change your mind where people can see it. Not in a DM. Not “on reflection.” In the room, out loud, with the thing you were wrong about named. Every time you do that in front of a team, you give them permission. Every time you don’t, you give them a warning.

Humility isn’t the absence of conviction. It’s conviction with an expiry date. Not “I have no opinion” — “my opinion is a hypothesis, and it has a budget.”

I don’t care about your ego. I care about it even less the moment it starts costing the product and the people building it.

Nobody cares that it’s your baby. The users don’t. The market doesn’t. The data certainly doesn’t.

The only person who ever did was you.

Share

There’s a conference happening 270 kilometers from my desk in September, and most of the product people I know in Poland still book flights to London or Lisbon to see the same speakers. I want to fix that.

WaysConf is in Kraków on 16–17 September, with a separate workshop day on the 15th. Marty Cagan and Brad Frost are keynoting. Debbie Levitt and Petra Wille are running workshops. Then it goes deep on working practitioners rather than circuit speakers — staff designers from Meta and Shopify, research leadership from Wise, a Principal AI Designer from Intercom, product leaders from Kraken and IKEA.

Here’s my one piece of unsolicited advice, and it’s the contrarian one: do not build your two days around the keynotes. You can watch Cagan on YouTube tonight, in your pajamas, for free. What you cannot get on YouTube is the case study about what actually broke, the roundtable where people who run research orgs argue about maturity, or the speaker who contradicts the one you heard ninety minutes earlier. Conferences that only platform one worldview are just expensive newsletters read aloud. WaysConf isn’t that. The speakers counter-argue each other, and that’s the point.

The theme this year is “Building What Matters” — good decisions under real constraints, in an AI-accelerated, ship-faster world. I spend most of my time on the gap between what the AI-product discourse claims is happening to teams and what is actually happening to teams. A room full of practitioners presenting real case studies, in our region, is exactly where that gap gets argued out by people living it rather than people monetizing the narrative about it.

And the unglamorous part: if you run a small studio or sit on a Polish team budget, a Kraków train ticket and a conference pass is a different financial event than flying a team to San Francisco. Same caliber of speaker, fraction of the friction. The room is people you’ll keep running into for the rest of your career, mostly within a few hours of where you already work.

So go. Block the 16th and 17th, decide whether the workshop day is worth the extra ticket, and spend your time in the case-study rooms and the hallway — not the front row of every keynote.

The keynotes will be on the internet by October. The conversations won’t.

WaysConf 2026 · 16–17 September (workshops 15 Sept) · EXPO Kraków + online · waysconf.com

Do you need support with recruitment, career change, or building your career? Schedule a free coffee chat to talk things over :)

  1. Product Manager - Asana

  2. Product Manager - Ampstek

  3. Product Manager - Allegro

  4. Senior Product Manager - Ideals

  5. Product Lead - BJAK

Refer a friend

Bundling gets treated as a packaging choice. Someone in the room asks whether the three products should be sold together or separately, the conversation turns to what customers seem to prefer, and a decision emerges from intuition and competitive comparison.

This is a mistake with a well-established alternative. Bundling has been formally analyzed in economics for decades, and the analysis produces a clear answer to when it works, when it fails, and why. The determining factors are not customer preference or competitive positioning. They are marginal cost and the distribution of how much different customers value the individual components.

Get those two right and the bundling decision largely makes itself. Get them wrong and no amount of packaging creativity will rescue the outcome.

A SaaS company has three products. Sales has been selling them separately, and the data shows a pattern: most customers buy one, some buy two, almost nobody buys all three. Each product has a passionate constituency and an indifferent majority.

The team debates bundling. The argument against is intuitive: if most customers only want one product, forcing them to buy three will feel like paying for things they do not need. The argument for is vaguer — something about simplicity and higher contract values.

The economics resolve this cleanly, and they favor bundling, for a reason that is not obvious from the sales data. Precisely because customers disagree about which product is valuable, the total value each customer places on all three is more predictable than their value for any one of them. That predictability is what makes bundling profitable — not customer enthusiasm for the bundle.

Bundling is the practice of selling multiple goods together for a single price, and the decision of whether to do so is governed by a body of economic analysis with clear, testable predictions.

The foundational modern treatment for software and digital products comes from Yannis Bakos and Erik Brynjolfsson, whose 1999 paper in Management Science analyzed bundling of information goods. Their central finding is what they called the predictive value of bundling: the law of large numbers makes it substantially easier to predict what consumers will pay for a bundle of goods than for the individual goods sold separately.

The mechanism is worth stating precisely. Individual customers vary enormously in how much they value any single component — one loves the analytics module and ignores the collaboration features, another the reverse. But when you sum across many components, those individual variations average out. The distribution of total valuations across customers is much tighter than the distribution for any one item. A seller facing a tight distribution can price close to the average and capture most of the available value. A seller facing a wide distribution must choose between pricing high and losing most customers, or pricing low and leaving money on the table with the enthusiasts.

Carl Shapiro and Hal Varian made the same point accessible in Information Rules, noting that information goods are characterized by very low marginal cost of production and by buyer uncertainty about value until the good is obtained — two properties that both favor bundling.

The single most important variable is marginal cost, and it operates as a threshold rather than a gradient.

When marginal cost is zero or near zero, bundling increases efficiency whenever it increases the fraction of customers who purchase, because each additional purchase creates value at no additional cost. Bakos and Brynjolfsson’s analysis shows that bundling very large numbers of unrelated information goods can be surprisingly profitable under these conditions.

When marginal cost is high, the logic reverses. If the marginal cost of a component exceeds the average customer’s valuation of it, bundling reduces profit — because the seller is now paying to deliver components that most recipients do not value. Bakos and Brynjolfsson are explicit that their main results do not extend to most physical goods, precisely because production costs for unused components negate the predictive benefit.

This is why the bundling conversation looks so different in software than in hardware, and why software companies that reason about bundling using retail intuitions reach wrong conclusions.

The second determinant is how much customers disagree about the value of individual components.

High dispersion — customers who value the same component very differently — is what bundling is good at solving. The bundle averages across the disagreement. Low dispersion means everyone values the component similarly, which means separate pricing already captures most of the value and bundling adds little.

This produces a counterintuitive rule: bundling works best when components appeal to different constituencies, not when they appeal to the same one. Bundling three products that the same customers all want is much less valuable than bundling three products that different customers each want one of. The instinct to bundle “things that go together” is often backwards.

The literature distinguishes several structures. Pure bundling offers only the bundle. Mixed bundling offers both the bundle and the individual components. Customized bundling lets the customer select a fixed number of components from a larger set.

The research finds that pure bundling captures nearly the entire available value when marginal costs are near zero and the number of goods is large. But when marginal costs are non-trivial or when demand for components is correlated, mixed bundling can be substantially more profitable — because it allows the seller to capture the enthusiast who would pay a premium for one component while still capturing the moderate buyer through the bundle.

Most SaaS pricing pages are implementing mixed bundling without naming it. The tiers are bundles; the add-ons are the unbundled options.

A less-discussed advantage: information goods are experience goods, meaning buyers cannot assess value until after use. Bundling introduces components to customers who would never have bought them separately, because the marginal decision cost of trying an included component is near zero.

Bakos and Brynjolfsson note that periodically updating bundle composition — so that any given customer faces an attractive mix of new and familiar goods — can be more effective at overcoming buyer resistance than selling separately. For product companies, this means bundling is not only a pricing mechanism but a distribution mechanism for new capabilities.

Microsoft 365 is the clearest large-scale demonstration of the predictive value of bundling. The suite contains applications with radically different constituencies — heavy Excel users who never open Publisher, Teams users who rarely touch Access, and so on. Almost no customer values every component. This dispersion is exactly the condition under which bundling maximizes capture: the total value each organization places on the suite is far more predictable than their value for any individual application, allowing Microsoft to price near the average and capture most customers. Attempting to sell each application at its individually optimal price would leave substantial value uncaptured across the population.

Adobe’s transition from perpetual licenses to Creative Cloud restructured both the payment model and the bundle simultaneously. Previously, a photographer bought Photoshop and a video editor bought Premiere, with each transaction capturing only that customer’s valuation of that product. The Creative Cloud bundle captures the aggregate across a professional’s whole toolkit, and — critically — exposes users to applications they would never have purchased separately. Adobe reported significant usage of secondary applications by subscribers who had never bought them under the perpetual model, which is the discovery mechanism operating as the theory predicts.

Spotify demonstrates bundling at extreme scale with the marginal cost condition fully satisfied. The catalogue contains tens of millions of tracks, and no subscriber values more than a tiny fraction of them. Under per-track pricing, the seller faces enormous valuation dispersion for each individual song. Under bundle pricing, the aggregate is highly predictable — most subscribers value unlimited access at roughly similar levels, despite listening to entirely different music. The zero marginal cost of an additional stream is what makes the arrangement work; the same bundle economics would be impossible if each play cost the seller money.

The bundling decision matters because it is made frequently, made intuitively, and has large revenue consequences that the intuitive process does not consider.

The most common error is bundling components that appeal to the same customers, which adds little value, while unbundling components that appeal to different constituencies, which forfeits the averaging benefit. This happens because product logic groups things by functional similarity, while bundling economics groups things by valuation dispersion. The two produce different answers, and product logic usually wins the argument.

A second error is applying bundling reasoning across a marginal cost boundary. A software company that bundles its services offering — where each engagement has real delivery cost — using the logic that works for its software may be systematically destroying margin. The threshold effect in the research is sharp: below it, bundling adds value; above it, bundling subtracts.

There is also a strategic dimension around unbundling. The same analysis that explains why bundles are profitable also explains their vulnerability: a competitor who serves one high-value component exceptionally well can attract the customers whose valuation of that component was subsidizing the rest. This is the mechanism behind most successful unbundling attacks, and it is predictable from the bundle’s own economics.

Start with marginal cost, before anything else. For each component under consideration, determine the true marginal cost of serving one additional customer. If it is near zero, the economics favor bundling and the remaining question is structure. If it is substantial relative to what an average customer would pay for that component, bundling will likely reduce profit regardless of how attractive the package looks.

Measure valuation dispersion rather than assuming it. The relevant question is not whether customers like a component but how much they disagree about it. Willingness-to-pay research, usage distribution data, and standalone purchase patterns all provide signal. High dispersion favors bundling; low dispersion means separate pricing already captures the value.

Bundle across constituencies, not within them. Components that serve different user types produce more averaging benefit than components that serve the same user type. When constructing bundle tiers, deliberately combine capabilities whose enthusiasts do not overlap — this is the opposite of the functional grouping instinct, and it is what the economics recommend.

Consider mixed bundling when marginal costs are non-trivial. Offering both the bundle and standalone components lets the seller capture the enthusiast willing to pay a premium for one thing while still capturing the moderate buyer through the package. The research shows this can substantially outperform pure bundling when costs are not zero or when component demand is correlated.

Use bundle composition as a distribution channel. Adding a new capability to an existing bundle exposes it to the entire customer base at zero marginal decision cost for them. For products where adoption is the bottleneck rather than willingness to pay, this is frequently a stronger launch mechanism than standalone pricing.

The reason bundling feels like a marketing decision is that its effects are visible on the pricing page, which is where marketing lives. But the pricing page is the output. The input is a question about cost structure and preference distribution that has a right answer independent of how the page looks.

What Bakos and Brynjolfsson identified is genuinely counterintuitive: the fact that customers disagree about component value is not an obstacle to bundling but the source of its power. Disagreement averages out. Averaging produces predictability. Predictability lets the seller price near the mean and capture nearly everyone.

The corollary is equally useful. When customers agree about value — when everyone wants the same thing to the same degree — bundling has nothing to average, and separate pricing is likely correct.

Check the marginal cost. Measure the disagreement. The packaging follows.

Leave a comment

In the late 1960s, Daniel Kahneman was lecturing Israeli Air Force flight instructors on the psychology of training. He presented the research: rewarding good performance produces better learning outcomes than punishing bad performance. The evidence was strong and the conclusion was uncontroversial among psychologists.

One of the instructors objected, and his objection was grounded in years of direct observation. When he praised a cadet for a beautifully executed maneuver, the cadet’s next attempt was usually worse. When he screamed at a cadet for a poor execution, the next attempt was usually better. His experience contradicted the research, and the room agreed with him.

Kahneman checked. The instructors were right about the pattern. They were entirely wrong about the cause.

Performance on a difficult maneuver varies. An exceptionally good execution reflects skill plus a favorable run of luck; the next attempt will likely be closer to the pilot’s actual average, which means worse. An exceptionally bad execution reflects skill plus an unfavorable run; the next attempt will likely be better. This happens regardless of what the instructor says or does.

The instructors had spent years being systematically rewarded for punishing and punished for praising, by a statistical artifact.

A product team notices that activation rate dropped sharply in March — the worst month in over a year. They investigate, identify a friction point in the signup flow, and ship a fix in April.

May’s activation rate is substantially better than March’s. The team writes it up as a win. The fix goes into the quarterly review as a case study in responsive product work. Confidence in the diagnostic process increases, and the same approach is applied to the next anomaly.

The problem is that March was an outlier. Activation rate fluctuates month to month around a stable mean, and an unusually bad month is followed by a better one whether or not anyone intervenes. The team has no counterfactual — they never observed what May would have looked like without the fix — and the improvement they measured is fully consistent with the fix having done nothing at all.

They may have fixed something real. They have no way to know, and they have concluded that they did.

Regression to the mean is the statistical tendency for extreme measurements to be followed by measurements closer to the average, whenever the measured quantity contains any random variation.

The mechanism requires no explanation beyond arithmetic. Any observed value is a combination of a stable underlying component and random noise. When a value is extreme, it is likely that the noise component was extreme in the same direction — because that is what makes a value extreme. The noise component does not persist. The next measurement will have its own independent noise, which is on average zero, so the next value will tend to sit closer to the underlying level.

The critical property is that this occurs with no causal intervention whatsoever. It is not a force pulling values toward the average. It is a consequence of how extremes are selected.

The associated error has a name: the regression fallacy, which is the attribution of regression-driven change to some intervention. Kahneman and Amos Tversky used the flight instructor case in their 1974 Science paper and Kahneman returned to it in Thinking, Fast and Slow, drawing out the uncomfortable general implication: because we tend to intervene when performance is bad and relax when performance is good, we are exposed to a lifetime schedule in which we are most often rewarded for punishing and punished for rewarding.

Regression appears whenever a group or period is selected for measurement because it was extreme. This is the operative condition, and it describes a large fraction of product investigation.

Teams do not investigate average weeks. They investigate the week the metric cratered, the cohort that churned unusually fast, the segment with unusually poor conversion, the customer whose usage collapsed. Every one of these is a selection on an extreme value, which means every one of them will show improvement at the next measurement regardless of what is done.

The investigative reflex — look at what went wrong, fix it, observe improvement — has regression built into its structure. The improvement is guaranteed by the selection method.

The organizational version compounds this. A poor quarter triggers a response: new initiatives, leadership attention, process changes, sometimes reorganization. The next quarter is better. The response is credited and institutionalized.

A good quarter triggers relaxation. The next quarter is worse. The lesson drawn is that complacency is dangerous and vigilance is required, which reinforces the pattern.

Over multiple cycles, this produces an organization confident in a set of interventions that have never been tested against a counterfactual, and a folk theory that pressure produces results. The theory is supported by consistent observation, and the observation is an artifact.

Properly randomized A/B tests are protected against regression, because both arms experience it equally. The protection disappears in several common situations.

Pre-post comparisons — measuring before a change and after — have no control group and inherit the full regression effect whenever the “before” period was selected for being unusual. Tests targeted at a struggling segment inherit it, because the segment was selected on an extreme. Tests launched in response to a metric drop inherit it, because the launch timing was determined by the extreme.

The pattern to watch for is any experiment whose existence was triggered by bad numbers. The trigger is a selection on an extreme, and the subsequent improvement is partly guaranteed.

Regression also manufactures apparent failures. A feature that launches immediately after an unusually good period will appear to have hurt performance, because the following period returns to normal. Teams have killed working features on this basis.

The asymmetry in how these errors are received matters. A false positive gets celebrated and institutionalized. A false negative gets quietly shelved. Organizations accumulate confidence in interventions that did nothing and abandon interventions that worked, and both errors trace to the same statistical mechanism.

Netflix’s experimentation infrastructure is built around randomized control as the default rather than the exception. The company runs thousands of concurrent A/B tests, and the organizational insistence on control groups — rather than measuring before and after a change — is precisely what protects against regression contamination. Netflix engineers have written publicly about the risks of drawing conclusions from non-randomized comparisons, including the specific problem of evaluating changes launched in response to metric movements. The infrastructure investment exists because the alternative produces confident wrong answers at scale.

Booking.com’s experimentation culture, which routinely runs over a thousand concurrent tests, reflects a similar structural commitment. The company has published on the frequency with which intuitively-obvious improvements fail to replicate under proper randomization. A meaningful portion of that gap is regression: changes that appeared to work in observational analysis, because they were made in response to unusual periods, and that show no effect when tested against a genuine control.

Sports analytics provides the cleanest illustration of the general phenomenon. The Sports Illustrated cover jinx — the observation that athletes featured on the cover subsequently perform worse — has been repeatedly analyzed and is fully explained by regression. Cover appearances follow exceptional performance; exceptional performance is followed by more typical performance. There is no jinx, and the same structure explains the rookie-of-the-year sophomore slump and the reliability with which a manager hired after a losing streak appears to turn a team around.

Regression to the mean matters for product teams because it manufactures evidence for interventions that did nothing, and it does so most reliably in exactly the situations where teams are most motivated to see results.

The learning consequence is the serious one. A team that credits regression to its own intervention has not learned that the intervention works — it has learned a false thing, and it will apply that false thing again. Over time, the organization’s accumulated body of product knowledge fills with practices that were validated by an artifact.

There is also a resource allocation cost. Effort directed at investigating and fixing extreme observations is effort not directed at the systematic patterns that persist across normal periods. The extreme observation is by definition unusual; the average performance is what determines outcomes. Optimizing the tails while the mean goes unexamined is a predictable consequence of investigating only what looks alarming.

The reputational dynamic compounds this. Teams that respond visibly to bad numbers and observe improvement build organizational credibility, which increases the resources directed at reactive investigation. The credibility is real; the underlying causal claim frequently is not.

Treat any analysis triggered by an extreme observation as regression-contaminated by default. When an investigation begins because a number was unusually bad, the subsequent improvement carries no evidential weight on its own. This does not mean the intervention was useless — it means the observed improvement cannot distinguish between a working intervention and no intervention. Name this explicitly in the writeup rather than reporting the improvement as a result.

Require a control group for any claim of causal effect. The only reliable protection is randomization, where both arms experience regression equally. Where a full A/B test is impossible, a matched comparison group experiencing the same period without the intervention provides partial protection. Pre-post comparison provides none.

Establish the baseline range before interpreting any single period. Most metrics fluctuate within a band. Knowing the historical variance makes it possible to distinguish an unusual value from a meaningful change. A month that is bad but within the normal range is not an event requiring explanation, and treating it as one initiates the regression cycle.

Predict the counterfactual before intervening. Before shipping a fix in response to a bad period, write down what the metric would likely do with no intervention at all, based on historical variance. This forecast becomes the comparison. An improvement that matches the no-intervention prediction is not evidence of anything.

Watch for the reverse error in feature evaluation. When a launch follows an unusually strong period and the subsequent numbers decline, check whether the decline is consistent with regression before concluding the feature caused harm. Features have been killed on the basis of a return to normal.

What makes the flight instructor story enduring is that the instructors were not careless. They had extensive direct observation, a consistent pattern, and a plausible causal story. Everything about their reasoning was sound except that the pattern had a different cause than the one they assigned to it.

Product teams occupy the same position routinely. The dashboard shows a real pattern. The intervention preceded the improvement. The causal story is plausible. And underneath, a statistical process is generating the same observations that a working intervention would generate, with no way to distinguish between them from the data alone.

The only reliable defense is structural: a control group, or an explicit prediction of what would have happened anyway. Everything else — better analysis, more data, more careful reasoning about mechanism — leaves the ambiguity intact.

The metric got better. The question is whether you did that.

Most of the time, without a control, there is no honest way to know.

Leave a comment

Trust is discussed in product organizations as an outcome. We say a product feels trustworthy, that users trust the brand, that trust was damaged by an incident. The language treats it as a property that emerges from doing good work over time — something earned rather than designed.

This framing is partly true and practically useless. It offers no guidance about what to build, no way to audit the current state, and no mechanism for improvement beyond continuing to be good and hoping.

The more useful framing is that trust is produced by specific, identifiable, designable signals — and that a product’s trust position is the sum of those signals, most of which are shipped without anyone considering their trust function at all.

A B2B SaaS product has a strong conversion rate on its self-serve tier and a persistent problem converting mid-market prospects. The product is good, the feature set is competitive, and the sales team reports that prospects consistently like the demo.

The team assumes the problem is feature gaps and builds toward parity with the enterprise incumbents. Conversion does not move.

An audit of the buying experience reveals something different. The pricing page shows two tiers and a “Contact us” for anything larger, with no indication of range. The security page consists of one paragraph. The customer logos are all early-stage startups. The permissions requested during integration setup are broad and unexplained. Nothing on the site indicates how many people work at the company or who they are.

Every one of these is a trust signal, and every one of them is transmitting the same message: this is a small company that may not be around in three years, and we cannot tell what you will do with our data. The feature gap was never the obstacle.

A trust signal is any element of a product experience that a user reads — usually unconsciously — as evidence about whether the product and the organization behind it can be relied upon.

The research on this has become substantially more granular than the general advice to “be trustworthy.” Studies of credibility judgment consistently find that users form initial trust assessments within tens of milliseconds, based primarily on visual design quality, before evaluating any substantive claim. Design polish operates as a proxy for organizational competence.

Beyond first impressions, the literature identifies distinct trust dimensions that operate somewhat independently: security perception, privacy transparency, expertise demonstration, interface quality, fulfillment reliability, human presence, and error handling. A product can be strong on several and weak on one, and the weakness in a dimension that matters to a particular buyer will dominate.

What makes this designable rather than merely descriptive is that each dimension has concrete, shippable expressions. Security perception is affected by whether a security page exists and what it contains. Privacy transparency is affected by how permission requests are worded. Human presence is affected by whether the team is visible anywhere. These are product decisions, currently made without reference to their trust function.

The most important finding in the trust research is that different trust signals matter in different contexts, and applying the wrong ones wastes effort or backfires.

Research on security badges illustrates this sharply: they increase form completion rates meaningfully in financial services contexts, where users are anxious about data exposure, and have minimal or negative effect on content subscription signups, where the anxiety is about recurring charges rather than security. The Baymard Institute has found that a substantial share of e-commerce cart abandonment traces to security concerns — which is exactly why security signaling belongs at checkout and is largely wasted on a blog.

For consumer subscription products, research indicates that prominent cancellation clarity and transparent pricing reduce friction more effectively than customer testimonials, because the operative anxiety is about being trapped rather than about legitimacy. Applying testimonial-heavy social proof to a subscription anxiety problem addresses the wrong fear.

The design implication is that trust architecture starts with identifying the specific anxiety present at each decision point, not with deploying a standard trust package.

A consistent finding across social proof research is that specific, relevant references outperform impressive aggregate numbers. A statement naming a count of companies resembling the prospect outperforms a much larger total customer number, because the reference group relevance is what does the work.

This extends beyond customer counts. A testimonial from a Fortune 500 company may actively intimidate a small business buyer, signaling that the product is built for someone else. Volume-based social proof optimizes for a magnitude that does not address the buyer’s actual question, which is whether people in their situation have succeeded here.

Trust badges, certifications, and third-party seals have become sufficiently ubiquitous that their differentiating value has eroded. Every site displays them regardless of underlying practice, and users have learned to discount them accordingly.

The research suggests a shift toward substantive transparency as the primary mechanism: explaining how the product works, how data is used, how the company operates. Companies that explain their processes clearly build stronger trust than those relying primarily on borrowed credibility markers. This is a meaningful strategic change, because transparency is harder to fake and therefore carries more information.

There is a documented failure mode sometimes called defensive design anxiety: crowding an interface with trust signals produces cognitive overload and reduces conversion. A checkout page covered in badges, guarantees, testimonials, and security assurances signals not confidence but the opposite — that the seller anticipates being doubted.

Trust architecture is therefore a selection problem rather than an accumulation problem. The right signal at the right moment, and restraint elsewhere.

Some of the strongest trust signals are not trust-labeled at all. Load speed operates as a competence signal — users read slow or glitchy behavior as evidence of organizational unreliability. Accessibility signals an ethical stance and a standard of care. Error handling quality signals whether the organization thought about what happens when things go wrong.

These are typically owned by engineering and evaluated on functional criteria alone. Their trust contribution is real and unmeasured.

Stripe’s approach treats documentation and transparency as the primary trust mechanism rather than badges. The company publishes detailed security practices, maintains a public status page with historical incident data, and documents its API behavior — including failure modes — in unusual depth. For a company asking developers to route customer payments through its infrastructure, the operative anxiety is reliability and comprehensibility. Stripe’s answer is substantive disclosure rather than borrowed authority, and the developer community’s trust in Stripe is regularly attributed to exactly this.

Airbnb’s trust architecture addresses a genuinely hard problem: convincing strangers to sleep in each other’s homes. The company’s signal set is unusually explicit — verified identity, two-sided reviews that both parties must complete, host response rates, cancellation policies stated plainly, and a guarantee structure that names what happens when something goes wrong. Each element targets a specific fear rather than generic reassurance. Airbnb’s research and design work on trust has been extensively documented, and the through-line is that marketplace trust required designing for the specific anxieties of both sides rather than deploying general credibility markers.

Basecamp’s public commitments function as trust signals through constraint. The company states plainly that it does not track users across the web, does not sell data, and does not use engagement-optimizing dark patterns — and its pricing page shows a single price with no hidden tiers. These are not badges; they are statements that would be costly to violate publicly, which is precisely what makes them credible. The signal derives its value from the constraint it imposes, not from the assertion.

Trust signal architecture matters because trust operates as a gate on every other product quality. A user who does not trust the product will not enter data, will not connect integrations, will not upgrade, and will not recommend — regardless of how good the underlying capability is.

The organizational problem is ownership. Trust signals are distributed across marketing, design, engineering, legal, and security, and no one is accountable for the aggregate. The pricing page belongs to marketing, the permission dialogs to engineering, the security page to security, the error messages to whoever wrote them. Each is optimized for its local objective, and the composite trust position is whatever emerges.

This is why trust problems are so often diagnosed as something else. The team sees a conversion problem and builds features. The actual obstacle is a set of signals that no one is measuring because no one owns them.

There is also a compounding asymmetry. Trust signals accumulate slowly and can be destroyed quickly. A single incident, a single deceptive pattern, a single unexplained data practice can undo years of consistent signaling. This asymmetry argues for treating trust architecture as risk management rather than as growth optimization.

Inventory the trust signals currently being transmitted. Walk the full user journey — landing page, pricing, signup, permission requests, first use, error states, billing, cancellation — and at each point, name what the experience communicates about reliability. This is an audit, not a design exercise, and it typically reveals that the strongest negative signals are in places nobody considered part of the trust surface.

Identify the specific anxiety at each decision point. Trust signals only work when matched to the fear actually present. At checkout, the fear is usually security. At subscription signup, it is usually entrapment. At integration setup, it is usually data access. At enterprise evaluation, it is usually organizational durability. Deploying the wrong signal is wasted effort and sometimes worse.

Prefer specific references to aggregate claims. Replace large customer counts with counts of comparable customers. Replace generic testimonials with cases from organizations the prospect would recognize as similar. Relevance of the reference group is what produces the effect.

Treat functional quality as trust infrastructure. Load speed, error handling, uptime transparency, and accessibility are trust signals owned by teams that do not think of them that way. Making the trust contribution explicit changes how they are prioritized.

Audit for signal overcrowding. Count the trust elements on high-stakes pages. If the checkout or signup page is dense with badges, guarantees, and reassurances, the cumulative message may be defensive rather than confident. Removing the weakest signals often improves the strongest ones.

The reason trust is treated as an outcome rather than an object of design is that it cannot be shipped directly. There is no trust feature. There is no sprint in which trust gets built.

What can be shipped is every individual signal that users read as evidence — and those are shipped constantly, by teams optimizing for other objectives, with no one asking what the composite communicates.

The research is clear that users make trust judgments fast, from visual and structural cues, before evaluating substance. This means the trust position is largely set by decisions that were never framed as trust decisions: how the pricing page is structured, how a permission dialog is worded, whether a status page exists, how an error message assigns blame.

None of these is a trust project. All of them are trust architecture.

You cannot build trust. You can build the things people read it from.

Leave a comment

The Minimum Lovable Action (MLA) is a tiny, actionable step you can take this week to move your product team forward—no overhauls, no waiting for perfect conditions. Fix a bug, tweak a survey, or act on one piece of feedback.

Why it matters? Culture isn’t built overnight. It’s the sum of consistent, small actions. MLA creates momentum—one small win at a time—and turns those wins into lasting change. Small actions, big impact

An email lands in your inbox that reads as terse, dismissive, almost hostile. A stakeholder ships a decision that seems to deliberately undercut your roadmap. A colleague doesn’t loop you into a conversation you clearly should have been part of. In each case, a story forms almost instantly — and it’s rarely a charitable one. They’re playing politics. They’re protecting their turf. They don’t respect my work.

The story feels like insight. It’s usually a mistake.

Hanlon’s Razor is a mental model with a deceptively simple instruction: never attribute to malice that which can be adequately explained by incompetence, ignorance, or simple oversight. It’s called a razor because, like Occam’s Razor, it shaves away unnecessary complexity — except where Occam cuts surplus assumptions in explanations, Hanlon cuts surplus assumptions about people’s motives. The reasoning behind it is statistical, not naive. Malice requires intention, planning, and effort. Error requires none of those things — it’s the natural state of busy, distracted, imperfect people operating with incomplete information. In any given negative interaction, the base rate of ordinary human error vastly exceeds the base rate of deliberate hostility. So when you don’t have evidence to distinguish between the two, the simpler explanation is almost always the more accurate one.

For product managers, this matters more than for almost any other role — because the PM sits at the intersection of more competing agendas than anyone else on the team. Engineering, design, sales, marketing, finance, leadership, customers: every one of these has different incentives, different information, and different pressures. Friction is constant. And each moment of friction is an opportunity to misread ordinary error as intentional opposition.

The cost of that misreading is specific and compounding. The person who consistently reads incompetence as malice tends to generate real conflict in the process of imagining conflict that wasn’t there. Your defensive response — the sharp reply, the escalation, the quiet withdrawal of goodwill — becomes visible to the other person, who now genuinely does have a reason to be difficult with you. The suspicion becomes self-fulfilling. Meanwhile, the actual root cause — a miscommunication, a missing piece of context, a competing priority nobody named — goes unaddressed, because you were busy solving the wrong problem.

This effect is amplified in exactly the environment most product work happens in: text-based, asynchronous, tone-absent communication. An email that reads as cold is usually just efficient. A message left unanswered for two days was usually lost in an inbox, not withheld as a power move. The interpretive gap between what was intended and what was received is one of the most reliable generators of unnecessary conflict in professional life.

This MLA is a single application of the razor to a real situation. Not to make you naive — malice does exist, and pretending otherwise is its own error. But to make the charitable explanation your starting point, requiring evidence before you upgrade to the hostile one.

Step 1: Identify one current situation where you’ve assumed bad intent

Think about the last week or two. Is there a stakeholder interaction, an email, a decision, or a behavior that you’ve filed under “they’re being difficult,” “they’re playing games,” or “they don’t respect what we’re doing”? Pick one specific instance — a concrete moment, not a general grievance about a person.

Write it down in one sentence: what happened, and what intention you attributed to it.

Step 2: Write out the malice explanation in full

Give your suspicion its full voice. What’s the story you’ve been telling yourself? They deliberately excluded me because they want to control this decision. Or: They shipped that without consulting me because they don’t think my input matters. Write the uncharitable interpretation completely, without softening it. You need to see it clearly to test it.

Step 3: Now generate three non-malicious explanations

This is the core of the exercise. For the same behavior, write three alternative explanations that don’t require any bad intent. Common categories:

Incompetence or error: They forgot. They made a mistake. They didn’t realize the implication. The system failed them. They’re not as organized as you assumed.

Missing information: They didn’t know you needed to be involved. They weren’t aware of the dependency. They were working from an outdated understanding of who owns what. They never received the context you assumed they had.

Competing pressure: They were under a deadline you don’t know about. They were solving a different problem that made your concern invisible to them. They were responding to a priority from above that you’re not aware of. They were as overwhelmed as you sometimes are.

Force yourself to write all three, even if the malice explanation still feels most true. The point isn’t to conclude that malice is impossible — it’s to see that plausible alternatives exist.

Step 4: Assess the actual evidence

Now ask the hard question: What actual evidence do I have that distinguishes the malice explanation from the innocent ones?

Not vibes. Not tone. Not a pattern you’ve assembled from other interactions that might themselves be misread. Concrete evidence: something they said explicitly, a documented action that only makes sense if the intent was hostile, a direct statement of adversarial intent.

Most of the time, when you look for this evidence, you find that it’s thinner than the confidence of your suspicion suggested. The story felt certain. The evidence is ambiguous. That gap is the finding.

Step 5: Choose your response based on the most likely explanation, not the most threatening one

Given what you found, what’s the most probable explanation? If it’s error, missing information, or competing pressure — respond to that. This usually means a curious, non-accusatory question rather than a defensive move:

  • “I noticed I wasn’t in that decision — was there a reason, or did it just move fast?”

  • “I might be missing context here. Can you walk me through how this came together?”

  • “I want to make sure I understand the priorities driving this — help me see what you’re seeing.”

These questions do something a defensive response can’t: they gather the information that would actually tell you whether malice is present, while leaving the relationship intact if it isn’t. If there genuinely is bad intent, a curious question loses you nothing. If there isn’t — the far more likely case — you’ve avoided manufacturing a conflict.

Step 6: Note what you learned

After the interaction, write one sentence about what the actual explanation turned out to be. Over time, keeping a mental (or literal) tally of how often your malice attributions turn out to be correct versus mistaken is calibrating. Most people discover that their suspicion has a poor track record — which makes it easier, next time, to reach for the razor before reaching for the defensive reply.

For you: The immediate benefit is emotional: assuming error instead of malice defuses anger before it drives a reaction you’ll have to walk back. But the deeper benefit is to your judgment and your reputation. If you respond to every ambiguous negative outcome as though it were a deliberate attack, you signal reactivity and poor judgment to everyone watching. The PM who stays curious under friction — who asks rather than accuses — is read as more senior, more trustworthy, and easier to work with. That perception opens doors that defensiveness closes.

For your team: A PM who models Hanlon’s Razor changes the emotional climate of a team. Misattribution is contagious — one person’s suspicion, expressed as a sharp comment or a pointed escalation, tends to provoke defensiveness that spreads. The inverse is also true: a PM who consistently gives the benefit of the doubt, who investigates before judging, creates space for others to do the same. Teams that assume competence rather than conspiracy spend less energy on manufactured conflict and more on the actual work.

For your organization: Beyond relationships, Hanlon’s Razor is a resource-allocation principle. Instead of spending time and energy planning for and defending against imagined hostile actors, teams that default to the simpler explanation divert that energy toward solving the real, more common problems — the miscommunications, the process gaps, the information asymmetries that actually cause most organizational friction. Charlie Munger has long advocated this kind of fact-based, emotion-resistant reasoning as essential to solving business problems clearly. Organizations that build a culture of charitable interpretation — while remaining genuinely alert to real bad faith — make better decisions and waste less energy on conflicts that never needed to exist.

Find the situation. Write the malice story, then write three innocent ones. Check the actual evidence.

Then respond to the most likely explanation — usually with a question, not a defense.

Doing this challenge? Share what you found on LinkedIn or X and tag it #MLAChallenge. The gap between how certain your suspicion felt and how thin the evidence actually was is usually the most worth naming out loud.

Leave a comment

Somewhere around the fourth AI tool in your stack, you stopped designing and started supervising, and nobody sent a memo about it. One week you were sketching flows. A few months later you were mostly reading things other systems made, deciding if they were good enough, fixing what wasn’t. Somebody in your all-hands called this 10x productivity. Somebody selling a $497 “AI Workflows for Designers” course called it something even better, with a bigger number attached, because bigger numbers sell more seats. What happened is you got pushed into a job nobody gave you a title for, and you didn’t exactly say no: the reviewer.

You can hand off the doing, but not the watching, and that second part was never the free half of the job - it just used to come bundled in, so nobody had to name it, let alone pay for it.

The delegation illusion

Boston Consulting Group asked 1,488 workers, in a study published this March, whether AI leaves them mentally fried, and 14% said yes - worst in marketing, at roughly one in four. Sure, BCG sells AI transformation for a living, so read the exact numbers the way you’d read a headline from someone with a horse in the race. But the shape of the finding doesn’t need BCG to be true: people whose job is mostly overseeing AI reported meaningfully more mental effort, more fatigue, more information overload than people whose job isn’t - the tool got faster, but the watching never caught up.

None of this is new, which is the part that should worry you. In 1983, an ergonomics researcher named Lisanne Bainbridge wrote a paper about chemical plants and other automated control rooms that’s still one of the most-cited things in her field, called “Ironies of Automation.” Her point: better automation makes the human’s leftover job harder, not easier, because you stop practicing the skill it replaced. Your judgment about it goes soft. And the one job left for you - deciding whether the machine is right - is the job you’re now worst prepared for.

There’s a study I trust more than the BCG one, because nobody’s selling anything in it. Researchers put sixteen drivers on ~150 kilometers of real road, part of it on autopilot, and measured their mental load two ways: what the drivers said they felt, and what an actual reaction-time test picked up. What the drivers said didn’t match what the test found: they said it felt easier, while their measured load went up, especially in traffic. They were wrong about their own effort, and wrong in the specific direction that matters - they underestimated the cost, not overestimated it. Vigilance has been known to be hard mental work since at least 2008. Watching just doesn’t feel like work, which is exactly why it gets skipped first.

Where it actually gets complicated

I’d be lying if I stopped there, though. Microsoft ran a survey of 319 knowledge workers across 936 real tasks last year, and the honest finding cuts the other way half the time: on routine, low-stakes work, people genuinely think less with AI in the loop, and that’s fine - that’s the delegation working as advertised. It’s only on the tasks where being wrong costs something that people reported putting in more effort with AI than without it.

That’s the annoying part - those are exactly the tasks you can’t afford to phone in, and the ones your calendar keeps scheduling back to back with the low-stakes ones, as if your brain can tell the difference on autopilot. It can’t, and that’s rather the point. Checking your third AI-generated onboarding flow before lunch is lighter than building it from scratch. Checking your ninth one, once you’ve stopped actually reading and started pattern-matching for “looks about right,” is a worse kind of tired, and it’s the kind that produces the bug you find in QA three weeks later.

Nobody puts any of this on a roadmap, because nobody’s measuring it. Your company tracks how many AI tools shipped this quarter. It doesn’t track the hour you spent unpicking a beautifully wrong flow at nine at night, or the review you rushed because you’d stopped believing your own tiredness counted as information.

I’ve watched this exact bet get made. A client’s leadership pushed teams onto early-stage AI tools - Make, whatever the pitch deck was selling that quarter - while cutting headcount at the same time, because the tools were supposed to cover the gap. The tools weren’t ready for anything production-grade. The designers who stayed ended up rebuilding most of that output by hand, after hours, because what shipped simply didn’t work. Upstairs called it a success. It had, after all, cut the budget.

Do the math yourself too, since no one up there will: BCG’s own data shows fried reviewers make 39% more major errors and report 33% more decision fatigue than reviewers who aren’t. A major error caught in QA costs an afternoon. The same error caught after ship costs a hotfix, a postmortem, and someone else’s afternoon too. Multiply either one by every reviewer on your team, every sprint, and you get a number that’s never once made it into a slide - not because it’s small, but because putting it there would mean admitting where it came from.

If anyone upstairs wanted the real cost of the rollout, they’d ask how many of those bugs trace back to a reviewer too fried to catch them. Nobody’s asking, because the honest answer would mean admitting the tool didn’t do what the slide deck said it would - and that’s a worse quarter than just letting you absorb it quietly.

I’ve done this too, more than I’d like to admit. Some days I’m four or five rounds into iterating something with AI, tweaking, re-prompting, switching between two or three tools that are all half-handling different pieces of the same problem, and somewhere around round six my standards quietly drop. I stop caring whether it’s good and start caring whether it’s done. Context-switching between tasks was already hard. Context-switching between the tasks and the tools supposedly doing them for me is a different order of tired, and most days that end like that, I close the laptop more drained than I ever felt doing the same work by hand - which was supposed to be the whole point of not doing it by hand.

So the honest version of this letter’s whole argument is narrower than “supervising AI always costs more than doing it yourself.” It costs more when the stakes are high, the output is unfamiliar, the tool is unreliable, or you’re four hours past when you should’ve stopped - and you’ll underestimate that cost, the same way those drivers did, because attention spent catching nothing feels like attention you never spent.

It’s also why your tool count matters more than your tool choice. Open Product Hunt on any given morning and you’ll find dozens of new AI tools launched that week alone, each one promising to be the thing that finally saves you time. Designers surveyed this year reported their kit more than doubling, from three tools to seven, in twelve months. (That number comes from a VC-backed report, and even their own methodology note admits the sample skews toward people who’ve gone all-in on AI.)

But it rhymes with what BCG found independently: productivity climbs through your second and third tool, then falls off a cliff. You didn’t decide seven was the number. It happened one plugin, one “just try this too” at a time, the same way scope creep happens to a project, except this scope creep happens behind your own eyes, and there’s no PM around to flag it.

What actually changes in practice

I’m not telling you to use less AI. If that’s the letter you were bracing for, close the tab - I have no interest in writing you a eulogy for doing everything by hand. What I’m telling you is narrower: budget the cost of watching the way you already budget the cost of doing, because right now you don’t, and the costs nobody budgets for are the ones that eventually eat the project.

Cap how many AI tools you’re actively supervising at once to two or three, roughly where BCG’s own curve turns over; past that point you’re not reviewing more, you’re just missing more. Batch your review passes into two or three fixed windows instead of checking as things trickle in. Research on notification batching, not AI-specific but close enough, found batched checking cut daily stress, while both constant checking and going fully dark made things worse.

Protect one real block of your day too, ninety minutes if you can get it, where you’re doing the work yourself instead of grading someone else’s attempt. Call it practice, not nostalgia. It keeps your judgment sharp enough to catch the machine when it’s confidently, fluently wrong, which, quietly, has become the actual job.

Be honest, too, about which of your reviews are the ones that don’t matter much and which are the ones where a miss costs something, and treat them differently on purpose - skim the first kind, and slow down for the second, on purpose, before you’re tired enough that slowing down stops being a decision and turns into something that just happens to you late in the evening.

You can delegate the sketch. You can delegate the first draft, the variants, the boilerplate. What you can’t delegate is the moment you decide whether any of it was right. That moment was always yours. All that’s changed is how easy it’s gotten to skip it - and call the skipping trust.

Leave a comment

A Scrum Master’s Perspective for Product Managers

Michał was coming back to the office on Wednesday after lunch when Katarzyna, a PM with six years of experience, caught up with him in the kitchen. She had three days of quarterly planning work behind her and a two-hour meeting with the board, where she had shown the new version of the roadmap. “I need to tell you what just happened. Because either I don’t understand this, or nobody in our company understands it.”

Katarzyna had prepared a roadmap for Q3. Six weeks of work — customer interviews,

Read the original on producttribe.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.