RSS Amplifier

Product Tribe 🔥 · Jun 9, 2026

Crafting Product Teams | The Exhaustion | The Selfish Gene Is Not the Whole Truth: On the Evolutionary Biology of Groups and Why Your Bonus System Kills What the Values Poster Promises

0
Sign in to vote or save

Destare Foundation, Alex Dziewulska, Sebastian Bukowski, Jakub Sirocki, Łukasz Domagała, Katarzyna Dahlke · Product Tribe 🔥

💜 Crafting Product Teams (by Alex Dziewulska)

💜 The Exhaustion (by Alex Dziewulska)

💜 The Selfish Gene Is Not the Whole Truth: On the Evolutionary Biology of Groups and Why Your Bonus System Kills What the Values Poster Promises (by Łukasz Domagała)

💪 Interesting opportunities to work in product management

🍪 Product Bites - small portions of product knowledge

🔥 MLA week#51

Join Premium to get access to all content.

It will take you almost an hour to read this issue. Lots of content (or meat)! (For vegans - lots of tofu!).

Grab a notebook 📰 and your favorite beverage 🍵☕.

DeStaRe Foundation

Every product lives inside a tension it can’t resolve. Business wants revenue. Users want their problem solved. Technology is either a constraint or an opportunity — sometimes both in the same sprint.

Three perspectives. Three legitimate priorities. They will never fully align.

And that’s the point.

Most organizations treat this tension as a disease. Something to cure. Get everyone on the same page. Align. Harmonize. Build cross-functional empathy until the friction disappears and we’re all one big product-shaped hug.

Wrong direction. The friction is doing something. Remove it and you don’t get alignment — you get groupthink in a nicer outfit.

When a product team stops arguing, one of two things has happened. Either they’ve genuinely arrived at the same conclusion through independent analysis of the same evidence — which is rare enough to be essentially theoretical. Or someone has stopped talking.

Usually it’s the second one.

The engineer who sees the technical debt piling up but has learned that raising it gets filed under “not a team player.” The researcher whose findings contradicted the roadmap and got thanked politely and ignored. The business person who flagged that the margins don’t work but got outvoted by enthusiasm.

Each of these people held information the group needed. Each chose silence over friction. And the product got worse — not because of what was said, but because of what wasn’t.

This isn’t a thought experiment. It’s the default state of most product teams I’ve worked with. The tension between user, business, and technology doesn’t disappear when you stop having the argument. It goes underground. And underground tensions don’t resolve — they metastasize.

Here’s what makes this counterintuitive. Research on group decision-making consistently shows that teams perform better when members disagree before the discussion — not after, not during, before. When people walk into a room already holding different positions, the group surfaces more information, challenges more assumptions, and arrives at better decisions.

The finding that should genuinely unsettle anyone who optimizes for harmony: this effect holds even when none of the dissenting members have the correct answer. Nobody in the room is right. But the disagreement itself breaks the confirmation loop that homogeneous groups fall into reflexively. It forces people to say things they wouldn’t say in a room full of nodding heads. The mechanism isn’t “someone had the right answer and was heard.” The mechanism is “disagreement itself changed what information entered the room.”

That’s the product triangle working as designed. The business person doesn’t need to be right about the margin concern. The engineer doesn’t need to be right about the technical risk. The researcher doesn’t need to be right about the user signal. What matters is that each of these perspectives forces the others to account for something they would have otherwise ignored. The tension between the vertices isn’t noise. It’s signal processing.

Now — before someone takes this as an argument for turning every standup into a cage match.

The research draws a line that matters. Task conflict — disagreement about the work, the approach, the direction — can improve decisions. But the moment it leaks into relationship conflict — when “I disagree with your approach” becomes “I don’t respect your judgment” — everything collapses. Performance drops. Trust erodes. The conditions that made productive disagreement possible get destroyed by the very friction that was supposed to help.

And in practice, the two bleed into each other constantly. Especially in product, where people’s identities are wrapped around their professional positions. The PM who championed the feature. The engineer who designed the architecture. The researcher who ran the study. Disagreeing with the work feels like disagreeing with the person. That’s the edge most teams fall off.

This is why most organizations default to removing the tension entirely. Not because they’re stupid. Because they’ve watched task conflict curdle into hostility, and they concluded — reasonably — that harmony is safer.

Reasonable. Also wrong.

The answer isn’t eliminating tension. It’s building the architecture that holds it. Clear decision rights. Explicit norms for how disagreement gets processed. And a shared understanding that the purpose of the tension isn’t to determine who’s right — it’s to surface what each perspective sees that the others can’t.

Here’s the distinction most teams collapse.

Agreement is everyone thinking the same thing. It feels great. It’s usually a sign that someone stopped talking or the group filtered for people who already think alike. Agreement is the output of a system that has stopped processing new information.

Compromise is everyone giving something up. The business accepts lower margins. Engineering accepts more debt. The user gets a diluted version. Everyone loses a little. Nobody loses a lot. And the product is mediocre — because mediocrity is what you get when every strong perspective gets filed down until it stops making anyone uncomfortable.

Consensus — real consensus — is a decision everyone can commit to, built from perspectives that genuinely competed. Not because everyone agrees it’s the best path. Because the disagreement was heard, the trade-offs were surfaced, and the decision is stronger for having been tested against each perspective’s resistance. You might still think you’re right and the decision is wrong. You commit anyway — not because you were overruled, but because the process earned it.

“Disagree and commit” isn’t losing gracefully. It’s the operating system of every product decision that actually holds up in the real world.

There’s a parallel here that I keep circling back to.

In product, we’ve accepted that certainty is impossible. You never have enough data. The user research captured a sliver. The market analysis has blind spots. The technical feasibility assessment has assumptions you won’t discover until you’re three sprints deep. We’ve built entire methodologies around this — lean, build-measure-learn, hypothesis-driven development. Nobody waits for certainty before acting. We call that professional maturity.

We haven’t extended the same thinking to disagreement.

When a team faces uncertainty about the market, we call it normal. When a team faces disagreement about the direction, we call it dysfunction. But the structure is identical. Both are states of incomplete alignment — one between what we know and what we need to know, the other between what different perspectives believe about the right path. Both are uncomfortable. Both are permanent features of the work, not temporary failures of the team. And both are where the actual craft lives.

The craft of product management isn’t waiting for certainty. It’s making good decisions while uncertainty is still in the room.

Same applies to tension. The craft isn’t resolving the disagreement between business, users, and technology. It’s making good decisions while the disagreement is still in the room. Not despite the tension — through it.

The teams that scare me aren’t the ones arguing about the roadmap.

They’re the ones where every planning session ends in twenty minutes because everyone already agrees. Where the PM presents a direction and nobody pushes back — not because it’s obviously right, but because pushing back costs more than staying quiet. Where the culture rewards harmony and calls friction “not constructive.”

Those teams feel great. They move fast. They build the wrong thing with confidence.

A product team without tension is a team where one perspective won and the others surrendered. The information those silenced perspectives carried — the constraints, the risks, the alternative paths — never entered the decision. The group solved the hidden profile wrong and felt aligned doing it.

Product management doesn’t live in agreement. It lives in the sustained, uncomfortable, productive space between people who see different things and refuse to stop saying so.

Share

Web Summer Camp 2025 — and Alex will be there

Web Summer Camp is a three-day, hands-on event for Europe’s web professionals — workshops, small group sessions, and real conversations, taking place July 3–5 in Opatija, Croatia. Not the kind of conference where you collect slides and forget about it by Monday. The kind where you actually work through problems with people who build digital products for a living.

This year they added an AI track — not because everyone else is doing it, but because they wanted to do it right. There’s also a dedicated Founders program and tracks for JavaScript, PHP, UX, and product.

Alex will be joining as a speaker. If you’re heading to Opatija this summer — or considering it — this is where to find her in person.

More about the conference: [LINK]

Ready to grab a ticket? [TICKETS]

Do you need support with recruitment, career change, or building your career? Schedule a free coffee chat to talk things over :)

!!! HOT OFFER !!!

Senior Product Manager

Product Manager

  1. Product Manager - AMS

  2. Product Manager - Worldline

  3. Product Manager - LHH

  4. Product Manager - InPost

  5. Senior Product Manager - Zendesk

Refer a friend

The user research session goes smoothly. Users are thoughtful, articulate, and engaged. They walk through the onboarding flow, flag a few minor confusions, and rate the experience positively. The team takes notes, updates the design, and ships with confidence.

Three months later, support tickets tell a different story. Users are getting lost at the exact steps that tested fine. Dropoff in the activation funnel is concentrated in moments that felt obvious in the research setting. The gap between what was tested and what is happening in production is significant — and the team has no clear explanation for it.

One part of the explanation is the hot-cold empathy gap. The users in the research session were calm, focused, and undistracted. The users in production are distracted, rushed, and often confused before they even start. These are not the same user in the same state. They are the same person in profoundly different psychological conditions — and the product was designed for one of them.

A fintech team is designing the first-time funding flow for a new investment account. The flow is tested extensively in moderated sessions. Participants are recruited, scheduled, and seated in a comfortable room or a quiet home environment. They walk through the flow deliberately, ask clarifying questions, and complete it successfully.

The product launches. Abandonment in the funding flow is 60 percent higher than the testing predicted. Exit surveys reveal the primary reasons: users were unsure whether the transfer was reversible, didn’t understand the timeline, and felt anxious about committing funds to an unfamiliar platform. Several noted they started the flow during a lunch break or while managing other tasks.

None of these issues appeared in testing. The participants in testing were calm. They could afford to read carefully. They were not anxious about the commitment because they were in a session context that removed the stakes from the decision. The real users encountered the same flow in a hot state — with real money, real uncertainty, and real time pressure — and experienced it as something fundamentally different.

The hot-cold empathy gap is a cognitive bias, formally identified by behavioral economist George Loewenstein in his 1996 paper “Out of Control: Visceral Influences on Behavior,” describing the systematic failure of people in one psychological state to accurately predict or understand their own behavior and preferences in a different psychological state.

Loewenstein’s framework distinguishes between “hot” states — conditions involving heightened visceral arousal such as stress, anxiety, hunger, time pressure, excitement, or fear — and “cold” states — conditions of relative calm, reflection, and low arousal. His central finding was that people in cold states dramatically underestimate how much hot states influence behavior, and people in hot states dramatically underestimate how different their cold-state preferences will be. The gap operates in both directions, but the cold-to-hot direction is most consequential for product design: we design in cold states, for users we imagine will also be in cold states, when the actual moment of use is frequently hot.

Research following Loewenstein’s foundational work has documented hot-cold empathy gaps across a range of visceral states — hunger, sexual arousal, physical pain, social anxiety, financial anxiety, time pressure — and consistently found that the magnitude of the gap is substantially larger than the cold-state subject predicts. The gap is not a small miscalibration. It is a systematic failure of cold-state cognition to model hot-state experience.

Product design happens almost entirely in cold states. Designers and PM-owie work in comfortable offices or home setups, with adequate time, full attention, and no stakes attached to the design decisions they are evaluating. User research, even when well-executed, typically replicates the design environment more than the use environment. Participants are scheduled, seated, and briefed. The session context removes the ambient anxiety, distraction, and time pressure that characterize real use.

The use environment is frequently a hot state. A user setting up a new financial account is anxious about the commitment. A user onboarding a new tool during a busy workday is distracted and time-constrained. A user completing a health-related form is emotionally activated by the context of the decision. A user troubleshooting an error in production is under pressure. The cold-to-hot empathy gap means that the product team’s model of what this user needs — formed in calm reflection — is systematically miscalibrated relative to what the user actually experiences in the hot moment.

Loewenstein classified the hot-cold empathy gap along three dimensions: direction, time, and person. Direction refers to cold-to-hot versus hot-to-cold gaps, both of which are real and both of which produce predictable errors. Time refers to whether the gap involves predicting one’s own future states or recalling past states — both are inaccurate in the direction of underestimating visceral influence. Person refers to the interpersonal dimension: people in cold states also underestimate the influence of hot states on other people’s behavior, which is the dimension most directly relevant to product design. The designer in a cold state cannot accurately model the user in a hot state, even with the best intentions and extensive user research.

A specific and underappreciated manifestation of the gap is financial and commitment anxiety. Users making first-time purchases, first-time account connections, or first-time data-sharing decisions are in a hot state characterized by uncertainty and risk aversion that is qualitatively different from the cold-state evaluation of those same decisions. Research on financial decision-making under emotional arousal consistently shows that hot-state financial decisions exhibit higher risk aversion, greater sensitivity to ambiguous terms, and lower tolerance for complexity than cold-state evaluations of the same choices.

Product teams that design payment flows, account connection flows, and commitment steps in cold-state test conditions will consistently underestimate how much anxiety reduction, explicit reversibility information, and simplification these flows require. The flow that felt clear in testing feels opaque and frightening in production — not because the design changed, but because the user’s state did.

Headspace’s approach to emotional-state calibration in its onboarding flow demonstrates deliberate attention to the hot-cold gap. The app is used primarily by people in moments of stress, anxiety, or emotional difficulty — hot states that are the primary motivation for opening the product. Headspace’s onboarding was designed to function for users in these states: short sessions that can be completed in under five minutes, minimal decision burden in the initial setup, and an explicit framing that the first session requires no prior knowledge or commitment. The design reflects an understanding that the user opening the app for the first time during an anxious moment is in a qualitatively different state from the user evaluating meditation apps in a calm research session. Headspace’s activation rate, which has been discussed publicly as substantially above industry benchmarks for wellness apps, correlates with this state-aware design approach.

Stripe’s error handling and API documentation reflects cold-to-hot empathy gap awareness at a developer experience level. Developers integrating payments in production are frequently in hot states: they are working under deadline, debugging in a live environment, or handling a critical customer-impacting issue. Stripe’s error messages are designed to be actionable in these conditions — they include specific error codes, plain-language descriptions of the problem, and links to relevant documentation that directly address the issue rather than requiring the developer to navigate from a general documentation home. The design assumption is not that the developer will have time for careful reading. It is that the developer is stressed and needs the minimum information required to act immediately. This assumption reflects an accurate model of the hot state rather than the cold-state reading experience that most documentation is designed for.

Robinhood’s design of first-time investment flows became one of the most studied cases of hot-cold empathy gap failure in consumer fintech. The app’s streamlined, gamified interface was designed to feel exciting and accessible — an accurate reflection of the cold-state experience of evaluating investment apps. In production, first-time investors activated in hot states of excitement and FOMO made trades they subsequently reported regretting, including options trades they did not understand. The cold-state design of the flow — which felt clear and empowering in calm evaluation — did not adequately account for the hot states in which real investment decisions would be made: excitement about market movements, anxiety about missing opportunities, and the cognitive narrowing that comes with financial arousal. Subsequent regulatory scrutiny and product redesign focused substantially on introducing friction and information that the cold-state design had optimized away.

The hot-cold empathy gap matters for product teams because it describes a systematic blind spot that standard research and design processes do not automatically correct for. Usability testing in calm settings produces accurate models of user behavior in calm settings. It does not produce accurate models of user behavior in the hot states that characterize many of the most important product moments: first use, first payment, first commitment, troubleshooting, and high-stakes decisions.

The gap also explains a recurring pattern in product post-mortems: the feature that tested well but performed poorly in production. When the test environment does not replicate the emotional conditions of the use environment, the test is measuring a different experience than the one being shipped. The gap between test performance and production performance is not random error. It is the predictable consequence of designing in cold states for hot-state users.

Research on mobile onboarding behavior specifically has documented that the majority of mobile app first-use sessions occur in environments with high ambient distraction — commutes, brief breaks, concurrent tasks — and that the cognitive load available for onboarding tasks in these conditions is substantially lower than the cognitive load available in controlled test settings. Product teams that calibrate onboarding complexity to test-session performance will routinely ship experiences that are too complex for the conditions in which they will actually be used.

Explicitly identify the emotional state of the user at each key product moment. For every critical flow — onboarding, first purchase, account connection, error recovery, cancellation — characterize the likely emotional state of the user in that moment. Is the user likely to be anxious, rushed, excited, confused, or frustrated? This characterization should precede design decisions about information density, decision complexity, and required reading. Cold-state design calibrated to hot-state use requires knowing which moments are hot.

Design redundant legibility into hot-state moments. In moments where users are likely to be in hot states, reduce the cognitive work required to understand the next step. Shorter sentences, more explicit CTAs, visible reversibility information, and reduced optional complexity are not design compromises — they are calibrations to the user’s actual cognitive capacity in the relevant state. Information that is easy to process in a calm test session may be effectively invisible in a stressed or distracted production session.

Test in conditions that approximate the use environment. When possible, conduct research sessions that introduce realistic environmental conditions: time pressure for flows that users will complete under deadline, mobile-only sessions for flows that will primarily be completed on mobile in variable conditions, and ambient distraction for flows that users complete while managing other tasks. The closer the research environment is to the use environment, the smaller the gap between test performance and production behavior.

Use behavioral data to identify hot-state failure patterns. When a flow that tested well underperforms in production, analyze where users exit relative to the sequence of steps. Exit patterns concentrated in specific steps — particularly steps involving commitment, payment, or irreversible action — are often signatures of hot-state friction that was not visible in cold-state testing. These patterns indicate where the design needs to account for anxiety, time pressure, or reduced cognitive capacity rather than the calm deliberation that testing assumed.

Introduce explicit state-awareness into design review. During design reviews, include an explicit question: “In what emotional state will the user encounter this?” If the answer involves anxiety, stress, time pressure, or high stakes, the design should be evaluated against that state rather than against the calm-evaluation state of the design session itself. The question does not require radical design changes, but it surfaces calibration decisions that would otherwise be made invisibly.

George Loewenstein’s foundational insight was that visceral states are not merely emotional coloring on top of rational decision-making. They are a fundamental driver of preferences, attention, and behavior — one that cold-state reasoning cannot accurately simulate. We know this about ourselves. We know that we eat differently when hungry, decide differently when anxious, and communicate differently when stressed. What the hot-cold empathy gap describes is the systematic failure to apply this knowledge when modeling other people’s behavior in states we are not currently experiencing.

For product teams, this is not an abstract epistemological problem. It is a design problem with a specific and predictable shape: the designs we build in calm offices, evaluated in calm research sessions, will consistently underestimate what users in hot states need. They will include more complexity than hot-state users can process, less anxiety reduction than hot-state users require, and fewer explicit reassurances than hot-state users need before committing.

The corrective is not to redesign everything for worst-case emotional states. It is to build an accurate model of which moments in the product are hot and to calibrate the design of those moments to the user who will actually be there — not the user we imagined while building.

The user in the research session is one of the people using the product. They are rarely the most important one.

Leave a comment

In October 2004, Chris Anderson, then editor-in-chief of Wired magazine, published an article that reframed how the digital economy works. His observation was precise: the internet had not simply made existing markets more efficient. It had made an entirely different type of market possible — one in which an almost unlimited number of niche products, each selling to a small audience, could collectively generate more revenue than a small number of blockbusters selling to a mass market.

The article was titled “The Long Tail.” The insight it contained has since been applied to media, e-commerce, software, and entertainment. It has also been incompletely applied to product management — where its implications reach further than most teams have drawn them.

A music streaming service in 2004 had a catalogue problem. The physical constraints of radio and retail created a brutal filter: if a song wasn’t likely to be a hit, it wasn’t worth the distribution cost. Of the millions of songs ever recorded, physical distribution economics meant that only a few thousand were reliably accessible at any given time. The result was a market dominated by a small number of popular songs — the “head” of the distribution curve — with an enormous body of existing music effectively inaccessible to most listeners.

Spotify launched in 2006 with a catalogue of millions of tracks. Within years, the service had documented what Anderson’s theory predicted: a substantial portion of its streams came from songs outside the top ten thousand most played tracks. The demand for niche music was not absent before. It was suppressed by the cost of access. Remove the constraint, and the tail came to life.

This is the core insight of the Long Tail. And it applies well beyond music catalogues.

The Long Tail theory, developed by Chris Anderson in his 2004 Wired article and expanded in his 2006 book of the same name, describes a statistical distribution pattern in digital markets: a small number of products (”the head”) generate high individual sales volume, while a vast number of products (”the tail”) each sell in small quantities. The key insight is that in digital markets — where marginal distribution costs approach zero — the aggregate revenue from the tail can equal or exceed the revenue from the head.

The theory rests on a power law distribution, first described mathematically in the context of market concentration by Vilfredo Pareto, and familiar in many natural and social phenomena. The Long Tail concept’s contribution is not the mathematical observation but the economic implication: when distribution costs are near-zero and discovery infrastructure (search, recommendation, filtering) is sufficient, the economically viable market is not limited to products with mass-market appeal. The product that serves 50,000 people with a highly specific need is viable in a way it was not when serving those 50,000 people required physical shelf space, broadcast distribution, or retail presence.

For product managers, the implications extend beyond catalogue economics. The Long Tail describes a new relationship between product specificity and market viability — one that changes how products are positioned, how features are prioritized, and how product organizations think about the value of serving minority use cases.

Anderson identified three enabling conditions for Long Tail economics. The first is democratization of production: the cost of creating content, software, or products has fallen dramatically, enabling a far larger number of producers to enter markets than was previously viable. The second is democratization of distribution: digital channels eliminate the shelf space constraint that historically limited what products could reach consumers at all. The third is connection of supply and demand: search engines, recommendation algorithms, user reviews, and social sharing make it possible for the small number of people who want any specific niche product to find it.

All three forces operate simultaneously in digital product markets. The product team that understands them can identify where their product sits in the distribution, how the tail is shifting in their category, and whether they are serving the head at the expense of economically viable tail segments.

Within a single product, the Long Tail dynamics apply to feature usage. In most products, a small number of features — the head — are used by the majority of users in the majority of sessions. A large number of features — the tail — are each used by small minority segments, but in aggregate may represent a substantial portion of the value the product creates.

The implication for feature prioritization is counter-intuitive: the features with the highest per-feature usage may not be the features creating the most aggregate value. A power user segment that relies heavily on a set of niche features may contribute disproportionate revenue, retention, and word-of-mouth advocacy precisely because those features are unavailable in mainstream alternatives. Removing or neglecting tail features to simplify the product can eliminate the specific value that makes the product worth switching to for the segments that use them.

Anderson’s analysis identified discovery — the mechanism by which users find niche products or features — as the primary constraint on Long Tail economics. A product that exists in the tail but cannot be found generates no value. This is why Amazon’s recommendation engine, Spotify’s Discover Weekly, and YouTube’s algorithmic home page are not incidental features — they are the infrastructure that makes Long Tail economics viable. Without discovery, the tail is economically equivalent to non-existence.

For product teams, this means that investing in the tail without investing in discovery infrastructure is incomplete. The niche feature that serves a valuable minority segment delivers its value only if that segment can identify and access the feature. Discoverability of tail features within a product is the equivalent of search and recommendation in a content marketplace — it is the mechanism that converts potential tail value into actual user value.

The Long Tail provides a strategic alternative to the default product logic of building for the median user. In markets where the head is dominated by an established player optimizing for the majority, the viable entry strategy is often to own a niche that the head player cannot efficiently serve — and to do so at a depth that mass-market players are structurally unable to match. The niche player who serves 50,000 users with exceptional depth outcompetes the mass-market player who serves those same users adequately, because switching to the niche product resolves a need the mass-market player cannot meet without compromising its core position.

This is the strategic logic underlying a recurring pattern in technology market evolution: specialized tools that serve a niche with exceptional depth displace general-purpose tools that serve the same niche adequately, then expand from the niche to adjacent segments once their depth creates a defensible position.

Amazon’s early growth is the foundational business case for Long Tail economics. Barnes and Noble’s largest physical stores stocked approximately 130,000 titles — the economic limit of physical retail shelf space. Amazon’s digital catalogue in 2004 contained millions of titles, and Anderson documented that more than half of Amazon’s book revenue came from titles outside Barnes and Noble’s top 130,000. The demand for books outside the head of the market had always existed. Physical distribution suppressed it. Amazon’s digital catalogue liberated it, and the aggregate value of the tail proved to be substantial relative to the head.

Etsy’s market entry and growth illustrate Long Tail dynamics in physical goods. Individual artisans producing handmade goods at low volumes had no viable retail market before digital platforms — their production volumes were too small to justify physical retail relationships, and their customer bases were too geographically distributed to support local sales. Etsy created the discovery infrastructure — search, category browsing, and recommendation — that connected makers with buyers who wanted exactly their specific products but had no mechanism for finding them. The market for handmade goods at Etsy’s scale was not created by Etsy; it was revealed by the infrastructure that made the tail navigable.

Figma’s growth in the design tools market illustrates Long Tail niche strategy applied to product competition. Adobe’s design tools — Photoshop, Illustrator — are head products built for the mass market of design use cases. They serve the majority of users adequately across a wide range of tasks. Figma entered through the specific niche of collaborative, web-based UI design, where it could offer depth the Adobe suite could not match without abandoning its position in other segments. The niche was real and underserved; the depth Figma offered within it was substantially better than what the mass-market alternative provided. Figma grew from the niche outward as its depth and ecosystem created switching costs that extended beyond the original niche use case.

Spotify’s Discover Weekly feature provides an internal product example of Long Tail infrastructure. Launched in 2015, Discover Weekly is a personalized weekly playlist that exposes users to artists and tracks outside the head of the catalogue — the obscure, the niche, the artists who would never reach users through radio or editorial curation. The feature has been credited with driving meaningful streams to long-tail catalogue content, increasing artist visibility, and creating a product experience that users describe as genuinely surprising in ways that head-content curation cannot match. Spotify has reported that Discover Weekly significantly increased streaming of music outside the top charts, demonstrating that the demand for tail content was constrained by discovery, not by preference.

The Long Tail matters for product teams because it challenges the dominant logic of building for the median user. The median user is real, and building for them is necessary. But in markets where digital economics apply — which is most markets that product teams work in — building exclusively for the median user leaves substantial value uncaptured in the tail and creates a vulnerability to niche competitors who serve specific segments with greater depth.

The feature prioritization implication is particularly important. Teams that evaluate features by average usage volume will systematically underinvest in tail features — the features used by small but valuable segments whose aggregate contribution to retention, revenue, and advocacy is underrepresented in average-usage metrics. The power user who relies on three niche features is not well-captured in a feature usage report that shows those features at two percent adoption. Their value to the business — in terms of willingness to pay, churn risk if features are removed, and word-of-mouth influence in their professional community — is substantially higher than the usage metric suggests.

There is also a competitive dynamics implication. In mature markets, the head is typically well-served by established players. The tail — the segments with specific needs that mass-market products address inadequately — is often the only viable entry point for a new product that does not have the resources to compete on every dimension simultaneously. The Long Tail provides a strategic framing for why niche entry is often not a limitation but a deliberate and sound approach to market entry.

Audit the product’s tail users explicitly. Beyond average feature usage, identify which user segments rely on features with below-average average usage and evaluate their contribution to retention, revenue, and advocacy relative to their size. The features that serve 5 percent of users but retain 20 percent of revenue deserve a different prioritization conversation than average-usage metrics surface.

Build discovery infrastructure for tail features within the product. Features that exist but cannot be found by the users who would value them are economically equivalent to features that don’t exist. For any feature serving a minority segment, invest in the discoverability mechanism — contextual suggestions, onboarding paths for specific user types, in-product search, and use-case-based documentation — that connects the minority with the feature designed for them.

Evaluate competitive positioning for tail opportunities. In any product category, identify the segments that the dominant players serve adequately but not exceptionally. These are tail opportunities where depth is achievable and defensible. A product that serves a niche segment with the depth that mass-market players cannot match creates a position that is hard to dislodge once switching costs accumulate.

Use recommendation and personalization to shift user behavior toward the tail. Anderson’s analysis identified recommendation as the mechanism that converts latent tail demand into actual usage. For products with deep feature sets or content catalogues, personalized recommendations — “users like you also use” or “based on your workflow, this feature” — replicate the Long Tail discovery infrastructure that makes the tail economically viable in content markets.

Anderson’s core insight was deceptively simple: the internet does not just make things cheaper. It makes things possible that were not possible before — specifically, the viability of an almost unlimited number of niche markets that physical distribution economics had made unviable. The Long Tail is not a story about the death of hits. Hits remain important. It is a story about the revealed scale of what was always below the surface: the aggregate demand for specificity that mass markets had never been able to serve.

For product teams, the translation is equally simple and equally easy to underestimate: the median user is not the only user who matters, the median use case is not the only use case worth serving, and the product that serves a niche with exceptional depth is not less valuable than the product that serves everyone adequately. In many markets, it is the only viable product strategy — because the head is already occupied, the median user is already served, and the open territory is in the long tail of unmet specific needs.

The future of business, Anderson wrote in 2004, is selling less of more. The future of product, by the same logic, is building deeper for more specific users — and trusting that the aggregate of many niches served well is a more defensible position than one median user served adequately.

The hits will take care of themselves. Build for the tail.

Leave a comment

The first week goes well. The new PM has a laptop, calendar access, and a full schedule of introductory meetings. By the end of week two, they’ve met the engineering lead, the design team, and most of the key stakeholders. By the end of the month, they’ve sat in on sprint reviews, read the product documentation, and started forming opinions about the roadmap.

By month three, they’ve shipped their first spec and made their first significant recommendations. The organization calls the onboarding a success.

And then, six months later, the PM makes a decision that creates a significant problem. It turns out they misunderstood the most important constraint on the product. They had the right information available. Nobody helped them see which information was most important. The onboarding covered everything. It prioritized nothing.

This is the most common failure mode in PM onboarding. Not a lack of information. A lack of a structured model for what to learn, in what order, and why it matters before making consequential decisions.

A senior PM joins a Series B company to lead the growth product area. She comes from a company where she built significant expertise. She is experienced, smart, and motivated. The onboarding consists of a week of meetings, a set of onboarding documents, and a Confluence space with product history. She reads everything and schedules conversations with the key players.

By month two, she is frustrated. The context she needed wasn’t in the documents. The decisions that matter aren’t written down anywhere. The constraint she’s now working around — a technical decision made eighteen months ago with implications for everything her team can ship — was mentioned in passing in one meeting but never explained. She built a strategy that ignored it because she didn’t know it existed.

Her manager sees a PM who isn’t moving fast enough. She sees a company that didn’t teach her what she needed to know. Both are partially right. The actual problem is structural: the onboarding was designed to transmit information, not to build the contextual model that makes information meaningful.

PM onboarding fails in a specific and predictable way. It provides coverage without depth, access without context, and information without the map that tells you which information is load-bearing.

The standard onboarding script is well-intentioned: read the product documentation, meet the team, understand the roadmap, learn the metrics, attend the rituals. These are all necessary. None of them is sufficient. They transfer the surface of the product without transferring the organizational archaeology — the understanding of why the product is the way it is, which decisions are revisable and which are fixed, where the real leverage points are, and who in the organization actually shapes what gets built.

A new PM who has read all the documentation and attended all the rituals still does not know which features were built because of a customer commitment that no longer exists, which technical constraints are structural and which are legacy, which stakeholders have disproportionate influence on what gets prioritized, and which metrics are genuinely diagnostic versus politically important. These are the things that make the difference between a PM who makes good decisions in their first year and one who makes expensive mistakes while learning the terrain.

The best PM onboarding programs are not more comprehensive. They are more structured about what must be understood before the new PM is expected to make significant decisions — and honest about the time it takes to build that understanding.

The most important discipline of the first thirty days is the discipline of not proposing. A new PM with strong prior experience has immediate opinions. The opinions feel grounded — and they are, in the context the PM comes from. In the new context, they are pattern-matched against an incomplete model. The new PM who proposes in week two is proposing from a map that is missing most of its territory.

The first thirty days should be structured around listening to understand, not listening to respond. The goal is not to collect opinions about the product and the team. It is to build a model of why the product is the way it is — the historical decisions, the constraints, the organizational dynamics, and the user realities that produced the current state. This requires a different type of conversation than the introductory meeting circuit typically produces.

Specifically, the first-thirty-days listening agenda should include: the history of the three most significant product decisions in the past two years and the reasoning behind them; the constraints that are truly fixed versus those that are negotiable; the gap between what users say they need and what the team understands them to actually need; and the organizational dynamics that shape which initiatives get resources and which don’t. These are questions that require not meetings but conversations — extended, exploratory discussions with people who have been there long enough to have the context.

The listening tour, when done well, produces a picture of the product’s recent past that no documentation captures: the decisions that were made under pressure, the constraints that were accepted as permanent, the opportunities that were evaluated and declined. This is the archaeology that a new PM needs before they can propose anything that accounts for the real terrain.

By day thirty, the new PM has a partial model of the product’s history and context. The second thirty days are for stress-testing and completing that model through direct experience rather than through conversation.

This means spending time with users — not in structured research sessions, but in the environments where users actually use the product. It means reviewing the support ticket history to understand where users experience friction that the team has normalized. It means reading customer success call notes from the past six months. It means understanding the full technical architecture at a level sufficient to know which product bets are feasible in six months and which require eighteen.

It also means understanding the metric system at a level that goes beyond knowing the definitions. Which metrics are genuinely diagnostic of user value? Which metrics are tracked because they were historically important and have become organizational fixtures? Which metrics are moving in the wrong direction but are not being discussed? A new PM who understands the metric system at this depth can make proposals that are grounded in what the product actually needs rather than what the current reporting cadence surfaces.

By day sixty, the PM should have a working model of the product that is detailed enough to support proposals with consequences — that is, proposals that involve trade-offs, constrain other possibilities, and commit organizational resources to a direction. Before this model is complete, significant proposals are premature.

The third thirty days is when the new PM makes their first consequential contribution. This is different from shipping something. A new PM in many organizations can ship something in the first thirty days without having a good model of what they’re doing. The first consequential contribution is the first thing the PM does that they could only have done with the context acquired in the prior sixty days.

It might be a prioritization decision that resolves a tension the team had been deferring. It might be a user research initiative that surfaces a problem the team had normalized. It might be a stakeholder alignment effort that creates organizational clarity where there was ambiguity. Whatever its form, it should be traceable to the context the PM built — something that demonstrates that the PM understands not just what the product does but why it matters and what it needs.

The ninety-day mark is also the right moment for the new PM to document their understanding of the product’s strategic context and share it with their manager and key stakeholders. This document — sometimes called a “listening tour synthesis” or a “product context brief” — serves two functions. It creates a legible record of what the PM has learned, which surfaces gaps and corrections. And it demonstrates to the organization that the PM has invested in understanding before acting — a signal of judgment that is as important as the quality of the eventual proposals.

The best onboarding programs we have observed share three structural properties that distinguish them from standard onboarding.

First, they sequence the PM’s access to different types of information deliberately. Documentation and product history come first. Direct user engagement comes second. Strategic and organizational context comes third. This sequence builds the model from the bottom up — from the product as it is, to the users it serves, to the organizational dynamics that shape its future. Reversing the sequence — starting with strategy before understanding the product — produces a PM who has opinions about strategy without the context to evaluate them.

Second, they assign the new PM a structured learning agenda that is distinct from their operational agenda. The operational agenda — attending the rituals, participating in sprints, meeting the stakeholders — is necessary but not sufficient. The learning agenda is a set of specific questions the PM should be able to answer by specific dates, covering the historical decisions, technical constraints, user realities, and organizational dynamics that constitute the load-bearing context. The learning agenda makes the invisible work of context-building explicit and legible.

Third, they protect the new PM from consequential decisions for the first thirty days. This is the most culturally difficult element for organizations that reward initiative and urgency. A new PM who is given the space to listen and learn before being asked to propose will make substantially better proposals than one who is pushed to demonstrate value immediately. The thirty-day protection period is not an accommodation for the PM’s comfort. It is an investment in the quality of the decisions that follow.

The product organizations that take PM onboarding most seriously — Amazon, Stripe, and a number of companies that have invested in explicit PM development programs — share a common orientation: they understand that the cost of a bad decision by an under-contextualized PM substantially exceeds the cost of a structured, slower onboarding. The urgency to ship is real. The urgency to get the ship right is more important.

PM onboarding matters because the decisions PMs make in their first six months set trajectories that are expensive to reverse. A feature shipped from an incomplete model may commit engineering resources to an approach that will require rework. A prioritization decision made without understanding the organizational dynamics may create stakeholder conflicts that take quarters to resolve. A strategic proposal made before the constraints are fully understood may move an organization in a direction that is costly to correct.

The cost of bad early decisions is not just the decision itself. It is the accumulated credibility cost of proposals that miss the context, the engineering cost of work done based on misunderstood constraints, and the organizational cost of alignments that had to be undone and redone. A PM who takes ninety days to build context before making consequential proposals will, in almost every case, produce better outcomes in their first year than a PM who begins proposing in month one from an incomplete model.

There is also a retention dimension. PMs who feel they are operating without adequate context are more likely to experience early-stage frustration that leads to attrition. The organizations that invest in structured onboarding retain new PMs at higher rates — not because the onboarding is comfortable, but because it produces the context that makes early performance possible, and early performance is the primary driver of early-stage retention.

Create a structured learning agenda, separate from the operational agenda. For every new PM, define the specific questions they should be able to answer by the thirty-, sixty-, and ninety-day marks. The questions should cover product history, technical constraints, user realities, metric systems, and organizational dynamics. Make the agenda explicit and review it in the regular one-on-one cadence. The learning agenda transforms invisible context-building work into legible organizational progress.

Assign an onboarding guide who is not the PM’s direct manager. The most useful onboarding guide is a peer with sufficient tenure to have the organizational archaeology and sufficient proximity to the new PM’s level to communicate candidly about the things that documentation doesn’t capture. The manager relationship is too formal and too evaluative to support the kind of candid context-building conversations that a peer guide enables.

Protect the first thirty days from consequential decision-making. Define explicitly, at the start of onboarding, which decisions are consequential — meaning they involve trade-offs, resource commitments, or strategic direction — and establish that the new PM will not be the primary decision-maker on these for the first thirty days. This requires organizational cover from the PM’s manager. It produces a PM who has the context to make those decisions well when the protection period ends.

Conduct a synthesis review at the ninety-day mark. At the end of the first ninety days, have the PM present their understanding of the product’s strategic context, the key constraints, and the priority opportunities to their manager and a small group of senior stakeholders. The synthesis review surfaces gaps, corrects misunderstandings, and creates a shared document of what the PM understands. It also signals to the organization that the PM has invested in context before action — a signal of judgment that builds early credibility.

The difference between PM onboarding that works and PM onboarding that merely covers the surface is the difference between transmitting information and building contextual judgment. Information is available. Judgment takes time to develop — and the time it takes is not wasted. It is the investment that makes the subsequent decisions more likely to be right.

The PM who spends the first thirty days listening before proposing is not slower than the PM who begins proposing in week two. They are building the model that will make their proposals more accurate, their stakeholder relationships more durable, and their first-year performance more likely to compound into sustained impact.

There is no shortcut to context. There is only the discipline to build it before acting on it.

The best time to learn is before the decisions have consequences. The second-best time is immediately after they do — but by then, the cost is already on the table.

Leave a comment

The Minimum Lovable Action (MLA) is a tiny, actionable step you can take this week to move your product team forward—no overhauls, no waiting for perfect conditions. Fix a bug, tweak a survey, or act on one piece of feedback.

Why it matters? Culture isn’t built overnight. It’s the sum of consistent, small actions. MLA creates momentum—one small win at a time—and turns those wins into lasting change. Small actions, big impact

In the 1990s, Amy Edmondson at Harvard Business School set out to test a straightforward hypothesis: better-performing hospital nursing teams would report fewer medication errors. The data showed the opposite. The better-performing teams reported more errors than the lower-performing ones.

The explanation, which became the foundation of twenty-five years of subsequent research, was disarmingly simple: the better teams weren’t making more mistakes. They were more willing to talk about them. That willingness — the shared belief that it’s safe to speak up, admit uncertainty, and raise problems without fear of punishment or embarrassment — is what Edmondson named psychological safety. And it turns out it predicts almost everything that matters in team performance.

Google’s Project Aristotle studied 180 teams over two years looking for what separates high-performing teams from the rest. They examined team composition, individual expertise, management styles, shared interests. None of it predicted performance reliably. Psychological safety did — by a significant margin. Teams with high psychological safety had lower turnover, generated more diverse ideas, brought in more revenue, and were rated as effective twice as often by senior management. The mechanism isn’t comfort or niceness. It’s learning behavior: teams that feel safe speak up early, surface problems before they compound, seek feedback rather than avoid it, and treat mistakes as information rather than liability.

The implication for product teams is direct. An engineer who stays silent about a design flaw because they fear being seen as the person who delays the project is a psychological safety failure — not a character failure. A designer who withholds concerns about conflicting user research because the PM seems committed to the decision is the same. A PM who doesn’t challenge an unrealistic timeline in front of leadership is the same again. These silences accumulate. They produce preventable failures, slower iteration cycles, and roadmaps built on assumptions nobody felt safe questioning.

What makes psychological safety difficult to manage is that it’s invisible until it’s absent — and by the time the absence is obvious, the cost has already been paid. The teams that score lowest on psychological safety measures are rarely the ones having visible conflicts. They’re the ones where meetings run smoothly, nobody pushes back, and problems surface for the first time in production.

The Psychological Safety Pulse is a one-question diagnostic. It doesn’t measure everything. But it creates a moment — brief, low-stakes, and often surprisingly revealing — where the team gets to notice where it actually stands.

Step 1: Choose a regular team meeting happening this week

This works best in a standing meeting where the full team is present — a sprint planning, a weekly sync, a retrospective. Not a 1:1, not an all-hands. The group context is what makes the question meaningful: you’re observing how the team responds together, not just what individuals say.

Step 2: Introduce the question at the start of the meeting — before any agenda items

Frame it simply, without over-explaining: “Before we get into the agenda, I want to try something quick. I’m curious about one thing: when did someone on this team last share a mistake or an uncertainty openly — in a meeting, in Slack, anywhere?”

Then wait. Don’t fill the silence immediately. The silence itself is informative.

Step 3: Notice what happens in the first thirty seconds

This is where the real data is, and it comes before anyone says anything. Pay attention to:

  • Does anyone answer immediately, or is there a pause? A long pause in a team that meets regularly suggests people are either struggling to recall an example or uncertain whether to share it.

  • Who speaks first? Is it always the same person — typically the most senior or most extroverted? Or do different people contribute?

  • What’s the body language? In a video call, do people look at each other or away? Does anyone visibly check whether it’s safe to say something before saying it?

  • Is the example recent or historical? An answer that reaches back three months in a team that talks every week tells you something about the current climate.

Step 4: Ask one follow-up — and mean it

After the first answer, ask: “What made it feel okay to share that?” Or, if nobody answered: “What would need to be true for that to happen more easily?”

Don’t jump to solutions. Don’t reassure. Just listen to the answer and reflect it back: “So what I’m hearing is that it’s easier when [X] — is that right?” The act of naming what makes sharing feel safe is itself a small contribution to making it safer.

Step 5: After the meeting, write down three observations

Not a formal report — three sentences. What happened when you asked the question? What did the responses (or the silence) suggest about where the team is? What’s one thing you noticed that you want to pay attention to in the next two weeks?

These observations are for you, not for a presentation. They’re the beginning of a diagnostic, not the conclusion of one.

Step 6: Decide on one small signal to send this week

Psychological safety isn’t built through surveys or workshops. It’s built through repeated small moments where people observe that speaking up didn’t have a negative consequence. Edmondson’s research identifies leader behavior as the single strongest driver of team psychological safety — specifically, what leaders do when someone raises a problem, admits a mistake, or challenges a direction.

Based on what you observed, choose one signal to send this week. It might be:

  • Sharing a mistake or uncertainty of your own in a meeting, briefly and without drama.

  • Explicitly thanking someone for raising a concern — not for being right, but for raising it.

  • Responding to a challenge to your decision with curiosity rather than defense: “That’s interesting — tell me more about what you’re seeing.”

  • Naming something that didn’t go well in the sprint without assigning blame: “Here’s what I think went wrong and what I’d do differently.”

One signal. One week. The pattern it creates over time is what actually moves the dial.

For you: Running this exercise makes visible something most PMs sense but rarely name: the difference between a team that performs and a team that merely executes. A team that executes does what’s asked. A team that performs tells you when what’s being asked is wrong, surfaces problems before they become crises, and iterates faster because nobody’s hiding the things that aren’t working. Knowing where your team sits on that spectrum — and having some evidence rather than just a feeling — changes how you run your meetings, how you respond when problems surface, and how you model the behavior you want to see.

For your team: The act of asking the question is itself an intervention. When a PM asks “when did someone last share a mistake openly?” — and then genuinely listens — it signals that openness is valued, not just tolerated. Edmondson documents what she calls the “amplification effect”: a single negative response to someone speaking up outweighs ten positive ones in its impact on psychological safety. The inverse is also true — a single moment where a leader responds to vulnerability with curiosity rather than judgment can shift the team’s sense of what’s safe to say. You are creating that moment by asking the question.

For your organization: Most organizations measure psychological safety once a year in an engagement survey that nobody reads in time to act on it. The Psychological Safety Pulse is a different model: lightweight, frequent, embedded in the work rather than separate from it. Teams that develop a habit of checking in on how safely they communicate tend to surface talent that was previously quiet, catch problems that were previously hidden, and iterate faster because the feedback loops actually work. That compounds across sprints, across quarters, across product cycles. Google’s data suggests the performance gap between high- and low-psychological-safety teams isn’t marginal — it’s the difference between a team that learns and one that doesn’t.

Ask the question in your next team meeting. Write down what you notice in the first thirty seconds.

Then send one signal this week — something small that shows your team that speaking up doesn’t cost anything.

Doing this challenge? Share what you found on LinkedIn or X and tag it #MLAChallenge. The silence — or the answer — you got when you asked is almost always worth naming out loud.

Leave a comment

A Scrum Master’s Perspective

Joanna had been the Scrum Master of this product team for two years. Five developers, plus Product Owner Marta, plus the ever-present tech lead Adam, plus two designers shared with another team. A standard configuration in a Polish software house that had recently entered the top 50 largest IT companies in the region.

In October, after the publication of financial results, the leadership introduced a new bonus system. Each developer had quarterly individual targets — number of tickets closed, number of code reviews written for others, no critical bugs. Each designer had their own KPIs — number of components completed, NPS from PO. Product Owner Marta had her own set — feature delivery dates, customer satisfaction score. The entire system was presented as “performance-driven culture,” “ownership,” “accountability.” It came with a new version of the values cards in the conference room — the first value read “Team First. Always.”

The first weeks were innocent. People read the new bonus system, did some math in their heads, returned to work. Joanna noticed nothing. After six weeks, things started happening that were hard to connect to each other. Marek, one of the developers, stopped helping Karolina, the junior programmer, with her tasks. Before that, he had sat with her twice a week, explaining architecture, reviewing code. Now he said he “had his own things to do.” Karolina didn’t want to escalate, so she started documenting herself alone on Sunday evenings from their shared repo. She took on tickets within her competence range rather than the more ambitious ones, because “I need to close a certain number of them.”

Adam, the tech lead, twice refused to participate in the refinement of a neighboring team, to which he had previously been invited as a consultant. He told Joanna it “didn’t affect his KPIs.” Marta started pushing features that hadn’t been delivered in the current quarter to the next — but in a way that the formal delivery date still held. Joanna saw all this from the position where she could see it — from facilitating daily standups and retros — but only by mid-December, when the team that formally delivered all quarterly sprints somehow lost two important clients and had three serious bugs in production, did Joanna understand that she was observing something no one in the organization could name.

The bonus system had changed the team. Not in one day. Gradually, almost imperceptibly, in a way no individual manager had designed. It changed what the team was willing to do beyond the minimum of the contract — which is what had previously made it a team. And it did so in a way that evolutionary biology has taught us to recognize for a hundred and fifty years, but that HR rhetoric refuses to see.

The question of why people and animals help each other is one of the most difficult in biology. Because from the standpoint of naively read evolutionary theory, altruism shouldn’t exist. An organism that helps others at the cost of its own reproductive chances should be outcompeted by those who don’t. And yet altruism is everywhere — from ant colonies through wolf packs to human societies.

Charles Darwin himself noticed this problem. In The Descent of Man of 1871, he suggested that selection might act not only on individual organisms but also on groups. A tribe in which many people are willing to sacrifice for the common good wins competition against a tribe where everyone is selfish — even if those selfish ones individually do better within their own group. Darwin didn’t have the mathematical tools to operationalize this intuition, but he clearly stated it.

For most of the twentieth century, the idea of group selection was, as David Sloan Wilson wrote in 2007, “in disarray.” In the 1960s, William Hamilton formulated the theory of inclusive fitness — kin selection — which explained many altruistic phenomena through genetic relatedness. A worker bee sacrifices her life not for the hive as such, but for the queen, who is her genetic relative. Your genes can “prefer” altruism toward a brother, because a brother shares half of the same genes. Hamilton’s theory was mathematically elegant and explained a lot. Richard Dawkins in 1976 popularized it in The Selfish Gene — one of the most influential popular science books of the twentieth century. The conclusion was: selection acts at the level of the gene, organisms are merely “vehicles,” and altruism toward non-relatives is illusory or paradoxical.

And this consensus held for decades — until 2007, when The Quarterly Review of Biology, one of the leading journals in the field, published an article by David Sloan Wilson and Edward O. Wilson titled “Rethinking the Theoretical Foundation of Sociobiology.” Not relatives by name, but collaborating scientists. The article — in volume 82, issue 4, pages 327-348 — argued that group selection is not only possible but in many cases probably the mechanism that best explains observations. The authors introduced the term multilevel selection (MLS) and showed that selection acts at multiple levels simultaneously: gene, cell, organism, group. Some behaviors can be explained by individual selection. Others — only by group selection. Most — by the interaction of both.

I need to be honest here, because this is not scientific consensus. Steven Pinker in 2012 wrote a sharp critique titled “The False Allure of Group Selection.” Jerry Coyne, the well-known evolutionary biologist from the University of Chicago, published “The Demise of Group Selection.” Critics argue that multi-level selection is mathematically equivalent to inclusive fitness — meaning the same phenomena can be explained by one model or the other, but the models do so with different elegance and sometimes one is more practical. Wilson himself, in a 2010 article in Nature by Nowak, Tarnita, and E.O. Wilson, received a critical letter from about 150 biologists. The dispute continues to this day and is unlikely to be resolved unambiguously.

But for our purposes — for a Scrum Master thinking about a team — the key observation is one that has support on both sides of the dispute. Regardless of whether we call the mechanism multi-level selection or inclusive fitness — there is a structural tension between what is optimal for an individual organism and what is optimal for a group. Within a group, selfish behaviors are sometimes adaptive. Between groups, the advantage goes to the group in which more members are willing to be altruistic. This tension doesn’t disappear. It can be managed, it can be shifted in one direction or the other through environmental conditions — but it cannot be removed by value declarations.

And here begins the bridge to products.

For a team working on a product, incentives are what environmental conditions are for the biologist. They define which behaviors get reinforced. What punishments are borne. What rewards are achieved. Every bonus system, KPI, periodic review is — using biological language — a form of selection pressure. And the question rarely asked explicitly is: at what level does this pressure act?

Joanna’s bonus system was constructed almost exclusively at the individual level. Number of tickets closed — an individual metric. Number of code reviews written — individual. No critical bugs — individual. Each element of the motivation package measures the employee in isolation from colleagues. What’s more, some of them are structurally competitive — if the bonus pool is finite, your gain is my loss. The entire selection regime tells the employee, in a way clear to any organism responding to environment: do what maximizes your individual results.

In parallel to this system, the organization articulates values at the group level. “Team first.” “Collective delivery.” “Together as one team.” These sentences don’t have structural reality until they find reflection in selection conditions. They can hang in a conference room in beautiful framing and still not influence behavior — because selection is stronger than rhetoric. Marek didn’t stop helping Karolina because he suddenly stopped liking her. Marek stopped helping because the system clearly communicated to him that an hour spent with Karolina was an hour in which he wasn’t closing his ticket, wasn’t writing his code review, wasn’t getting closer to his bonus.

This is exactly what Edward Deci and Richard Ryan in their series of studies from the 1970s onward called “crowding out” of intrinsic motivation. In their meta-analysis of 128 experimental studies published in Psychological Bulletin in 1999, they showed that tangible rewards — money, bonuses, material prizes — systematically undermine internal motivation for tasks that are intrinsically satisfying. The mechanism is paradoxical but solidly documented. If someone does something out of curiosity or sense of craft, and then starts getting paid for it, internal motivation weakens. When the money disappears, they do it less than they did before.

Helping Karolina was not something for which Marek had previously been paid. He did it because he wanted to. Because he was experienced, because he liked her, because he saw potential in her, because it was a form of professional pride. After the new system was introduced, Marek could still help. But now helping had a clearly expressed cost — the cost of his own bonus. And that was enough for it to disappear. Not explicitly, in a communicated way. Just quietly, through withdrawal.

Bruno Frey, in his works of the 1990s and 2000s formalizing “motivation crowding theory,” showed that this effect is stronger the more personal the relationship between worker and task is, and the more autonomy the worker had before. In other words — the crowding-out effect hits hardest those teams that were previously most engaged and most autonomous. Ironically, the teams closest to the Scrum ideal.

The structural problem Joanna observed manifested in five recognizable mechanisms. It’s worth distinguishing them, because each has different dynamics — and each would require different organizational intervention.

The first is counting individual production units. Tickets, story points, lines of code, delivery days — everything that can be counted on the side of an individual worker. These are units natural for accounting, but anti-natural for teamwork. Because in a real team, the greatest value is in the interfaces — in code someone wrote, someone else reviewed, a third person integrated, a fourth tested. Value is emergent, not summative. Counting an individual contribution to such emergent value is exactly what Wilson in the Quarterly Review of Biology warned about — atomizing selection to a level at which essential information about group dynamics is lost.

The second mechanism is internal competition for limited resources. The bonus pool is limited — your gain is my loss. Promotions are limited — your promotion blocks mine. The best performance review reserved for the top 20% — for you to get it, I have to get a worse one. Each of these mechanisms is, biologically, identical to intraspecific competition for limited territory. And as dozens of studies show, such competition selects for short-term, dominance-signaling, defensive behaviors — exactly the opposite of what is expected from a mature product team.

The third mechanism is the un-disclosure of mutual help costs. Marek helped Karolina on Wednesdays and Fridays for two years. That time was not accounted for anywhere. It wasn’t his task. It wasn’t rewarded. It wasn’t even logged in any tool. It functioned outside the visible system — like the informal gift economy that Mauss described in 1925. When a system appeared that started accounting for everything an employee does against their KPIs, this invisible gift economy became an overt cost. And what became an overt cost stopped being worthwhile. The invisibility of mutual help was the condition of its existence.

The fourth mechanism is “performance management” as an individualizing tool. Most modern periodic review systems — individually-applied OKRs, performance reviews, 360 calibrations — have a feature that is rarely named. They extract individual contribution from the group context in which that contribution arose. The question “what did Marek bring this quarter” implicitly assumes that there is a meaningful answer abstracted from what Karolina, Adam, Marta, and the others brought. In a system that is truly group-level, such an answer is impossible to fully formulate — because contribution is emergent. Performance management that tries to formulate it anyway commits an act of atomization.

The fifth mechanism, the subtlest, is rhetorical asymmetry between the language of values and the language of structure. Joanna’s organization had “Team First. Always.” on the wall. It also had a bonus system that said exactly the opposite. This asymmetry is not neutral. It creates a particular kind of cognitive dissonance in employees who see both messages simultaneously. Some resolve the dissonance through cynicism — they stop believing in the values. Others through ritualistic maintenance of declarative culture while structurally behaving anti-group. In both cases the true value — authentic group action — disappears.

I have to be honest here about a serious limitation. The bonus system in Joanna’s organization was not her decision. Neither was it Marta’s or Adam’s. The decision was made at the level of the HR director and CEO. The Scrum Master in such a situation is in a position where they observe a structural problem they don’t have the tools to solve. This requires diagnostic humility — the understanding that not everything we see can be changed in the time and with the resources we have.

But there are things an SM realistically can do, worth listing.

The first is naming the phenomenon at the team level. Not necessarily critically. Not necessarily with accusation of specific people. But directly, using language that lets the team see what is happening to it. “I’ve noticed that mutual help in the team has dropped over the past two months. This isn’t personal — it’s a structural consequence of the new bonus system. I want us to think about this together.” Just naming it de-individualizes what otherwise looks like “Marek changed” or “Karolina is lazy.” And de-individualization of diagnosis is the first step toward a conversation that is not accusatory.

The second thing is documenting costs. Most organizations don’t see these costs because they have never been counted. Joanna could begin to create — even informally, for herself and her team — a register of what stopped happening in the team after the new system was introduced. Marek’s help for Karolina no longer exists. Adam no longer participates in refinements of neighboring teams. Designers stopped sitting together because “everyone has their own components to finish.” These costs are real, but they’re in the organizational shadow ledger. Bringing them to light — even in one conversation with the product director — is a political act that requires courage, but is possible.

The third thing is rituals that strengthen the group level inside the team, counterbalancing organizational pressure. A retrospective whose key question is “what did we do today together that won’t appear in any ticket.” A Sprint Review in which a story is told — not who did what, but how it came together as a product. A Daily in which time is deliberately left for “help someone,” not just “report your own.” These rituals alone won’t defend the team from structural individualizing pressure, but they can counterbalance it somewhat. Every hour spent together over a shared artifact is an hour strengthening the group level.

The fourth thing — and here I need to be honest — is escalation. Joanna won’t fix the bonus system alone. But she can present observations to the product director, the HR Business Partner, a mentor, anyone with a voice in the organizational conversation. This requires building relationships before a crisis, not during it. It requires language — and this is exactly where the literature of evolutionary biology and motivation crowding theory helps. The ability to say “this is not an individual character problem of Marek’s, it’s a phenomenon described in a meta-analysis of 128 experiments” changes the conversation’s status from anecdote to structural observation.

The fifth and final thing is protecting one’s own psyche. Because a Scrum Master who sees the atomization of a team and cannot fully stop it themselves experiences erosion. Every week of observing increasingly individualistic behaviors of people one likes and roots for is a week of emotional cost. You need someone to talk to. You need distance that allows you to maintain diagnostic autonomy. And sometimes you need to be able to admit that there are organizations that structurally don’t want to operate as a group — and then perhaps it is not the organization for us.

Wilson and Wilson wrote a 2007 article that remains disputed in evolutionary biology to this day. But one of their key insights stands firm regardless of whether we call the mechanism multi-level selection or inclusive fitness. There is a structural tension between what is optimal for an individual organism within a group and what is optimal for the group as a whole. This tension cannot be removed by value declarations. It has to be addressed structurally, through selection conditions — that is, in an organization, through what is really measured, rewarded, and punished.

Most bonus systems in modern product companies select at the individual level, while the organization’s rhetoric talks about teamwork. This contradiction is structural, not cultural. It cannot be fixed by values on the wall. It cannot be fixed by another offsite. It cannot be fixed by a new framework. It can only be fixed by changing the selection conditions — by what really counts in the organization when quarter-end comes and bonuses are negotiated.

A Scrum Master who knows multi-level selection has in their hands something few Agile Coaches in the market have — a biological argument for group rituals. It is not an aesthetic or philosophical argument. It is an empirical argument: when selection pressure favors individualism, the invisible gift economy — mutual help, mentorship, knowledge sharing, care for juniors — predictably shrinks. Not because people become worse. Because organisms respond to environmental conditions, regardless of how those organisms value their own honesty.

Joanna didn’t change the bonus system in her company. She didn’t have that power. But after the first quarter of observation, she began writing a monthly memo to the product director documenting observable costs of atomization. After the second quarter, this memo reached the HR Business Partner. After the third quarter, the bonus system was partially modified — a team component appeared, constituting twenty-five percent of the bonus pool. This didn’t fix everything. But it shifted the equilibrium point. A year later, Marek again sat with Karolina some Wednesdays. Adam returned to neighboring team refinements. Designers worked again in a room where they could look at each other and ask questions without scheduling a Zoom.

The Agile Manifesto says we value individuals and interactions over processes and tools. But people respond to selection conditions. If conditions reward interaction — interaction flourishes. If conditions reward individual productivity isolated from the group — productivity flourishes, interaction shrinks. This is not a matter of sincerity of declarations. It is a matter of structure. The empiricism we invoke involves also seeing this connection — and not pretending that good intentions suffice where structure says otherwise.

The selfish gene is not the whole truth. But in an organization where everything is measured individually, the selfish gene will win against any value hanging on the wall. Unless someone changes the conditions — or at least sees their structural consequences clearly enough to name them.

Sources:

  • Wilson, D. S., & Wilson, E. O. (2007). Rethinking the theoretical foundation of sociobiology. The Quarterly Review of Biology, 82(4), 327–348. doi:10.1086/522809

  • Darwin, C. (1871). The Descent of Man, and Selection in Relation to Sex. John Murray.

  • Hamilton, W. D. (1964). The genetical evolution of social behaviour, I & II. Journal of Theoretical Biology, 7(1), 1–52.

  • Dawkins, R. (1976). The Selfish Gene. Oxford University Press.

  • Pinker, S. (2012). The false allure of group selection. Edge.org, June 18, 2012.

  • Nowak, M. A., Tarnita, C. E., & Wilson, E. O. (2010). The evolution of eusociality. Nature, 466(7310), 1057–1062.

  • Deci, E. L., Koestner, R., & Ryan, R. M. (1999). A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. Psychological Bulletin, 125(6), 627–668.

  • Frey, B. S., & Jegen, R. (2001). Motivation crowding theory. Journal of Economic Surveys, 15(5), 589–611.

  • Mauss, M. (1925). Essai sur le don: Forme et raison de l’échange dans les sociétés archaïques. L’Année Sociologique. [English: The Gift: Forms and Functions of Exchange in Archaic Societies. Cohen & West, 1954.]

  • Lepper, M. R., Greene, D., & Nisbett, R. E. (1973). Undermining children’s intrinsic interest with extrinsic reward: A test of the “overjustification” hypothesis. Journal of Personality and Social Psychology, 28(1), 129–137.

  • Kohn, A. (1993). Why incentive plans cannot work. Harvard Business Review, 71(5), 54–63.

  • Schwaber, K., & Sutherland, J. (2020). The Scrum Guide. Scrum.org.

  • Edmondson, A. C. (2019). The Fearless Organization: Creating Psychological Safety in the Workplace for Learning, Innovation, and Growth. Wiley.

Leave a comment

We obsess over product-market fit. We prototype. We validate. We run discovery sprints, test hypotheses, iterate on solutions until they fit the problem like a key in a lock. We’d never ship a product without understanding the environment it’s supposed to survive in.

Then we hire a product team by posting a job description on LinkedIn and hoping.

I’ve run end-to-end recruitment for product organizations for a long time. Mapped team structures, defined roles, run technical interviews, matched people to problems. And the single most expensive mistake I see — across companies, across industries, across continents — isn’t hiring the wrong person.

It’s hiring the right person for the wrong architecture.

Last month I sat across from a PM candidate — sharp, strategic, fifteen years in the field. She talked about discovery with the kind of fluency that only comes from actually doing it. Hypothesis-driven. Evidence-grounded. She’d killed features her

Read the original on producttribe.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.