💜 Hiding Behind Neutrality (by Alex Dziewulska)
💜 Heat Death of the Team: On the Second Law of Thermodynamics, Organizational Entropy, and the Work No One Wants to See (by Łukasz Domagała)
💪 Interesting opportunities to work in product management
🍪 Product Bites - small portions of product knowledge
🔥 MLA week#52
Join Premium to get access to all content.
It will take you almost an hour to read this issue. Lots of content (or meat)! (For vegans - lots of tofu!).
Grab a notebook 📰 and your favorite beverage 🍵☕.
Green face. Yellow face. Red face. Press one on your way out, and have a nice day.
I never press. And for a long time I assumed that made me the problem — the difficult one, the woman who can’t perform basic gratitude on cue. Then I started paying attention to why my hand never moves toward the buttons, and I realised the problem isn’t me. The problem is that I’m being asked a question that doesn’t exist.
Let me reconstruct the crime scene.
I have just left a Biedronka. Stay with me, because every element matters. The cashier was genuinely lovely — patient with the older gentleman ahead of me who paid for forty zloty of groceries in five-grosz coins. The floor looked like a small agricultural incident. The queue took nineteen minutes, three of which I spent watching a single open register while four closed ones gleamed in the distance like a taunt. The bread was stale. And the corporation that owns the whole operation I have opinions about — opinions that would not fit on a smiley face if you printed it the size of a billboard.
And now, at the exit, a cheerful plastic totem asks me to compress all of that into one tap.
Which one is the green face for? The cashier? Then I’m lying about the floor. The floor? Then I’m punishing a woman who did nothing wrong. The queue? The bread? The abstract dread I feel about late-stage retail capitalism? You handed me one button and an experience with eleven moving parts and you want a single tap to mean something coherent.
It can’t. And — this is the part that should embarrass somebody — they know it can’t.
Here is the tell. On the more expensive kiosks, the ones with a touchscreen instead of three sad buttons, watch what happens when you tap the unhappy face. A second screen appears. Was it the queue? The price? The staff? The selection? They built a follow-up question. They built an entire second step whose only purpose is to recover the information the first tap threw away. Read that back slowly. The people who manufacture these things designed the device around the knowledge that one tap tells you nothing — and then sold the version without the second screen to your local supermarket anyway, because it’s cheaper, and because a wall of faces looks like listening, and looking like listening is the actual product.
So that’s the foundation. Now let me add the Polish surcharge, because here it gets funnier.
We do not give tens. This is not a personality flaw, it’s a national setting, baked in somewhere between the partitions and the queues for toilet paper. A ten is a unicorn — it requires a service experience so transcendent that angels descend and refund your shopping. Seven is a good day. Eight means you’re suspicious of yourself for being so happy. So before the kiosk has measured anything at all, it’s running on a currency that doesn’t convert. The green face, in Polish, quietly means “fine, I suppose, nothing went catastrophically wrong, why are you looking at me.” The machine reads that as delight. The machine is an optimist. The machine has never lived here.
And then there’s me, and people like me, standing at the exit refusing to touch a row of cartoon faces because it feels like being asked to rate my school dinner with a sticker. Not because I’m above feedback — I’d happily tell you in three precise paragraphs everything wrong with that shop — but because the instrument is insulting. It assumes I’m a toddler. So I opt out. And here’s where the whole thing quietly collapses.
Because the people who opt out are not random.
The furious press red — they’ve got rage to discharge and a button is cheaper than therapy. The genuinely delighted press green, all four of them. And the enormous, reasonable, articulate middle — the people whose opinion you would actually pay money for — walk straight past, the way I do, because the format is beneath the feedback. So you are not measuring your customers. You are measuring your most emotional extremes, minus a correction factor for the children who treat the terminal as an arcade game. Which the manufacturers also know about, by the way — they sell a built-in filter to screen out button-mashing kids and bored employees. They have a feature for that. Sit with the fact that “bored eight-year-olds are corrupting our customer data” was a known enough problem to become a product spec.
So now we have a single-button instrument, measuring an eleven-variable experience, calibrated in a currency it can’t read, collecting answers only from the angry and the ecstatic, pre-filtered for toddlers. And somewhere there is a dashboard turning all of this into a trendline, and somewhere above the dashboard there is a person in a meeting saying “satisfaction is up four percent this quarter” with a straight face.
But fine. Let’s pretend none of that mattered. Let’s pretend the faces were honest, the Poles were generous, the middle showed up and the children behaved. Let’s grant the machine perfect data.
I still have one question, and it’s the only one that’s ever mattered.
What do you do when it goes red?
Because if the floor is still a swamp next week — if the queue is still nineteen minutes, if four registers still sit closed like a museum exhibit, if not one single person’s schedule changes, not one cleaning rota, not one thing — then the kiosk was never a measuring instrument. It was a shrine. A little plastic chapel bolted to the wall where you go to confess your dissatisfaction to a god who was never installed. You press red, a number ticks up in a spreadsheet nobody reads, and the experience of being heard is the entire service being rendered. You didn’t give feedback. You performed the ritual of feedback, and the kiosk performed the ritual of caring, and everyone went home.
I don’t mind being measured. I’ll tell you exactly how I feel, in detail, at length, possibly more than you wanted.
What I mind is being handed three faces by an organisation that never once decided what it would do if I pressed the red one.
Web Summer Camp 2025 — and Alex will be there
Web Summer Camp is a three-day, hands-on event for Europe’s web professionals — workshops, small group sessions, and real conversations, taking place July 3–5 in Opatija, Croatia. Not the kind of conference where you collect slides and forget about it by Monday. The kind where you actually work through problems with people who build digital products for a living.
This year they added an AI track — not because everyone else is doing it, but because they wanted to do it right. There’s also a dedicated Founders program and tracks for JavaScript, PHP, UX, and product.
Alex will be joining as a speaker. If you’re heading to Opatija this summer — or considering it — this is where to find her in person.
More about the conference: [LINK]
Ready to grab a ticket? [TICKETS]
Do you need support with recruitment, career change, or building your career? Schedule a free coffee chat to talk things over :)
!!! HOT OFFER !!!
Senior Product Manager
Product Manager
Product Manager - Allegro
Product Manager - Booksy
Product Manager - Trimble Inc.
Senior Product Manager - Box
Product Manager - Tesco Technology
In 1975, economist Sam Peltzman published a paper in the Journal of Political Economy that irritated almost everyone who read it. The U.S. government had mandated a range of automobile safety features in the late 1960s: seatbelts, padded dashboards, energy-absorbing steering columns, dual braking systems. The expectation was straightforward — safer cars would produce fewer fatalities. Peltzman studied what actually happened. Fatalities did not decline. Drivers, feeling safer, drove faster and more aggressively. Pedestrian and cyclist deaths increased. The safety equipment had not reduced the total risk in the system. It had moved it.
Peltzman’s original finding has been extensively debated and partially revised. Subsequent research confirmed that safety regulations did eventually reduce fatalities when combined with enforcement mechanisms and cultural shifts. But the core phenomenon he identified — that people systematically adjust their risk-taking behavior in response to perceived changes in safety — is robust, documented across dozens of contexts, and directly relevant to how product teams design safety mechanisms, guardrails, and error-recovery systems.
A product team adds an auto-save feature to their document editor. The intention is clear: users will lose less work. The actual outcome is more complex. Users stop manually saving before attempting risky operations. Some begin editing critical documents in conditions they would previously have considered too unstable — on a slow connection, with multiple tabs open, while multitasking. The rate of lost work decreases per session. The rate of sessions involving risky behavior increases. Whether the net effect on data loss is positive, neutral, or negative depends on specifics the team didn’t model when they shipped auto-save.
This is the Peltzman Effect operating in a product context. The safety net changed behavior, and the behavior change partially or fully offset the intended safety gain. Not because users were irrational — they were making perfectly reasonable adjustments to a changed risk environment — but because the product team designed the feature without modeling the behavioral response it would generate.
The Peltzman Effect, also called risk compensation or the offset hypothesis, describes the tendency for people to increase their risk-taking behavior when they perceive that their safety has been improved, partially or fully offsetting the intended benefits of the safety measure.
Peltzman drew on the concept of risk homeostasis, developed by psychologist Gerald Wilde: the idea that individuals maintain a target level of subjectively acceptable risk. When environmental changes reduce perceived risk below the target, people adjust their behavior to restore the familiar level — by driving faster, skipping precautions, or attempting actions they would otherwise have avoided. The adjustment is not always conscious and is not necessarily irrational. It reflects a genuine optimization: people are trading reduced risk in one domain for increased benefit in another, calibrated to their preferred risk tolerance.
The research since Peltzman has found that risk compensation exists across a wide range of domains — cycling helmet use, anti-lock braking systems, avalanche airbags, antiretroviral medication adherence — and that the magnitude of the offset is typically partial rather than complete. Safety improvements generally produce net positive outcomes. But the behavioral adjustment consistently reduces the benefit below the level the safety measure would have delivered if behavior had remained constant.
For product teams, the implication is not that safety features are counterproductive. It is that they are never purely additive. Every safety feature, guardrail, or error-recovery mechanism changes the risk calculus of the users who interact with it — and that changed calculus produces behavioral adjustments that product teams rarely model when designing safety into their products.
A useful mental model for the Peltzman Effect is the risk budget: people behave as if they have a subjective budget of risk they are willing to spend in a given context. When a safety feature reduces the perceived cost of one type of risk, the freed-up budget tends to be spent elsewhere. This is not a conscious financial transaction. It is a behavioral adjustment that happens automatically as the risk environment changes.
In product design, the risk budget heuristic predicts that adding safety features — undo buttons, auto-save, confirmation dialogs, data backups — will generate behavioral adjustments in proportion to how much the feature reduces perceived risk. Features that eliminate catastrophic downside risk will tend to generate larger behavioral adjustments than features that reduce minor inconveniences. This is the context in which the Peltzman Effect is most significant for product teams: the features designed to prevent the worst outcomes are exactly the features most likely to generate the behavioral adjustments that bring users closer to those outcomes.
Peltzman’s most troubling finding was not that safety measures were ineffective overall — they reduced occupant fatalities. It was that they redistributed risk to people who had not been offered the safety improvement: pedestrians and cyclists who faced more aggressive drivers without the benefit of seatbelts and airbags. In product contexts, an analogous redistribution can occur: a safety feature that reduces risk for the primary user may generate behavioral adjustments that increase risk for secondary users, downstream systems, or organizational processes that depend on the primary user’s behavior.
A product that gives users powerful automation capabilities with easy rollback, for example, may generate more frequent and more aggressive use of those automation capabilities — which reduces net risk for individual users while increasing the rate of errors that reach production environments before being caught. The safety feature improved individual risk. The behavioral adjustment shifted risk upward in the system.
The Peltzman Effect is structurally related to moral hazard: the phenomenon in which access to insurance or protection from consequences changes behavior in ways that increase the likelihood of the insured event. The most direct product application is in financial products — lending, investment, and payment platforms where backstop features can change how users engage with financial risk. But moral hazard operates in any product context where a safety mechanism reduces the cost of failure: backup systems, test environments that shadow production, fraud protection, and account recovery mechanisms all create conditions for Peltzman-style behavioral adjustment.
Robinhood’s design of protective features for options trading illustrates the Peltzman Effect operating at the product level. The platform introduced educational overlays and risk disclosure steps before options trading — mechanisms intended to ensure that users understood the risks they were taking. Research on user behavior following the introduction of these features found patterns consistent with risk compensation: users who had completed the risk disclosure flows engaged with options trading at higher rates and with more complex strategies than users who had not. The safety feature — intended to reduce harm from uninformed risk-taking — may have served, for some users, as a signal that they were now sufficiently informed to take on the risks the disclosures described, generating behavior the feature was designed to prevent.
Git’s version control system demonstrates the Peltzman Effect as a positive force when correctly anticipated. The ability to commit code, branch freely, and roll back changes at any point removes the catastrophic downside from code experimentation. This was intentional: the design explicitly models the behavioral adjustment. Engineers who know they can roll back will experiment more aggressively. The safety feature is designed to generate the behavioral adjustment — more risk-taking in the form of more experimentation — because the experimentation is the desired outcome. The Peltzman Effect is not a bug in this design. It is the product’s primary value proposition.
Cybersecurity provides some of the most consistent documentation of Peltzman dynamics in enterprise software. Research published by ISACA on organizational security behavior consistently finds that users who have access to security tools — password managers, VPNs, multi-factor authentication — exhibit increased risk tolerance in adjacent behaviors: clicking links they would otherwise have avoided, using the same password across more accounts before changing it, connecting to public networks more frequently. The security features reduce specific risk vectors. The behavioral adjustment introduces risk in adjacent vectors. Net security outcomes are generally positive, but the Peltzman adjustment is a consistent and measurable component of the outcome that security teams underestimate when evaluating the expected impact of new security features.
The Peltzman Effect matters for product teams because it describes a systematic failure mode in safety feature design: assuming that the behavioral response to a safety improvement will be constant, when in fact safety improvements reliably change behavior in ways that partially offset the improvement.
This assumption error leads to consistent overestimation of safety feature impact. When a product team models the expected benefit of an auto-save feature, a confirmation dialog, or an undo mechanism, they typically calculate the benefit as if user behavior remains unchanged. The Peltzman Effect says it will not. Users will adjust. The adjustment will reduce — and in some cases reverse — the expected benefit.
There is also a design implication for how safety features are communicated. Features that are prominently signaled as safety mechanisms generate larger behavioral adjustments than features that provide equivalent protection without explicit safety framing. A product that tells users “your data is automatically backed up every minute” creates a different risk budget adjustment than one that provides the same backup silently. The explicit framing reduces perceived risk more dramatically, generating a larger Peltzman adjustment. This is not an argument against transparent communication — it is an argument for thinking carefully about how safety signaling changes user behavior, not just how safety mechanisms change user outcomes.
Model the behavioral adjustment explicitly when evaluating safety features. Before shipping a safety feature, ask: how will users behave differently because this feature exists? Which behaviors that were previously avoided will now seem acceptable? Who in the system bears the risk of those adjusted behaviors? The answers will not always suggest not shipping the feature — but they will produce a more accurate model of the feature’s actual impact.
Design safety features to enable valuable risk-taking, not just to prevent costly errors. When the behavioral adjustment is desirable — when more experimentation, more creative use, or more confident decision-making is the goal — design safety features that explicitly enable that adjustment. Git is the canonical example: the safety mechanism is designed to generate the behavioral adjustment that produces the product’s value. When the behavioral adjustment is undesirable, design safety features that are less perceptually salient, reducing the magnitude of the Peltzman response.
Measure behavior changes, not just outcome changes. Standard evaluation of safety features measures whether the bad outcome occurs less frequently. This misses the Peltzman Effect entirely if the frequency of the behavior generating the outcome increases in parallel. A complete evaluation measures both the rate of the negative outcome per behavior instance and the rate of the behavior itself. If the rate per instance decreases while the frequency of the behavior increases, the net effect may be smaller than the outcome-only metric suggests.
Watch for redistribution to adjacent systems and users. When a safety feature generates behavioral adjustments in primary users, identify whether those adjustments shift risk to secondary users, downstream systems, or organizational processes that the safety feature did not protect. The pedestrian-and-cyclist problem in Peltzman’s original research is an organizational and systems design problem as much as a behavioral one: the risk didn’t disappear, it relocated.
Sam Peltzman’s original finding was not that safety features are bad. It was that people are not passive recipients of safety improvements. They are active agents who continuously calibrate their behavior to their perceived risk environment, and who adjust that behavior when the environment changes. Safety features change the environment. The behavior adjustment follows.
For product teams, this is ultimately a systems thinking challenge. Safety features are not interventions in a static system — they are inputs to a dynamic system in which users respond, adjust, and optimize continuously. Modeling those responses requires the same rigor as modeling the direct effect of the feature itself.
The seatbelt made the driver safer. The driver drove faster. The pedestrian paid the difference.
Product design that does not anticipate this dynamic is not designing for safety. It is designing for the appearance of safety — and leaving the behavioral adjustment to chance.
In 1945, at the Trinity nuclear test site in New Mexico, physicist Enrico Fermi watched the first atomic bomb detonate. As the shockwave reached his position, he dropped a handful of small pieces of paper from chest height and watched how far the blast carried them. From the distance they traveled, he estimated the energy of the explosion at ten kilotons of TNT. The actual yield was later calculated at twenty-one kilotons. Fermi was within a factor of two — on the first nuclear weapon ever detonated, using paper scraps and arithmetic.
Fermi was known for this: his ability to arrive at defensible estimates for quantities no one could directly measure, using only available information, logical decomposition, and a tolerance for approximate answers. His approach — break a complex unknown into components that can each be estimated from what you already know, then combine them — became the foundational technique of what physicists call order-of-magnitude estimation, and what product teams desperately need more of.
A product team is evaluating whether to build a new feature that would allow users to connect their accounting software directly to the product. The feature would take approximately six weeks of engineering time. Leadership asks whether it’s worth it. The PM calls for a proper market sizing analysis. The analysis takes three weeks, involves two external data sources, and concludes with a range of potential revenue impact that spans an order of magnitude — from two hundred thousand dollars to two million dollars annually.
The analysis was more expensive than the information it produced. Three weeks of PM and analyst time for a conclusion that didn’t materially change the decision that intuition had already suggested: it depends on how many users actually have the target accounting software, and what the conversion rate from the feature to retained customers looks like.
Both of those estimates could have been produced in an afternoon, with reasonable confidence, using available data and structured reasoning. The Fermi approach would not have given the same precision as the formal analysis. It would have given 80 percent of the decision-relevant insight in 5 percent of the time.
Fermi estimation is a problem-solving technique for arriving at an approximate answer to a question that cannot be directly measured, by decomposing the question into components that can be estimated from available information and then combining those estimates systematically.
The technique is named for Enrico Fermi, who taught it at the University of Chicago and used it routinely in his experimental and theoretical work. Fermi problems — also called order-of-magnitude estimates or back-of-envelope calculations — are structured around a consistent logic: identify the unknown, decompose it into factors, estimate each factor from available knowledge, multiply or combine, and check whether the result is plausible against any anchor points available.
The classic Fermi problem asks how many piano tuners are in Chicago. The answer requires no data lookup. It requires estimates of the population of Chicago, the proportion of households with pianos, the frequency of piano tuning, and the number of tunings a piano tuner can complete per day. Each estimate can be made from general knowledge. The combined estimate — a few hundred piano tuners — is close to the actual number and is useful for any purpose requiring a Chicago piano tuner count.
For product managers, the technique is applicable to a far broader range of questions than are typically recognized. Market sizing, feature adoption estimation, opportunity sizing, customer lifetime value calculation, and competitive impact assessment can all be structured as Fermi problems — decomposed into estimable components and solved in hours rather than weeks.
The core of Fermi estimation is decomposition: breaking a question that cannot be answered directly into questions that can. This sounds simple and is not always easy. The quality of the decomposition determines the quality of the estimate. A decomposition that identifies the genuinely load-bearing variables — the ones whose uncertainty most affects the final answer — produces useful estimates. A decomposition that misses a critical factor or that mixes correlated variables produces estimates that are precise but wrong.
A useful test for decomposition quality is to identify the variables that most affect the final answer and ask whether each can be independently estimated. If two variables in the decomposition are highly correlated — both driven by the same underlying factor — the estimate will be overfitted to that factor and will miss the others. An independent check on whether each component estimate is plausible given general knowledge of the domain is also useful.
The essential discipline of Fermi estimation is accepting order-of-magnitude accuracy rather than pursuing false precision. An estimate that is correct within a factor of two or three is, for most product decisions, sufficient. Pursuing greater precision typically requires either more data collection or more modeling effort — and the question to ask before doing either is whether the decision changes depending on which end of the range the true answer occupies.
The Fermi approach makes this test explicit. A preliminary estimate that returns a range from one hundred thousand to one million users does not make the decision harder — it makes the critical uncertainty visible. If the decision is the same for both numbers, further analysis is not warranted. If the decision differs, the analysis should be targeted at the specific variable driving the range, not at the overall model.
Every Fermi estimate should be checked against at least one independent anchor. If the estimate of new accounting software integration users comes out to 50,000 in the first year, the plausibility check asks: does this seem reasonable given the total user base, current conversion patterns, and the nature of the feature? If the estimate implies that 40 percent of all users would adopt a feature in its first year, something is wrong with the decomposition. The plausibility check catches arithmetic errors, misplaced decimals, and decompositions that look correct internally but produce results that don’t survive contact with domain knowledge.
The technique has specific failure modes that define when more formal analysis is warranted. When the decision is genuinely sensitive to precision within an order of magnitude — when it matters whether the answer is 100,000 or 200,000, not just whether it’s 100,000 or 1,000,000 — Fermi estimation is insufficient and more rigorous data collection is warranted. When the variables in the decomposition are not independently estimable from available knowledge — when each depends on empirical data that hasn’t been collected — the technique produces outputs that look structured but rest on circular assumptions. And when the stakes of the decision are high enough that a factor-of-two error produces significantly different outcomes, the investment in more precise analysis is justified.
Amazon uses order-of-magnitude estimation explicitly in the early stages of the Working Backwards process. Before a product press release is written or a formal business case constructed, teams are expected to produce rough estimates of customer reach, expected usage, and business impact that are sufficient to assess whether the initiative warrants further investment. These estimates are deliberately rough — the goal is to identify whether the opportunity is in the range of interesting or trivially small — and are constructed through decomposition of available information rather than formal market research. The technique filters out a substantial fraction of potential initiatives that appear attractive qualitatively but are immediately revealed as economically marginal when the numbers are approximately quantified.
Stripe’s approach to product prioritization uses Fermi-style reasoning at the team level. When evaluating whether a feature serving a specific developer segment is worth building, the team estimates the size of the segment (developers using the target language or framework), the proportion who have the specific pain point, the conversion rate from pain point to feature adoption, and the revenue retention impact of serving that segment better. Each component is estimated from available developer ecosystem data, usage patterns, and known conversion rates from similar features. The result is not a financial model — it is a structured check on whether the opportunity is worth a deeper investigation. Most of the time, the rough estimate provides enough signal to either proceed or move on without the deeper investigation.
Intercom’s product team has publicly described using what they call “napkin math” before formal discovery on any significant initiative — a structured decomposition of the potential impact that is produced in a single working session before any research, user interviews, or data pulls are scheduled. The napkin math does two things: it identifies the variables most worth investigating (those that most affect the final answer and can be changed by research) and it establishes a baseline expectation against which research findings can be evaluated. When the napkin math suggests a potential reach of 50,000 users and the user research identifies a segment of 5,000, the discrepancy is itself informative — either the napkin math was wrong, or the research methodology missed a significant portion of the relevant segment.
Fermi estimation matters for product teams because it addresses one of the most persistent dysfunctions in product organizations: the gap between the speed at which decisions need to be made and the time required to make them well with formal analysis. Product decisions are frequent and often time-sensitive. Formal analysis is thorough and often slow. The gap between them is usually filled by intuition — which is sometimes excellent and often systematically biased.
Fermi estimation occupies the space between intuition and formal analysis. It is faster than formal analysis by orders of magnitude, more structured than intuition by a significant degree, and sufficient for the majority of product decisions that require only order-of-magnitude guidance. It does not replace rigorous research when research is warranted. It dramatically reduces the volume of decisions for which research is mistakenly treated as warranted when it is not.
The technique also changes the nature of the conversation between product teams and leadership. A PM who can produce a defensible order-of-magnitude estimate in a meeting — one that shows the reasoning, identifies the key uncertainties, and gives a range that is honest about its limits — is demonstrably more valuable than one who defers every quantitative question to a formal analysis process that takes weeks. The estimate is not the final word. It is the starting point for a conversation about which uncertainties matter enough to investigate.
Decompose before you research. When faced with a quantitative question that requires decision support, start with a decomposition before initiating any data collection. Write down the components of the estimate, make your best guess at each, and calculate the result. This exercise takes thirty minutes. It identifies which components are genuinely uncertain and which are stable — and therefore which research is worth doing and which isn’t.
Set an explicit threshold for “good enough.” Before beginning any quantitative analysis, define what level of precision would change the decision you’re facing. If the decision is the same for any answer in a given range, and your Fermi estimate falls clearly within that range, stop. Additional precision is not decision-relevant and is therefore not worth the cost of collecting it.
Document the reasoning, not just the estimate. A Fermi estimate that presents only the conclusion — “we think this will reach 50,000 users” — is indistinguishable from intuition. A Fermi estimate that presents the decomposition — “we estimate 2 million users in the segment, 40 percent with this type of workflow, 15 percent adoption in year one” — is auditable, correctable, and useful for identifying where better data would most improve the estimate. Always show the reasoning.
Use the estimate to identify what you don’t know. The most valuable output of a Fermi estimation session is often not the estimate itself but the list of variables that most affect the answer and are most uncertain. These are the specific data collection targets — the questions that research should answer — rather than the general topic area. Research that is targeted at specific uncertainties is dramatically more efficient than research that is structured around a general topic.
Check every estimate against an independent anchor before presenting it. The plausibility check is not optional. An estimate that cannot survive contact with available anchor data — industry benchmarks, historical conversion rates, comparable feature performance — has an error in the decomposition that needs to be found and corrected before the estimate guides a decision.
Enrico Fermi’s paper scraps at Trinity were not a precise instrument. They were a structured observation, combined with known physics and clear reasoning, that produced a useful approximation of an answer that could not be directly measured. The precision was not the point. The point was to have a defensible number rather than no number — to convert an unknown into an estimated range that could be acted on.
Product teams face the same challenge at a smaller scale every week. The market size of the target segment is unknown. The adoption rate of the proposed feature is unknown. The revenue impact of the strategic initiative is unknown. Waiting for precise answers to all of these questions before making decisions is a strategy for paralysis. Making decisions without any quantitative structure is a strategy for chronic surprise.
Fermi estimation is neither. It is the discipline of producing approximate answers to important questions, quickly, honestly, and with explicit acknowledgment of the uncertainty involved. It will not always be right. It will almost always be useful.
Drop the paper. Watch how far it flies. Make your estimate. Move forward.
There is a type of diversity that rarely appears in diversity reports, hiring metrics, or team retrospectives. It is not measured by gender, ethnicity, or professional background, though it correlates with all of these. It is harder to see, harder to count, and harder to intentionally build — and research consistently shows it is more predictive of team decision quality than any of the more visible dimensions.
Cognitive diversity is the variation in how people process information, frame problems, generate ideas, and evaluate solutions. It is not what people know that differs — it is how they think about what they know. And in product teams, where the core work is making good decisions under uncertainty, it may be the most important compositional variable that most teams have never deliberately addressed.
A product team is making a decision about whether to redesign a core workflow. The PM has assembled the relevant stakeholders: two senior engineers, a designer, a data analyst, and a customer success manager. The meeting goes smoothly. Discussion surfaces two or three approaches, the group converges quickly, and the decision is made with apparent confidence and consensus.
Three months later, the redesigned workflow has a serious problem that nobody anticipated during the planning process. Looking back, the issue was foreseeable — it was visible in a category of user behavior that the team’s mental model did not include. The engineers were thinking about technical elegance. The designer was thinking about visual clarity. The analyst was thinking about metric impact. Nobody on the team thought primarily about the operational context in which the workflow would actually be used.
The team was diverse by most measures. By cognitive profile — by the mental models, frameworks, and problem-solving approaches represented in the room — it was homogeneous. The blind spot was structural.
Cognitive diversity is the variation in perspectives, information processing styles, problem-solving approaches, and mental models that different people bring to shared problems. The concept was formalized in the work of mathematician and social scientist Scott Page, whose 2007 book The Difference provided mathematical grounding for the intuition that diverse groups often outperform homogeneous groups on complex problems.
Page’s central finding was that the performance advantage of cognitively diverse groups is not simply additive — it does not arise merely from having more information in the room. It arises from the interaction between different models of the problem, which generates insights that no single model would produce alone. A group that includes people who frame a problem differently is more likely to identify solutions that a group of similar thinkers would miss, even if the similar thinkers are individually more expert.
Research by Katherine Phillips, then at the Kellogg School of Management, extended this finding to the social dynamics of diverse groups. Her experiments demonstrated that groups including a socially distinct outsider — someone perceived as different from the core group — performed better not only because of the outsider’s ideas but because their presence changed the behavior of the existing group members. Homogeneous groups took their agreement for granted. Diverse groups worked harder to explain their reasoning, surface their assumptions, and examine their conclusions. The presence of difference raised the epistemic standards of the whole group.
For product teams, cognitive diversity is not a social benefit. It is a structural property that affects the quality of the decisions the team makes, the range of solutions it generates, and the blind spots it is able to identify before they become problems.
Cognitive diversity and demographic diversity are related but not identical, and conflating them produces confusion about both. Demographic diversity — variation in gender, ethnicity, age, national origin — tends to correlate with cognitive diversity because different backgrounds, experiences, and cultural contexts produce different ways of framing and approaching problems. A team that is demographically diverse is more likely, on average, to be cognitively diverse than a team that is demographically homogeneous.
But the correlation is not sufficient. A team can be demographically diverse and cognitively homogeneous if all its members were trained in the same discipline, worked in the same industry, and share the same professional frameworks. Conversely, a team that appears demographically homogeneous can achieve meaningful cognitive diversity through deliberate selection of people with different professional backgrounds, different reasoning styles, and different approaches to uncertainty. The goal is cognitive diversity. Demographic diversity is one reliable path toward it, but not the only one and not a guarantee.
Page’s framework identifies four distinct dimensions along which cognitive diversity operates: information, which refers to the different facts and domain knowledge people possess; heuristics, the problem-solving rules and shortcuts different people apply; perspectives, the mental representations or models people use to frame problems; and predictive models, the theories people use to forecast how systems will behave. Diversity across all four dimensions contributes to group problem quality. But diversity in perspectives and predictive models tends to be most valuable for the strategic and product decisions that matter most — and is least likely to be achieved by teams assembled primarily on the basis of functional role or professional expertise.
Cognitive diversity is not free. Research by Amir Goldberg at Stanford’s Graduate School of Business, analyzing hundreds of thousands of messages from distributed work teams, found that cognitively diverse teams are more effective during ideation and exploration phases of work but less efficient during coordination and execution phases. The variation in mental models that generates better solutions during problem-framing can slow convergence during implementation. Teams that do not actively manage this dynamic tend either to suppress diversity during early phases to gain efficiency or to maintain it during late phases and suffer coordination costs.
The practical implication is that cognitively diverse teams benefit from explicit phase management: creating conditions that actively encourage diverse reasoning during discovery and decision-making, and actively converge toward a shared model during execution. Teams that try to maintain equal cognitive diversity across all phases of work will tend to underperform on both dimensions.
Phillips’s research documented a consistent failure mode in cognitively homogeneous groups: the false consensus effect, in which group members overestimate how much other members agree with them or think like them. Homogeneous groups assume agreement without testing it, which means they spend less time examining their conclusions, less effort surfacing dissenting evidence, and less cognitive work on the problem overall. The result is not just worse decisions — it is a group that is more confident in those decisions, having done less to challenge them.
For product teams, this manifests as the confident retrospective error: the post-launch discovery that the team was unified in a shared assumption that nobody questioned, not because the assumption was sound, but because nobody in the room had a mental model that would have flagged it as an assumption at all.
Pixar’s creative review process — the Braintrust — is one of the most documented examples of institutionalized cognitive diversity in a product-adjacent context. The Braintrust brings together directors, writers, and producers with different creative backgrounds, different strengths, and explicitly different aesthetic sensibilities to review films in development. President Ed Catmull has described the process as deliberately including people who will see problems that others miss, not because they are more capable but because they approach the work from different frames. The process does not require consensus — the director retains creative authority — but it systematically surfaces perspectives that a more homogeneous review group would not generate. Pixar’s creative record during the Braintrust period suggests the approach works; its research and development failures that predate and postdate full Braintrust engagement provide the counterfactual.
Amazon’s product and strategy teams explicitly value what they call “intellectual curiosity across domains” in PM hiring — a proxy for the cognitive profiles of people who have built mental models in multiple fields and can apply them to new problems. The emphasis on written communication in Amazon’s culture also serves a cognitive diversity function: the practice of writing down reasoning forces articulation of the mental models underlying a recommendation, making them visible and therefore challengeable by people operating from different models. The six-page narrative memo, evaluated silently before discussion, is partly a mechanism for surfacing cognitive diversity — for ensuring that different readers’ different frameworks are applied to the same text before the social dynamics of the meeting suppress the differences.
At Figma, hiring for design tool development has deliberately included people with backgrounds in domains adjacent to design: game development, film production, architecture, and education. The cognitive profiles these backgrounds produce — different relationships to collaboration, iteration, and the relationship between creator and audience — have been explicitly cited as contributing to product decisions that tools built by teams of exclusively software and design backgrounds would not have made. The multiplayer design session, Figma’s most distinctive product innovation, emerged from a mental model about collaboration that was not native to how most design software teams thought about their product.
Cognitive diversity matters for product teams because it directly affects the quality of the decisions the team makes at the most consequential moments — when the product’s direction is being set, when user problems are being framed, when trade-offs between competing directions are being evaluated. These are exactly the moments when the range of mental models in the room most determines what gets seen and what gets missed.
The research is consistent: homogeneous groups are more comfortable, more efficient, and more confident in their conclusions. Cognitively diverse groups are less comfortable, less efficient, and more likely to be right. Product organizations that optimize for comfortable, efficient decision-making processes will tend to build teams that are good at executing a known direction and poor at identifying when the direction needs to change.
There is also an innovation dimension. Page’s mathematical work demonstrated that the performance advantage of cognitively diverse groups is most significant for hard problems — problems that require novel approaches, that do not have precedent, and where existing frameworks are insufficient. As products become more complex, markets more competitive, and user needs more sophisticated, the advantage of cognitive diversity compounds. Teams solving easy, well-understood problems can afford homogeneity. Teams solving hard, novel problems cannot.
Audit the cognitive profiles represented in key decision rooms. The question is not who is in the room — it is what mental models are represented. For any significant product decision, ask: what problem-framing perspectives are present? What disciplines, domains, or backgrounds are not represented? The goal is not to add people for every missing category — it is to identify the categories of perspective that are most relevant to the decision and to ensure at least one person in the room holds each.
Build cognitive diversity into hiring criteria explicitly. Beyond functional skills and domain experience, evaluate candidates for the mental models they bring — their problem-solving approach, the analogical domains they draw on, the way they decompose problems that are novel to them. Interviews that ask only about domain expertise will select for people who think about the domain the way existing team members do. Interviews that surface reasoning style, cross-domain pattern-matching, and approaches to ambiguity will select for cognitive diversity.
Create conditions for cognitive diversity to express itself. A team with cognitive diversity whose norms suppress dissent or penalize unconventional framing will not perform better than a cognitively homogeneous team. The research benefit requires that different perspectives actually surface during decision-making. Structured processes — written reasoning before group discussion, explicit devil’s advocate roles, pre-mortems that invite alternative framings — create the conditions for cognitive diversity to affect outcomes rather than being suppressed by social dynamics.
Manage the coordination cost deliberately. Cognitively diverse teams are less efficient during execution. Building explicit convergence mechanisms — clear decision authority, documented reasoning for chosen directions, norms that distinguish ideation phases from commitment phases — allows teams to benefit from diversity during exploration without paying the full coordination cost during implementation.
Scott Page’s core finding was deceptively simple: diverse groups of problem solvers can outperform groups of the best individual problem solvers. The mechanism is not that diverse groups contain more knowledge. It is that they contain more models — more ways of framing the problem — and that the interaction of those models produces insights that no single model would generate alone.
For product teams, this translates into a concrete organizational challenge: building teams whose cognitive profiles are as thoughtfully selected as their functional skills. Most product organizations hire for expertise and culture fit. The expertise requirement is necessary. The culture fit requirement, if it selects for cognitive similarity, may be the primary mechanism by which teams become blind to the problems they don’t already know how to see.
The teams that build the most durable products are not always the teams with the most domain expertise. They are often the teams that saw problems others missed — that asked the question nobody else thought to ask, that applied the framework from a different field that turned out to be exactly right. That capacity does not arise from individual genius. It arises from the deliberate cultivation of different minds working on the same problem.
Hire for the perspectives you don’t already have. Build the conditions for those perspectives to be heard. Manage the tension between diversity and efficiency as an organizational design challenge, not as a culture problem.
The blind spot you cannot see is the one that will cost you most.
The Minimum Lovable Action (MLA) is a tiny, actionable step you can take this week to move your product team forward—no overhauls, no waiting for perfect conditions. Fix a bug, tweak a survey, or act on one piece of feedback.
Why it matters? Culture isn’t built overnight. It’s the sum of consistent, small actions. MLA creates momentum—one small win at a time—and turns those wins into lasting change. Small actions, big impact
The word comes from Japanese gardening. Ne means roots. Mawashi means going around. Nemawashi — literally “going around the roots” — describes the practice of carefully preparing a tree for transplanting: loosening the soil, tending the roots, making the ground ready long before the move actually happens. Applied to business, it describes something most experienced product managers know intuitively but rarely practice deliberately: the real work of getting a decision made happens before the meeting where the decision is announced.
In Japanese organizational culture, nemawashi is a formal practice — proposals circulate through informal conversations before they ever enter a meeting room, so that by the time a decision is officially presented, everyone in the room has already shaped it. The meeting isn’t where decisions get made. It’s where decisions that have already been made get documented. This might sound slow. In practice, research on organizations that use the approach consistently shows faster implementation, stronger stakeholder alignment, and significantly fewer reversals after decisions are made — because the resistance that would have surfaced in the room has already been addressed, privately and without the defensive dynamics that public disagreement tends to produce.
The parallel for product teams is exact. Most PMs have experienced the meeting where a roadmap decision gets ambushed — a stakeholder raises a concern nobody anticipated, a priority gets challenged in front of senior leadership, a direction that seemed settled gets reopened. These moments feel like bad luck or bad faith. Often, they’re neither. They’re the predictable consequence of a process that waited until the official moment to surface the concerns that were always there.
The alternative isn’t to avoid decisions or to run every choice by committee. It’s to do the conversational groundwork before the room assembles — so that the meeting becomes a confirmation of alignment rather than an attempt to achieve it under pressure. Ideas that go through genuine pre-consultation are stronger, not weaker. They’ve been challenged privately, refined by people with different perspectives, and stress-tested by the concerns of the people who will ultimately need to implement them.
This MLA is one application of that principle. One decision. Two or three conversations. Before the meeting.
Step 1: Identify a decision coming up in the next two weeks that matters
This works best with something real — a roadmap priority that isn’t yet settled, a proposed change to how the team works, a recommendation you’re planning to bring to leadership. It doesn’t need to be large. It needs to be something where the outcome matters and where at least two or three people have a stake in the direction.
Don’t choose something already decided and communicated. Nemawashi is pre-decision groundwork, not retroactive buy-in.
Step 2: Map who needs to not be surprised
Write down the names of the people who, if they walked into the room without prior conversation, might raise a concern, push back on the direction, or feel that their perspective hadn’t been considered. These aren’t necessarily the most senior people. They’re the people whose reaction matters most to how the decision lands.
Aim for two to three names. If you have more than five, you’re mapping an all-hands, not a nemawashi.
Step 3: Have the conversations — individually, informally, before the official meeting
Reach out to each person separately. Not with a formal briefing or a structured agenda — with a question. “I’m working through a decision about [X] and I wanted to get your read on it before I bring it to the team. Do you have fifteen minutes this week?”
In the conversation, lead with the problem you’re trying to solve, not the solution you’ve already landed on. Listen for concerns that change your thinking. Listen for constraints you didn’t know about. Listen for the objection they would have raised in the room — and address it now, when there’s no audience and no pressure.
If you hear something that shifts your direction, update accordingly. That’s not weakness. That’s the point.
Step 4: Note what you learned — and what changed
After each conversation, write one sentence: what did you learn that you didn’t know before? It might be a constraint, a competing priority, a concern about timing, a piece of organizational context that changes how you’d frame the decision. It might be nothing — confirmation that your direction is sound is also useful information.
If you had three conversations and nothing changed, either your decision is genuinely well-grounded or you’re talking to the wrong people. Both are worth knowing.
Step 5: Walk into the official meeting differently
When the decision reaches the room, you’ll notice something: the people you spoke to will engage differently. They’ve already had their concerns acknowledged. They’ve already shaped the direction, even slightly. They’re not hearing this for the first time and looking for a way in. The dynamic is different — less defensive, more constructive, more likely to move toward closure rather than reopening settled ground.
You don’t need to announce that you did this. The quality of the conversation will show it.
Step 6: After the meeting, compare what you expected with what happened
Did the conversations you had beforehand change how the meeting went? Were there concerns you’d already addressed that might otherwise have derailed the discussion? Were there people who didn’t receive a pre-conversation whose reaction surprised you?
Write two sentences: what went differently than it would have if you’d gone straight to the meeting, and what you’d do differently next time.
For you: The most common source of roadmap frustration for product managers isn’t strategy — it’s execution friction that was entirely predictable but wasn’t surfaced early enough. Nemawashi builds the habit of treating stakeholder alignment as continuous groundwork rather than a one-time event at the moment of announcement. PMs who practice this consistently make fewer decisions that get reversed, spend less time in damage control after meetings, and develop a much more accurate model of how their organization actually makes decisions — as opposed to how it’s supposed to make them.
For your team: When a PM does the conversational groundwork before bringing a decision to the room, the quality of the discussion improves substantially. Concerns that would otherwise surface defensively — in a meeting where raising them feels like opposition — have already been processed. What’s left is the kind of productive disagreement that refines a direction rather than derails it. Teams that develop this norm tend to have shorter, more decisive meetings and fewer “we need to take this offline” moments, because the offline conversation already happened.
For your organization: The alternative to nemawashi isn’t efficiency — it’s the kind of slow-motion reversal that happens when decisions get made in rooms where not everyone who mattered was adequately consulted. Research on organizations using structured pre-consultation approaches consistently shows faster overall implementation timelines, despite the upfront investment in individual conversations. The time spent loosening the roots before the transplant is exactly what makes the transplant stick.
Find the decision. Identify the two or three people whose reaction matters most. Have the conversations before the meeting.
Then notice what’s different when the room assembles.
Doing this challenge? Share what you found on LinkedIn or X and tag it #MLAChallenge. The concern someone raised in a private conversation — that would have derailed a meeting — is usually the most useful thing to name out loud.
A Scrum Master’s Perspective
Marek had been the Scrum Master of this team for nine months. The team had four years of prior history. Three founding developers who remembered choosing the project’s name. A senior tester who joined after the first year. Two developers who came on board a year ago in a recruitment wave. A Product Owner who had changed three times. A designer who was new. Eight people in total who talked about themselves as “a team” but whose mutual respect was decreasing in a way Marek initially couldn’t name.
He read his predecessor’s notes. Notes from March of the previous year, when the team was described in every account as excellent — fast delivery, high quality, low conflict, people helping each other, retrospectives were productive. He read notes from June of this year, when Marek started — and he saw something different. Retrospectives are short, people don’t talk about problems, twice a month someone asks in a one-on-one “do I even fit here.” Sprints last the same length, but more things inside them don’t make sense — documentation is old, tests have exceptions no one remembers, code has fragments marked with a “refactor later” comment from two years ago.
In his third meeting with the team lead — Kasia, whom I’ll call tech lead for simplicity, though formally she is a “senior developer with extra responsibilities” — Marek asked the question that had been on the tip of his tongue for two months. “What happened? When did the team stop being what it was?” Kasia looked at him with the expression of someone who had repeatedly considered this and found no satisfying answer. “I don’t know. No one left. No one had a big fight. There was no crisis. It happened slowly. Week by week some thing disappeared. No one noticed the thing in the moment it was disappearing. After a year we look at each other and we’re not what we were.”
And in this seemingly innocent sentence of Kasia’s — “it happened slowly, week by week some thing disappeared” — lies something organizations rarely name clearly, although physics described it a century and a half ago, and cybernetics extended it to social systems over seventy years ago. I’m talking about the second law of thermodynamics — and the fact that a team is a system that structurally gravitates toward dissolution, not because of anyone in particular, but because that’s how complex systems behave when no one actively pumps energy into them.
This article is about how the second law of thermodynamics translates onto product teams. I have to start with an honest warning — this is a metaphor. The second law in physics is literal; in application to organizations it’s an analogy. Some physicists consider such analogies abuses. Others — from Norbert Wiener in 1948, through Stafford Beer in the 1970s, to contemporary knowledge entropy literature — accept them as legitimate, provided they are used with discipline. I will show why it’s worth deploying and where its reach ends.
In physics, closed systems gravitate toward a state of maximum entropy — meaning maximum disorder, maximum dispersion of energy, maximum uniformity. A cup of hot coffee placed in a room does not spontaneously become hotter; it becomes colder, until it equalizes with the room’s temperature. A sugar cube dropped in water does not reassemble itself into a cube; it dissolves into solution. A room no one cleans does not become cleaner; it becomes dirtier. All of these are manifestations of the same law — without an input of energy, order disintegrates, and entropy rises.
Importantly, the second law does not say order is impossible. It says order requires continuous supply of energy from outside. Life is the perfect example — living organisms are local “islands” of low entropy in a sea of rising cosmic entropy. They maintain that order by importing energy from outside (food, sunlight) and exporting entropy outside (heat, waste). When that flow stops, the organism dies — and then spontaneous decomposition processes bring it to thermodynamic equilibrium with its surroundings.
This thought is worth holding in mind, because it — though in different substrate — exactly describes what happens to a team when no one actively invests energy in maintaining its order.
In 1948 Norbert Wiener, mathematician at MIT, published a book titled Cybernetics: Or Control and Communication in the Animal and the Machine. The book defined a new discipline — cybernetics — that studied mechanisms of control, communication, and feedback in biological, mechanical, and social systems. In chapter seven, titled “Information, Language and Society,” Wiener takes a step that became foundational for all subsequent applications of entropy to organizations. He states that societies, like organisms, are local low-entropy systems that, to maintain themselves, must import information (analog to energy) and export disorder. Without this import — Wiener writes — societies decompose.
This was a bold intellectual move. Wiener took a concept from physics, joined it with Shannon’s concept of information (published in the same year, in parallel), and stretched both onto biological and social organisms. Some physicists to this day consider this an abuse. Some biologists have mixed feelings. But the fact is that Wiener’s framing opened the door for an entire tradition of thinking that proved fruitful in practice, regardless of methodological caveats.
The most important figure who carried this tradition directly into management was Stafford Beer. Beer, a British cyberneticist and consultant, built in a series of books from Brain of the Firm (1972) through Designing Freedom (1974) — based on Massey lectures at CBC — to Heart of Enterprise (1979) a model he called the Viable System Model. According to Beer, an organization is a system that, to remain “viable,” must have a minimal set of regulatory and informational functions that counterbalance the natural tendency toward disorganization. Beer wrote directly: an organization without active, conscious investment of energy in maintaining order heads toward “heat death” — a state of maximum entropy, where all differences equalize and the system stops functioning as a distinct entity.
Contemporary knowledge management and organizational learning literature continues this intuition with precision. Constantin Bratianu and collaborators, publishing from the early 2000s, use “knowledge entropy” as a formal category. Their thesis, summarized in works appearing in Knowledge Management Research & Practice and other journals: knowledge in an organization disperses over time if not actively maintained, documented, rotated among people, and structurally protected. This is not a loose metaphor — it can be measured (through knowledge map audits, through tracking who knows what, through documenting the cost of losing an employee with specific knowledge). Multiple studies show that knowledge loss is a real, measurable organizational cost, but one rarely accounted for explicitly. A 2023 systematic review of 91 empirical studies on turnover-induced knowledge loss shows that this cost is both widespread and serious, but almost always underestimated by leadership.
All of this is empirical superstructure for a metaphor that started with Wiener. The second law of thermodynamics is literal in physics; in organizations it’s an analogy. But the analogy is strong, because it’s supported by both theoretical coherence and observational evidence.
For a product team, entropy manifests in five recognizable forms. Each requires a different kind of “energy” for counterbalance. It’s worth distinguishing them, because a Scrum Master who sees only one of them fights a single symptom, not the totality of the phenomenon.
The first is knowledge entropy. Kasia from the opening scene, when she started on the team, remembered who wrote which component. She knew who to go to with a question about the payment module. She knew there was a specific exception in tests for a customer in Lithuania, because she had been there when that exception was added. Two years later, Kasia still remembers — but her two new developers do not, because they were not there then. The designer, who joined six months ago, does not know why certain screens look different in the B2B context than in B2C. Knowledge that was once evenly distributed across the team now concentrates in the heads of a few people — and parts of that knowledge are nowhere anymore, because the people who held it left the company long ago. This is knowledge entropy — the rise of disorder in the “who knows what” map, up to a state in which no one knows who knows what.
The second is process entropy. When the team was new, Sprint Planning lasted two hours and everyone knew why. Definition of Done was discussed every quarter, and each team member could recite it. Refinement had its rhythm. Four years later, Sprint Planning still lasts two hours, but no one remembers why — it just lasts, because “it has always been so.” Definition of Done is a document no one has read in two years. Refinement starts informally after Daily, because “there’s a meeting anyway.” Processes that were once living responses to specific needs become ritual forms no one understands anymore — this is the entropy of form: when structure survives the meaning that birthed it.
The third is relational entropy. Four years ago, three founding developers went together to lunch every day. They called each other by name not in a business sense but in a relational one. They remembered each other’s children, vacations, health problems. Today, one of them works remotely from another city, the second spends lunches with the new Python crew, the third is looking for a new job and no one knows. New developers have never seen them together at lunch, so they don’t even know what’s missing. The social capital that once existed disappears not through conflict, but through the slow withdrawal of people from small, invisible interactions that created that capital. This is relational entropy — the increase of social distance between members of the same formal organizational unit.
The fourth is goal entropy. When the team started, everyone knew what they were trying to build and why it mattered. The Product Vision was fresh, exciting, was explained to every new employee in the first week. Four years later, the vision exists in some document, but no one remembers exactly where. The roadmap changed so many times that it became an unreliable artifact. OKRs are written every quarter, but their connection to the larger strategy is not clear to anyone. The new developer, when asked what the team is trying to achieve, says “we do this project” — not “we solve this problem for these users.” Goal disperses in actions. Actions remain, but understanding of where they lead blurs. This is goal entropy.
The fifth is quality entropy. This is the most tangible and most-cited form. Code that was once clean accumulates layers of historical decisions, patches, workarounds, “TODOs” from five years ago. Tests that once provided real confidence have disabled individual cases “just for this sprint, we’ll fix later.” Documentation that was current is three versions of the system out of date. Each of these things is rational at the moment it appears — it was a time saving on some sprint. But the cumulative sum of these rational shortcuts gives a system that is harder every quarter to understand, to change, to test. Software engineers call this “technical debt,” but this is only one form of a more general phenomenon — entropy of the quality of a system not actively maintained.
Here a central observation appears. All the activities a good Scrum Master performs — and many of them described in the Scrum Guide as ceremonies and artifacts — are in essence anti-entropy mechanisms. Work whose purpose is not “to create something new” but “to prevent what already exists from dispersing.”
The retrospective is an anti-entropy mechanism against process entropy. Its function is not “continuous improvement” in the naive positivist sense, where each sprint we do something better. Its function is refreshing the connection between the form of the process and its meaning. When a team asks at retro “does this Sprint Planning still serve us?”, it is not searching for improvement — it is protecting the process from degeneration into empty ritual.
Onboarding is an anti-entropy mechanism against knowledge entropy. When a new developer joins the team, it’s not just about showing them the codebase. It’s about transferring knowledge from the heads of people who have been on the team for years into the head of the new person. Without that knowledge transfer, knowledge concentrates more and more in the heads of individuals, and when those people leave — which statistically they will inevitably do at some point — the knowledge disappears.
Knowledge sharing, code review, pair programming are anti-entropy mechanisms against dispersal of knowledge. Each has a cost in hours. Each looks like “double work” to someone counting only closed tickets. But each disperses knowledge across the team to enough people that the loss of one of them does not capsize the system.
Sprint Review, when well conducted, is an anti-entropy mechanism against goal entropy. It is a moment when the team returns to the question: “what are we actually doing and for whom?”. Without that moment — with the usual rhythms of work, where tickets roll from one sprint to the next — the team loses contact with the goal, even while productively working.
Team building, shared lunches, coffees, hallway conversations — which are most often cut first when an organization looks for savings — are anti-entropy mechanisms against relational entropy. Each has a cost in time. Each looks like “work that isn’t in Jira.” But each maintains the social capital whose reconstruction, when it disappears, is disproportionately costly compared to maintaining it.
Refactoring, technical debt, documentation, writing tests are anti-entropy mechanisms against quality entropy. Here we have the most developed literature, because software engineering understands these phenomena technically. But the principle is exactly the same: we work against the system’s natural tendency toward disorganization.
And here appears the second central observation we draw from physics. Anti-entropy work is invisible. By definition invisible. Because it does not create anything new — it only prevents something existing from disappearing. If a retrospective is good, the team doesn’t feel “something” changed — they feel “everything is as it was.” If onboarding is good, the new developer doesn’t bring “revolutionary quality” — they simply enter the existing structure smoothly. If pair programming is good, the code doesn’t become “better in some sense” — it simply doesn’t become worse. The value of anti-entropy work is negative in the mathematical sense: its measure is what did not happen thanks to it.
This is exactly the reason this work is cut first when organizations look for savings. Because nothing is lost immediately. Everything still works. After a week of work without retrospective, the team still delivers. After a month without code review, code still compiles. After a quarter without team building, people still communicate at work. Everything looks fine. Until after a quarter, or two, or three, the team is what Marek’s team is — everything ostensibly the same, but something has fallen apart. And no one remembers exactly when.
The second law of thermodynamics has one particularly unpleasant property when applied to organizations. It tells us that anti-entropy work is nonlinear in consequences. You can save once, twice, three times — and nothing terrible happens. But the cumulative savings eventually cross a threshold past which the system does not return to its prior state on its own. Kasia from the opening scene cannot be “fixed” by intensifying retrospectives for a month. The knowledge that disappeared from the heads of people who left has really disappeared. The relationships that withered through a year without shared lunches will not rebuild through one offsite.
This is the trap organizations fall into regularly. They have cost pressure — a quarter comes when something must be optimized. They look at team costs and see things they cannot weigh. Number of closed tickets — can be counted. Number of retrospective hours — can be counted. Number of knowledge sharing meetings — can be counted. And the answer seems easy: let’s cut these “soft” things, let’s keep the “hard” production. After three months, team metrics even improve — because time was saved on tickets. After a year, the same team no longer delivers them, because the system fell apart — and no one can say precisely why.
Here the Scrum Master has a role no one else in the organization has. The role of translator between two languages: the language of accounting, which measures what is visible, and the language of systems, which sees what is invisible. The SM should be able to say: “yes, the retrospective costs four hours of work for the entire team every two weeks. Yes, those four hours can be saved. Yes, in the first month we will not notice the difference. But in six months we will, and then it cannot be undone as quickly as it is cut.”
This requires language. That is why the second law of thermodynamics and the concept of organizational entropy are worth knowing. Not as academic ornament. As a political tool — an argument with which one can explain to the CFO why “soft” activities are not luxuries. An argument that is not based on “I know better” — it is based on “this is a property of systems, not just my conviction.”
Marek did something concrete when he understood what had happened to his team. He wrote a one-page document for the product director titled “Invisible Costs of Optimization.” In the document he described what had disappeared from the team over the last two years and, where possible, when specifically it had been optimized. Pair programming was eliminated last August after the decision “we have delivery pressure.” Knowledge sharing was shortened to fifteen minutes per sprint after November of the previous year. Shared lunches disappeared after the COVID year and no one brought them back. Each of these decisions was rational at the moment. The cumulative sum created what Marek was observing — a team where everything ostensibly was there, but nothing worked as it should.
The document reached the product director, who read it, was silent for a long time, and said one sentence: “I understand.” Then a six-month reversal optimization program — diplomatically named “Re-investing in Team Capability” — was introduced as a pilot. After six months, Marek’s team did not return to the state of last March, because some things do not return. But it returned to a state in which it stopped further falling apart. Kasia again remembered who to invite for help on a new module. The designer and tester started talking about what “good UX for this scenario” means. The three founding developers again showed up together at Friday lunch.
The second law of thermodynamics is not organizational science. It is physics. But Wiener, Beer, and generations of their students showed that certain properties of complex systems are fundamental enough to transfer between substrates. Systems that maintain internal order do so through active, continuous investment of energy. Without that investment they fall apart. Not through anyone’s fault. They just work that way.
A product team is such a system. Not a magical exception to laws governing other complex assemblies. It equally requires continuous investment of energy — in retrospectives, in knowledge sharing, in team building, in refactoring, in onboarding, in keeping the vision alive. Everything that a Scrum Master considers their daily work has deeper justification than “the Scrum Guide says so.” It has justification in the physics of complex systems — justification that most financial managers do not see, because it was not explained to them in their language.
A Scrum Master who understands this language has a role in the organization no one else has. They are not an organizational expert in Agile. They are a translator between two levels of description — between the accounting level, which measures what is visible, and the systemic level, which sees what works. Without this translator, organizations repeatedly fall into the trap that physics described a century and a half ago. They optimize visible costs. Accumulated entropy remains invisible until the moment the system falls apart — and then it is too late for the return to the prior state to cost less than maintaining it would have cost earlier.
The Agile Manifesto says we value individuals and interactions over processes and tools. But individuals and interactions are precisely what falls apart fastest without active, conscious, costly work to maintain them. Processes and tools persist regardless, even when empty. Individuals and interactions require investment that cannot be optimized away without consequences.
The empiricism we invoke in Scrum involves also seeing this nonlinearity. Seeing that saving four hours of retrospective in this sprint has a price that will be paid not in this sprint, but six quarters out. Seeing that the knowledge that disappears from the heads of departing people is a real cost, even if no one books it. Seeing that relationships that evaporate through slow withdrawal from small interactions are an asset that is not on any balance sheet.
The heat death of the team will not happen from one decision. It will happen from hundreds of rational micro-decisions distributed over time. And that is precisely why it is invisible — until it stops being. A Scrum Master who knows the second law of thermodynamics sees it earlier. And can speak about it before it becomes irreversible.
That is their work. That is the value they bring to the organization. That is the reason — if the organization listens to them — they can still save something.
Sources:
Wiener, N. (1948). Cybernetics: Or Control and Communication in the Animal and the Machine. MIT Press.
Shannon, C. E. (1948). A mathematical theory of communication. Bell System Technical Journal, 27(3), 379–423.
Beer, S. (1972). Brain of the Firm: The Managerial Cybernetics of Organization. Allen Lane.
Beer, S. (1974). Designing Freedom. CBC Learning Systems / John Wiley & Sons.
Beer, S. (1979). The Heart of Enterprise. John Wiley & Sons.
Beer, S. (1985). Diagnosing the System for Organizations. John Wiley & Sons.
Bratianu, C. (2019). Exploring knowledge entropy in organizations. Management Dynamics in the Knowledge Economy, 7(3), 353–366.
Sutter, S., Smith, R., & Massingham, P. (2023). Knowledge loss induced by organizational member turnover: A review of empirical literature, synthesis and future research directions. Journal of Knowledge Management. Systematic review of 91 empirical studies.
Ton, Z., & Huckman, R. S. (2008). Managing the impact of employee turnover on performance: The role of process conformance. Organization Science, 19(1), 56–68.
Clausius, R. (1865). On various forms of the fundamental equations of the mechanical theory of heat. (Formulation of the second law of thermodynamics.)
Senge, P. M. (1990). The Fifth Discipline: The Art and Practice of the Learning Organization. Doubleday.
Nonaka, I., & Takeuchi, H. (1995). The Knowledge-Creating Company. Oxford University Press.
Schwaber, K., & Sutherland, J. (2020). The Scrum Guide. Scrum.org.
Edmondson, A. C. (2019). The Fearless Organization. Wiley.
The neutral researcher doesn’t exist.
Not as a rare achievement. Not as an aspiration we’re all failing toward. Not at all. The neutral researcher is a fiction the research industry tells itself — and right now, the people clinging hardest to that fiction are the ones losing their jobs.
I want to take this apart properly, because the argument I keep hearing goes like this: researchers should be neutral. They shouldn’t have domain knowledge, because domain knowledge means assumptions, and assumptions mean bias, and bias means

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.