RSS Amplifier

Experimental Mind · Jun 8, 2026

'Maximize surprise ' — Digest 23, 2026

0
Sign in to vote or save

Kevin Anderson · Experimental Mind

This week: a new formula for powering switchback experiments, why good metrics need to be able to surprise you, and what "done" actually means when you're shipping AI features. Plus the UK government using experimentation to fix public services.

And as always: open roles (993!), upcoming events, and something that made me smile.

Convert—A/B tests & personalization for growth teams

Experimentation Jobs—Find your next role

Good metrics maximize surprise
By Zach Flynn — If your metric is guaranteed to go in the direction you want, you’re running propaganda, not an experiment. Good experiment metrics can go either way, and let customer behavior decide the outcome. READ

____

[Paper] Powered switchback experiments, or not?
By Sergei Pankratev (DoorDash) — Platforms use switchback experiments to avoid network effects, but had no formula to calculate statistical power. This paper fixes that. It also explains why switchback experiments require much more data than we think to achieve adequate statistical power. READ

____

Testing a different way to improve complex public services
By Tom Wynne-Morgan (UK Government Digital Service) — Most (public service) transformation fails not from lack of skill, but because no single team controls all the constraints that matter. GDS/CustomerFirst is testing a model: small, time-limited, multidisciplinary teams that surface constraints and test assumptions before committing to structural change. READ

____

What “done” means when you’re shipping AI features
By Jeff Gothelf — The definition of “done” needs to be rewritten: write acceptance criteria as probability ranges, define your failure triage before you launch, and make sure you can rollback if needed. READ

____

Last week’s favourite:

[Paper] Estimating the value of Evidence-Based Decision Making
By Abadie, Agarwal, Imbens et al. (MIT/Stanford/Amazon) — How much is running experiments actually worth? This paper builds a framework to calculate it. Using empirical Bayes on 4,8k+ Upworthy A/B tests, they show that standard p-value decision rules leave ~27–30% of attainable value on the table. READ

Axel Hommenga invited me to the Recruitment Experiments podcast. We talked about the role of job boards in general and Experimentation Jobs specifically. In Dutch, though.

If you like this episode, make sure to subscribe to the podcast. Or discover other podcasts via the Experimental Mind curated podcast feed.

Find 993 open roles on ExperimentationJobs.com. This week’s featured roles:

A running list of upcoming events. Subscribe here. (👋= join me, 🎁= discount)

We want innovation, but without the risk of failure. We want bold thinking, but look for the safe answers. We need to get better at dealing with uncertainty. Looking for growth? Become comfortable with uncertainty.

If this newsletter sparked something and you want to talk it through, a few ways I help:

  • Scale experimentation
    Strategy, setup, metrics, and ways of working for teams that want more impact.

  • Careers and hiring
    Support with your next role, or help finding and hiring strong experimenters.

  • Quick sparring
    A fresh outside perspective on your ideas, roadmap, or experimentation setup.

Interested?
Just hit reply and tell me what you’re thinking about.

Have a great week — and keep experimenting.

Thanks, Kevin

Share Experimental Mind

Read the original on kevinanderson.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.