Follow us: Slack | LinkedIn | YouTube | Spotify | Substack
Community thread: Your experiment shows Variant A wins on Desktop and Variant B wins on Mobile. Is it valid to launch A on Desktop and B on Mobile, or is that statistically questionable? Why?
Throwback: Ways to destroy an experimentation culture
Jobs: New job opportunities at LendingClub and Trivago
Save 50%: Off our Amazon Canada Bestseller, Prove It or Lose It
Photo credit: Adobe Firefly
Segmenting after the fact doesn’t automatically create new winners. If your experiment wasn’t designed and powered for Desktop and Mobile decisions, opposite winners should be treated as hypotheses to validate, not immediate launch decisions.
Don’t assume randomization holds within every subgroup. Different randomization, identity, and targeting approaches across platforms can weaken the assumptions behind a device-level analysis. Before acting on segmented results, verify that assignment behaved as expected.
Passing an SRM check isn’t enough. SRM confirms the allocation ratio, not that the treatment and control groups are truly comparable within each platform. If you can’t establish that, or the experiment wasn’t designed for device-level inference, a follow-up experiment is the safest path before launching different winners.
Ezequiel Boehler: Assuming you are doing this through a frequentist RCT, was your sample size calculation done separately for Desktop vs Mobile? Because if not, and you dice the total samples by Device Type, you are not looking at the enough sample size. So I at least would not consider it a legitimate outcome. Most likely would proceed with running the test separately by device with their corresponding required sample and see if the initial observations stand.
Shiva Manjunath:What if it was run separately by device? And the results weren’t stat sig on their own, but when you math up the aggregation (because you plan on rolling both desktop and mobile to prod) it’s stat sig?
Ezequiel Boehler: Well for me is not about what post processing you end up doing, is about what was your design intention in the first place and sticking to it. So to answer that, id say then that Id do a single test next and see if the results hold. The test design and its output are a relationship about the observed sample and its population so the moment you post process the samples , you drift away from your design so the guarantees of the method stop withholding to the same degree.
Eline Schepers: We always set up tests for desktop and mobile separate, so that we can analyse and make decisions on device level as well. If its a win on both devices it’s a general website learning, if it’s different results on both devices we see it as a device based learning. As the view of a page on a mobile device can be so different from desktop and the time/attention span on mobile vs desktop is different as well we always test seperate. This way we also know if a design is good on desktop and should be tweaked and tested further on mobile or the other way around.
Marko: How often do people actually test their test results? Someone above mentioned that after doing a test - They would additionally test the results to confirm if it’s going to be same. It seems like a fail of test setup and analysis if you have to test if your test results are correct? …if let’s say the rerun test shows different results. What do you do and how do you interpret that? And do you think it would show these different results if you simply kept running tests for longer
Ishan Goel: I would specifically consider the following things.
Firstly, the sample size as has been mentioned. Each segment should be properly powered.
Multiple Comparisons: Have you been randomly segmenting every possible way to find this insight or there is a proper reasoning (that you had before the test) as to why the two results are so off.
High Variance: Is the target metric something like revenue that has a lot of outliers and hence the results seem to be just randomly fluctuating on two segments.
Overall, the result should have some explanation or you should retest and confirm it.
If you’re looking for a place where you can connect with other CROs either to talk shop, or simply shoot the sh*t - consider signing up to join our Slack.
Our Slack is a place to:
Get answers in a judgement-free zone
Share learnings
Talk shop with other CROs
You can join by signing via this form and we will send you an invite within 48 hours.
In this episode, Gerda and Rommil dive into the realities of CRO today — from AI anxiety to experimentation culture and how teams should actually be structured. We discuss:
2:07 – AI concerns and job market anxiety
2:45 – The “everything looks the same” AI problem
4:00 – Too many AI tools, no real differentiation
5:38 – What people are really asking about AI
6:23 – AI fear vs reality (and layoffs)
7:08 – Fear-mongering vs signal
8:00 – Fear and decision-making
9:00 – CRO team structures: centralized, hybrid, embedded
9:31 – Why centralized works early
10:08 – When to move to hybrid
10:29 – Why embedded is hard
11:04 – The role of leadership buy-in
12:13 – Why experimentation gets cut first
13:00 – CRO’s “nice-to-have” branding problem
14:00 – Why CRO is easy to underestimate
15:45 – Future-proofing: combine CRO with other skills
16:54 – Building experimentation culture (the “laundry” analogy)
18:19 – How to destroy experimentation culture
19:23 – Why authority and buy-in matter
20:25 – How ICs can influence culture
21:34 – Reframing experimentation as empowerment
23:09 – Becoming a catalyst across teams
24:46 – Confidence and presence in CRO roles
25:59 – CRO vs Growth: what’s the difference?
27:24 – Where CRO fits inside growth teams
28:02 – Moving CRO upstream into systems thinking
As a THANK YOU for being an active subscriber of Experiment Nation’s newsletter, you get access to Optimization Jobs, a CRO-focused job board with 2000+ open roles from the last 30 days alone.
From Optimization Jobs by Experiment Nation
Marketing Growth and Experimentation Lead
📍 trivago – DüsseldorfData Scientist, Global Growth
📍 Stripe – SingaporeSr Web Organic Growth Manager
📍 LendingClubProduct Manager, SEO & Growth (Remote)
📍 Trellis – Anywhere in the USCRM & Lifecycle Marketing Analyst (B2B)
📍 Gympass – Brazil (Remote)
To see more jobs, visit the Optimization Jobs board and explore 3000+ jobs from 400 companies from just the past 30 days that other boards may have missed.
Are you ready to prove the value of your Experimentation Program? In this actionable guide, Rommil Santiago distills over 15 years of experience and shares all you need to know to survive managing an Experimentation Program. Inside, you’ll learn:
Frameworks for communicating value effectively
Strategies to win over stakeholders
How to build a brand for your program that lasts
Prove it or Lose it is laser-focused on delivering only the actionable advice you need. Your Experimentation Program is only as strong as the value you can demonstrate.
Get your PDF copy for 50% off today!

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.