RSSAmplifier

Blog

Atticus Li — Better Decisions

Know what to test, when to trust the result, and what to do next. Practical decision guides for analysts, growth teams, and founders.

atticusli.comRSS feed ↗811 posts

Latest posts

What VWO Gives an Experimentation Team—and What It Cannot Decide

A source-backed guide to VWO and Wingify statistical models, stopping approaches, approvals, health checks, team workflow, and program fit.

How Booking.com Runs 1,000 Parallel Experiments—and Measures Quality

Inside Booking.com experimentation: decentralized ownership, a central platform team, power and runtime controls, CUPED, and a quality-first KPI.

Apple Product Page Optimization Uses Bayesian Testing—and Where It Stops

How Apple Product Page Optimization uses empirical-Bayes shrinkage, sequential evidence, credible intervals, and human decisions—and what it cannot prove.

How Google Runs Experiments at Scale—and Where the Evidence Stops

What Google has publicly documented about experiment infrastructure, review, power, A/A calibration, Bayesian Conversion Lift, and decision-making.

How Netflix Matches Experiment Methods to Product Decisions

A source-backed analysis of Netflix experimentation: its hub-and-spoke team, test workflow, statistical methods, decision rights, and company fit.

The AI Hype Didn’t Die. It Was Waiting for a Closed Loop

Why earlier AI waves stalled, what terminal agents changed, and how verified closed loops may reshape builders, work, and the companies we create.

What Intelligence Analysts Know About Evidence That Growth Teams Don't

Intelligence tradecraft solved the problem growth teams face daily: weighing evidence when no single source is conclusive. Here is the playbook.

What Forecasting Tournaments Say About Trusting Your Gut

Twenty years of forecasting tournaments measured what actually produces good judgment. The most valuable habit is one almost no business leader practices.

The Skill That Matters More Than the Perfect Prompt

Being specific about what 'finished' looks like matters more than finding magic wording — Claude Code fills in any gap you leave, not always the way you meant.

Do You Need to Learn to Code Before You Start?

Most of what makes someone effective with Claude Code is clear thinking, not fluency in a programming language. Here's what actually matters instead.

Vibe Coding Is a Real Way to Build Software

Describing what you want in plain English and having an AI build it produces working software, not a toy version of programming. Here's why that holds up.

What Claude Code Actually Does for You

Claude Code writes, edits, and runs real code on your computer. Here's the difference that makes, and what it means if you've never written a line of code.

Why Your Spend Limit Doesn't Survive a Fresh CI Checkout

A per-day API spend cap enforced in code still failed, because every scheduled run started from a clean checkout with no memory of prior spend.

When Prompt Caching Costs You More Than It Saves

Prompt caching is supposed to be free money. On calls spaced further apart than the cache actually lasts, it's a straight surcharge with nothing recouping it.

The Env Var That Secretly Tripled Our AI Coding Bill

We named our cost-control setting something that collided with Claude Code's own environment. Every automated run quietly inherited the most expensive option.

The Cheapest Way to Run Scheduled Claude Jobs Isn't the API

A twice-daily job billed against the API drained an account balance for days. Moving it to a subscription runner fixed the bill and the blind spot it created.

Anchoring Effect Pricing: Can the Cheapest Plan Backfire?

A pricing test made the cheapest of three plans the visual anchor -- and conversion dropped. Why anchoring on price can backfire.

Checkout Optimization: Can a Countdown Timer Hurt Conversion?

A rate-lock countdown timer worked at ticket checkout. It backfired at checkout for a recurring service. Why urgency is category-conditional.

Statistical Significance in A/B Testing: Is a Big Lift Still Noise?

A -20% topline result looked like a clear loss. It wasn't statistically significant. Why a big number and a real result aren't the same claim.

What Can a Website Heatmap Reveal About the Wrong Homepage?

A heatmap showed most homepage visitors ignored the extra pathways offered to them. Removing those paths, not adding more, won.

Can Progress Bar UX Improve Conversion Twice?

A progress bar that won at checkout got re-tested earlier in the funnel, not assumed. What transferred, and why it wasn't automatic.

Does Choice Overload Really Reduce Conversion? Our Largest Test Said No

A well-powered test of 'choice overload' came back null. What a landmark behavioral-economics finding looks like when it doesn't transfer.

Should Pricing Page Design Make the Price Less Visible?

Sometimes making a price harder to notice outperforms making it easier to justify. A seasonal pricing experiment explains why.

Can Pricing Page Design Beat a Redesign by Reordering Prices?

Reordering three prices on a pricing page outperformed a full redesign -- a decoy-effect lesson in testing cheap before expensive.

Multivariate Testing or a Confounded A/B Test: Which Did You Run?

Four bundled changes in one experiment came back inconclusive, and couldn't have told us anything either way. A confounded-test-design lesson.

Can Mobile Conversion Optimization Be as Simple as Deleting Copy?

Deleting a few sentences from a mobile modal lifted conversion by double digits -- what cognitive load teaches about 'helpful' copy.

Can Simpler Mobile Navigation Produce a Double-Digit Lift?

A decade-old mobile UX principle got tested in production instead of assumed on reputation. It held up -- here's the discipline behind why.

Do Website Personalization Examples Help—or Just Add Friction?

See why a customer-selector pop-up can fail, how two first-party chooser tests compare, and how to test useful personalization without adding friction.

Which Lead Magnet Examples Actually Earn the Download?

See why brochure previews may increase downloads, how to grade the evidence, and how to test lead magnet clarity without mistaking images for proof.

Color Psychology in Marketing: Does Matching Beat Meaning?

See what a product-color matching test really suggests, what its source omits, and how to test color congruence without relying on folklore.

How Can You Tell Whether an A/B Testing Case Study Is Trustworthy?

Evaluate any A/B testing case study with a 12-point evidence checklist covering source, sample, metrics, stopping, SRM, limitations, and transfer.

Which Navigation A/B Test Wins: Visible, Collapsed, or Removed?

Run a cleaner navigation A/B test with visible, collapsed, and removed treatments, precommitted metrics, guardrails, SRM checks, and decisions.

How Do You Optimize a Pricing Page Without Hiding What Buyers Need?

Pricing page optimization should reduce decision work without hiding comparison context. See public evidence, portfolio patterns, and a test plan.

Should a Landing Page Have Navigation—or Is It Costing Sales?

Should a landing page have navigation? Compare the public evidence, missing methods, intent conditions, guardrails, and a safer A/B test plan.

How Should You Design a Checkout Page Without Removing Trust?

Design a focused checkout page without removing trust, recovery, or control. See the research, evidence limits, guardrails, and test plan.

Four A/B Testing Examples—and How Much You Should Trust Them

See four A/B testing examples graded by evidence quality, with missing data, limits, transferable lessons, and safer next-test plans.

Sequential Testing and the SPRT: How to Stop a Test Early Without Cheating

Peeking at a fixed-sample A/B test inflates false positives. Sequential testing lets you check results repeatedly and stop early without cheating.

Triangulation Over Isolation: Building Confidence from Weak, Convergent Signals

A single underpowered test never proves anything alone. How senior practitioners stack weak, independent signals until they converge into real confidence.

The Meta-Analysis Your Experimentation Program Is Missing

Most programs audit individual tests, almost none audit the program itself. A quarterly portfolio audit answers what leadership actually wants asked.

Why Most 'Wins' Don't Replicate: The Winner's Curse, Applied to Growth Teams

The winner's curse means shipped A/B test wins systematically overstate their true effect. The fix: track predicted lift against realized lift over time.

The Confidence Tier Model: How to Decide When Your Data Isn't Enough

Most testing programs are built for traffic they don't have. Three confidence tiers — proven, directional, speculative — each with its own bet-sizing rule.

Decide What Counts as a Win Before You Test

Medicine proved that picking your primary metric after seeing the data is a structural bias. The five-minute fix most experimentation programs skip.

How to Tell If a Growth Hire Understands Risk

A great win story tells you almost nothing about judgment. Two borrowed interview probes — from forecasting research and intelligence tradecraft — do.

Why AI Coding Agents Keep Duplicating Your API Keys

Isolated AI coding sessions can't see your main .env file, so they quietly mint duplicate API keys instead of asking. The mechanism, diagnostic, and fix.

Subscription or Pay-Per-Token? I Audited My Own Claude Code Usage to Find Out

I ran my real Claude Code usage through live API pricing to see if my $200/month subscription was actually a good deal. The gap was bigger than I expected.

7 Examples of Behavioral Economics That Explain Bad Choices

Real examples of behavioral economics, ranked by evidence: which biases replicate at scale and which collapse under scrutiny.

How to Resolve a Git Merge Conflict Without Guessing What the Other Side Was Trying to Do

A clean merge isn't proof it's correct. Here's how to investigate what changed on each side — and the one conflict type worth refusing to auto-resolve.

Why Trusting an AI Assistant's Memory Is the Wrong Default (And What to Check Instead)

An AI assistant answers fluently whether a fact is current or stale. Here's the rule for knowing what to verify live instead of trusting memory.

How to Get Your AI Coding Assistant to Catch Its Own Mistakes Before You Do

The AI that wrote your draft is the worst reviewer of it. Here's the independent-review technique that catches what a second read-through misses.

Behavioral Economics Examples That Explain Your Bad Decisions

Behavioral economics examples reveal why even Microsoft's experiments succeed only a third of the time. Learn what actually works and why.