पाठशाला Pathshala · उत्पाद Utpād, The product · Lesson 27 · Scale

An experimentation culture with real statistical power

Most product ideas do not work, so a test that cannot detect the effect it hopes for will mostly declare false winners. Size every test before it starts and read it once.

Pathshala, The Founder Library · 11 October 2026 · 7 min read

Three petri dishes with cultures in yellow and pink on a white bench.
Photograph: Edward Jenner · Pexels

A company that runs a hundred A/B tests a year and ships every winner can spend a quarter building on results that were never real. The cure is not fewer tests. It is tests that are large enough to tell a real effect from noise.

This lesson explains why most declared winners at a growing company are suspect, the arithmetic of statistical power in plain words, a figure to check any test before it runs, and the habits that keep an experimentation programme honest when traffic is thin, as it is for most Indian startups.

Why so many wins are not wins

Start with the base rate. Ron Kohavi, Alex Deng and Lukas Vermeer collected published success rates in A/B Testing Intuition Busters, presented at KDD in 2022. About a third of experiments succeeded at Microsoft, 15 per cent at Bing, around 10 per cent at Booking.com, Google Ads and Netflix, and 8 per cent in Airbnb search. The Microsoft figure goes back to a 2009 paper by Kohavi and colleagues, which found that “only about 1/3 of ideas improve the metrics they were designed to improve.” These are teams with years of practice. A startup’s ideas are not better on average.

Now combine that base rate with the test’s threshold. If one idea in ten is truly better, and the team calls a winner whenever the result is significant, some of those winners are false positives from the nine ideas that did nothing. The same paper computes the share: with a 10 per cent success rate and a test at 80 per cent power, 22 per cent of declared winners are false. At 20 per cent power, the figure is 52.9 per cent. More than half the wins are noise, and each one becomes a feature the team maintains and a belief it repeats.

Power, in plain arithmetic

Four numbers decide whether a test can work. The baseline: today’s conversion or retention rate. The minimum detectable effect: the smallest lift worth finding, written before the test, usually as a share of the baseline. The significance threshold: how surprising a result must be before it is called, conventionally a 5 per cent two-tailed test. And the sample: users in each arm. Power is the chance that a real lift of the stated size will be detected. The convention is 80 per cent.

Laboratory glassware casting long shadows across a white surface.
A small effect needs a large sample to see. Size the test before it starts or the shadow is all you measure. Photograph: Ron Lach · Pexels

The relationship that matters is that the sample needed grows with the square of how small the effect is. Halve the lift you want to detect and you need roughly four times the users. Lower baselines need more users too, because rare events are noisy. This is why a test of a button colour on a page converting at 3 per cent is hopeless for most companies, and why the [landing page lesson](/library/landing-page-and-first-conversion-test) asks for the sample before the headline.

At the defaults, a 5 per cent conversion rate, a 10 per cent relative lift and 20,000 users in each arm, the test has about 61 per cent power. With one idea in ten truly winning, about 27 per cent of declared winners are false, and twenty such tests find a little over one real win and about half a false one. Raise users to 31,000 per arm and power reaches 80 per cent. Lower the lift to 5 per cent and even 100,000 users per arm leave power near 72 per cent. Move the prior to 30 per cent, the rate of a team whose ideas come from good research, and the false share falls below 10 per cent at the same traffic. Better ideas are a statistical asset.

A test that cannot detect the effect it hopes for does not produce a weak answer. It produces a confident wrong one.

The habits that waste a quarter

Peeking. Checking a test daily and stopping when it looks significant inflates false positives. Evan Miller shows in How Not To Run an A/B Test that a test of a change that does nothing, checked after every observation and stopped at the first significant result, reports a false winner 26.1 per cent of the time against a nominal 5 per cent. His advice is to decide the sample size in advance and wait, or use a sequential design built for early looks.

Many metrics, one story. A test that tracks twenty metrics will usually find one that moved by chance. Name one primary metric before launch and treat the rest as diagnostics. The winner’s inflation. An underpowered test that does reach significance tends to overstate the effect, because only the lucky draws cross the line; expect the shipped lift to be smaller than the test said. Novelty. Users click on new things. Run tests across at least one full weekly cycle, two for products with a weekly rhythm, and check whether the effect fades in the second week.

Broken randomisation. If the split was meant to be fifty-fifty and one arm has noticeably more users, something in the assignment or the logging is wrong, and the result cannot be trusted however significant it looks. Check the counts in each arm before reading any metric. No guardrails. A change that lifts sign-ups can quietly raise refunds, support tickets or page load time. Name two or three guardrail metrics for every test and treat a significant harm on any of them as a loss, whatever the primary metric did.

Culture is what leaders do with a losing test

Tools are the easy part. Stefan Thomke’s Building a Culture of Experimentation in Harvard Business Review describes Booking.com running some 25,000 tests a year, and argues that the hard change is in leaders, who must confront the possibility that they are wrong every day and let data outweigh opinion. In a startup that means the founder’s own ideas go through the same test as everyone else’s, and a losing result on a founder’s favourite feature is reported in the same tone as any other.

It also means rewarding the right thing. A team praised for the number of winners it ships will learn to run underpowered tests and peek, because both produce more winners. A team praised for well-designed tests and for what it learnt, including from losses, will produce fewer and truer wins. Put the power calculation and the pre-registered metric in the review, not only the result.

When traffic is too thin

Many Indian B2B products have a few hundred customers and a consumer app may have a few thousand checkouts a week. Classical A/B testing on small effects is out of reach, and pretending otherwise is how teams ship noise. Four alternatives work. Test bolder changes: a new pricing page rather than a new button, because large effects need small samples. Run longer and accept fewer tests. Use holdouts for big launches, keeping 10 per cent of users on the old experience for a quarter and comparing retention cohorts. Combine evidence: a result that is directionally positive, matches what [users said in interviews](/library/talking-to-users-while-you-build) and moves a leading indicator is a decision worth making, even if it is not a statistical proof.

A worked example: an edtech checkout in Bengaluru

A Bengaluru company selling exam-preparation courses gets 1,200 checkout visits a day and converts 8 per cent of them. The growth team wants to test a redesigned UPI payment step and hopes for a 5 per cent relative lift, from 8.0 to 8.4 per cent. Two weeks of traffic gives 8,400 visits in each arm. Power is about 15 per cent. With a one-in-ten prior, nearly 60 per cent of any winner it declares would be false, and reaching 80 per cent power needs about 74,000 visits per arm: more than four months.

The team changes the question. Instead of polishing the payment step, it tests a different offer on the same page: an EMI option shown upfront with the monthly figure in large type. It expects a lift near 20 per cent if the idea works at all, which needs about 4,900 visits per arm for 80 per cent power, a little over eight days of traffic. The test runs fourteen days to cover two weekly cycles, is read once, and its result goes into the log whatever it says.

The fortnightly experiment review

Every two weeks, the product and growth leads review one log. Before a test starts its row holds: hypothesis, primary metric, baseline, minimum detectable effect, sample per arm, power and the fixed end date. If power is below 80 per cent, the test is redesigned or not run. After it ends: the result with its confidence interval, the decision, and a note on what was learnt. Each quarter count the tests run, the share that won and the share of past winners whose effect held in the following month’s data. That last number is the honest measure of an experimentation culture.


The figure uses the normal approximation for two proportions and the Bengaluru company is illustrative. Published rates are as stated by their sources, checked 11 October 2026.

Sources

  1. Ron Kohavi, Alex Deng and Lukas Vermeer, A/B Testing Intuition Busters, KDD 2022 — Success rates from 8 to 33 per cent; false positive risk of 22 per cent at a 10 per cent success rate and 80 per cent power.
  2. Evan Miller, How Not To Run an A/B Test — Peeking raises a nominal 5 per cent false positive rate to 26.1 per cent in the worst case.
  3. Stefan Thomke, Building a Culture of Experimentation, Harvard Business Review, March–April 2020 — Booking.com runs some 25,000 tests a year; leaders must accept being wrong.
  4. Ron Kohavi et al., Online Experimentation at Microsoft, 2009 — Only about one third of ideas improve the metrics they were designed to improve.