“How long should we run this A/B test?” is the most practical statistics question in product data science interviews, and it usually comes with a trap attached: “The dashboard already shows p = 0.03 on day four. Can we stop?” Interviewers ask both because they map directly onto the job. A candidate who can size a test from first principles, translate users into calendar days, and explain why early stopping inflates false positives is a candidate who will not ship noise. A candidate who says “run it until it’s significant” is describing exactly how teams fool themselves.
Before you start
You need the vocabulary of hypothesis testing: alpha (the false-positive rate, usually 0.05), power (the chance of detecting a real effect, usually 0.8), and a p-value. You should know that the standard error of a proportion is sqrt(p(1 - p) / n). The code uses the Python 3.14 standard library (random.binomialvariate was added in 3.12); outputs in comments come from running the blocks in order with the seeds shown.
The short answer
The per-arm sample size for comparing two conversion rates is n = (z₁₋α/₂ + z₁₋β)² × (p₁(1 - p₁) + p₂(1 - p₂)) / (p₂ - p₁)². For a baseline of 10%, a lift to 11%, alpha 0.05 and 80% power, that is about 14,750 users per arm. I convert that into days using eligible traffic and round up to whole weeks. Then I fix the duration in advance, because checking repeatedly and stopping at the first p below 0.05 pushes the false-positive rate from 5% to roughly 24% over 20 looks; if the team needs to look early, I use a sequential design with adjusted boundaries.
How it works
Sample size is a tug of war between signal and noise. The signal is the difference you want to detect, the minimum detectable effect (MDE). The noise is the standard error of the difference, which shrinks with sqrt(n). You need enough users that a real MDE-sized difference would usually land beyond the significance threshold: 1.96 standard errors for the threshold, plus another 0.84 standard errors so that it clears the threshold 80% of the time.
import math, random
from statistics import NormalDist, fmean
Z = NormalDist().inv_cdf
def sample_size(p1, p2, alpha=0.05, power=0.80):
z = Z(1 - alpha / 2) + Z(power)
return z ** 2 * (p1 * (1 - p1) + p2 * (1 - p2)) / (p2 - p1) ** 2
print(round(sample_size(0.10, 0.11), 1), math.ceil(sample_size(0.10, 0.11))) # 14748.0 14749Round sample sizes up. The memorable shortcut 16 × p(1 - p) / δ² comes from 2 × (1.96 + 0.84)² ≈ 15.7 and gives 14,400 here, close enough for a whiteboard.
Step-by-step walkthrough
Step 1: Choose the metric and the minimum detectable effect
The MDE is a business decision: the smallest lift worth shipping. Because it sits squared in the denominator, it dominates the answer. Be explicit about absolute versus relative lifts; “a 5% lift” on a 10% baseline is 10.5%, not 15%.
for target in (0.12, 0.11, 0.105):
print(target, math.ceil(sample_size(0.10, target)))
# 0.12 3839
# 0.11 14749
# 0.105 57760Halving the MDE from 1 point to 0.5 points almost quadruples the users. Picking an MDE from what the change can plausibly achieve, rather than from wishful thinking, is what keeps tests feasible.
Step 2: Account for power and the traffic split
More power and unequal splits both cost users. A 50/50 split is the most efficient for a fixed total, because the variance of the difference is proportional to 1/n₁ + 1/n₂.
print(math.ceil(sample_size(0.10, 0.11, power=0.90))) # 19744
print(round((1 / 0.1 + 1 / 0.9) / (1 / 0.5 + 1 / 0.5), 2)) # 2.78Going to 90% power costs a third more users. A cautious 90/10 split needs 2.78 times the total traffic of a 50/50 split to reach the same power, which is often the hidden reason a “small, safe” test never reaches significance.
Step 3: Convert users into days
If 3,000 eligible users a day reach the page (eligible means they actually see the changed experience, not all site visitors), the test needs both arms filled:
per_arm = math.ceil(sample_size(0.10, 0.11))
days = math.ceil(2 * per_arm / 3000)
print(days, math.ceil(days / 7) * 7) # 10 14Ten days rounds up to 14. Whole weeks matter because weekday and weekend users behave differently; a test that runs Monday to Thursday describes a different population than the one you will ship to.
Step 4: Simulate what peeking does
Now the trap. Suppose 750 users per arm arrive per day for 20 days (15,000 per arm, about the planned size), and someone checks the z-test every evening. The function below returns the first day a stopping rule fires, or None.
def z_stat(c1, c2, n):
p = (c1 + c2) / (2 * n)
se = math.sqrt(p * (1 - p) * 2 / n)
return (c2 - c1) / n / se
def first_stop(rng, p_treat, boundary, daily=750, days=20):
c1 = c2 = 0
for day in range(1, days + 1):
c1 += rng.binomialvariate(daily, 0.10)
c2 += rng.binomialvariate(daily, p_treat)
if abs(z_stat(c1, c2, day * daily)) > boundary(day):
return day
return None
rules = {
"fixed horizon": lambda d: 1.96 if d == 20 else math.inf,
"peek daily": lambda d: 1.96,
"Pocock": lambda d: 2.67,
"O'Brien-Fleming": lambda d: 2.12 * math.sqrt(20 / d),
}
rng = random.Random(3)
for name, rule in rules.items():
aa = [first_stop(rng, 0.10, rule) for _ in range(20_000)] # no real effect
ab = [first_stop(rng, 0.11, rule) for _ in range(5_000)] # true lift of 1 point
print(f"{name:16} false positives {sum(d is not None for d in aa) / len(aa):.3f}"
f" power {sum(d is not None for d in ab) / len(ab):.3f}"
f" mean stop day {fmean(d for d in ab if d):.1f}")
# fixed horizon false positives 0.050 power 0.817 mean stop day 20.0
# peek daily false positives 0.241 power 0.892 mean stop day 7.4
# Pocock false positives 0.052 power 0.663 mean stop day 10.4
# O'Brien-Fleming false positives 0.051 power 0.798 mean stop day 13.7With no real effect, daily peeking declares a winner 24% of the time. Early in a test the z statistic swings widely, and every look is another lottery ticket; once it crosses 1.96, stopping locks in the false positive. The fixed-horizon test keeps its promised 5%.
Step 5: Use a sequential design if you must look
Sequential designs raise the bar at each look so the overall false-positive rate stays at 5%. Pocock uses one constant, stricter threshold (about 2.67 for 20 looks), which stops early often but loses power at the end. O’Brien-Fleming sets a very high bar early (2.12 × sqrt(20 / day), so a z of 9.5 on day 1) that relaxes to about 2.12 at the end. It keeps nearly all the power of the fixed test while still allowing a dramatic early result to stop the test. Experiment platforms also offer always-valid p-values (mSPRT), which allow continuous monitoring.
Worked scenario
On day 4 of a planned 14-day checkout test, the dashboard shows a lift of 2.1 points (21% relative), p = 0.01. The PM wants to stop, ship and announce a 21% lift.
Two things are wrong. First, the stop inflates the false-positive rate as shown above. Second, early winners exaggerate: tests that cross the line early do so partly because noise pushed them high.
rng = random.Random(8)
def stop_and_lift(p_treat, daily=750, days=20):
c1 = c2 = 0
for day in range(1, days + 1):
c1 += rng.binomialvariate(daily, 0.10)
c2 += rng.binomialvariate(daily, p_treat)
if abs(z_stat(c1, c2, day * daily)) > 1.96:
return day, (c2 - c1) / (day * daily)
return None, None
runs = [stop_and_lift(0.11) for _ in range(5000)]
early = [lift for day, lift in runs if day is not None and day <= 5]
print(len(early), round(fmean(early), 4)) # 2015 0.0236When the true lift is 1 point, runs that stopped by day 5 reported an average lift of 2.36 points, more than double the truth. The fixed response: keep the test running to its planned end, or, if the team agreed an O’Brien-Fleming plan in advance, check the day-4 boundary, which is 2.12 × sqrt(5) = 4.74, far above the observed z. Early stopping for harm on guardrail metrics (crashes, payment failures) is still allowed with its own pre-agreed threshold.
Common mistake
- “Run until significant.” That is the peeking problem with no upper limit; given enough looks, an A/A test will eventually cross 0.05.
- “Peeking is fine if we only look.” Looking is harmless; deciding based on the look is the problem. Dashboards can be open as long as the stop rule is fixed.
- Computing post-hoc power from the observed effect. It is a function of the p-value and adds no information; plan power before the test using the MDE.
- Sizing on all visitors when only users who reach the changed page can be affected. Dilution makes the effect smaller and the test longer.
- Stopping at a calendar date that cuts a week in half, or extending a test “a few more days” because it is close: both are forms of optional stopping.
Verify the behavior
Before trusting a sample size, simulate the planned test at that size: about 80% of runs with the true MDE should reject, and about 5% of A/A runs should.
vrng = random.Random(21)
n = math.ceil(sample_size(0.10, 0.11))
def rejects(p_treat):
c1 = vrng.binomialvariate(n, 0.10)
c2 = vrng.binomialvariate(n, p_treat)
return abs(z_stat(c1, c2, n)) > 1.96
power = fmean(rejects(0.11) for _ in range(20_000))
alpha = fmean(rejects(0.10) for _ in range(20_000))
print(round(power, 3), round(alpha, 3)) # 0.802 0.047
assert 0.78 < power < 0.82 and 0.04 < alpha < 0.06Follow-up questions
- How does CUPED change the sample size? It removes the variance explained by pre-experiment behaviour; with correlation ρ between the pre-period and in-test metric, the required sample shrinks by a factor of
1 - ρ². - What if the metric is revenue, not conversion? Use the general formula
n = 2 (z₁₋α/₂ + z₁₋β)² σ² / δ²with the metric’s standard deviation σ, estimated from historical data; heavy tails make σ large, so consider capping. - Why not always use one-sided tests to save users? They save about 20% of users, but they cannot detect harm, and switching sides after seeing data doubles the real alpha.
- How do you handle many variants? Each extra comparison against control adds false-positive chances; apply a correction such as Bonferroni or Holm, and size for the corrected alpha.
Interview exercise
Your baseline checkout conversion is 4%. The team wants to detect a 5% relative lift with alpha 0.05 and 80% power. About 6,000 users per day reach checkout. How long will the test take, and what would you propose if the answer is unacceptable?
Answer and reasoning
A 5% relative lift moves 4% to 4.2%, a difference of 0.2 points. sample_size(0.04, 0.042) gives 154,302 per arm (rounded up), so 308,604 users in total, or 52 days at 6,000 a day, which I would round up to 56 days (8 weeks). That is usually too long: seasonality drifts, cookies churn and the team waits two months. Levers, in order: accept a larger MDE if a 5% lift is not the smallest change worth shipping (10% relative needs roughly a quarter of the users), use variance reduction such as CUPED on prior purchase behaviour, choose a more sensitive metric earlier in the funnel as the primary metric if it is a credible proxy, or test on a higher-traffic surface. I would not shorten the test by peeking or by moving to a one-sided test after the fact.
Continue learning
- Practise the theory in the Statistics & Data Science chapter and test yourself with the data science MCQs.
- Revisit what the threshold means in p-values and hypothesis testing.
- Check that the test you size is the right one with t-test vs z-test vs chi-square.
- Read the NIST/SEMATECH handbook on sample sizes required and the Python statistics module documentation.