A/B Test Sample Size and the Peeking Problem: Interview Guide
Size an A/B test from baseline, MDE, alpha and power, turn it into days, and see by simulation why stopping at the first p below 0.05 fails.
Data science interviews: probability, distributions, hypothesis testing and p-values, confidence intervals, A/B test design, regression, sampling bias and product metrics.
Official reference: NIST Engineering Statistics Handbook
Size an A/B test from baseline, MDE, alpha and power, turn it into days, and see by simulation why stopping at the first p below 0.05 fails.
Solve the classic 1% disease test question three ways, see why rare events wreck precision, and avoid the base rate traps interviewers set.
What the central limit theorem really promises, shown by simulation on skewed revenue data, and where it quietly breaks in A/B tests.
What 95% confidence really means, how to build t, Wilson and bootstrap intervals in Python, and why the textbook proportion interval fails.
The OLS assumptions interviewers expect, which violations bias coefficients and which only break standard errors, with Python checks for each.
What a p-value actually measures, how to run a two-proportion test by hand in Python, and the misreadings that sink data science candidates.
Why a variant can win in every segment yet lose overall, how to standardise for mix, and how to tell confounders from causes in product data.
Pick the right test from the data type and design: two-proportion z-tests, Welch and paired t-tests, and chi-square, each run in Python.
The mean uses every value, so a few extreme values pull it toward them. The median is the middle value, so it only cares about order and is robust to outliers. In a right-skewed distribution such as income, order values or session length, the mean sits above the median.
Example: nine customers spend 20 dollars each and one spends 2,000. The mean basket is 218 dollars, but the median is 20, which is what a typical customer does. If you report the mean you will tell the product team their users spend ten times more than they do.
I report the median (often with the 90th percentile) for skewed data and the mean when the total matters, for example revenue per user for forecasting, because the mean times the count gives the total. The mode is useful for categorical data or the most common basket size.
Likely follow-up: How would you summarise latency for an SLO? · What is a trimmed mean and when is it useful?
n - 1 instead of n?easyBecause the deviations are measured from the sample mean, not the true mean. The sample mean is, by construction, the point that minimises the sum of squared deviations for that sample, so squared distances to it are systematically a little smaller than distances to the population mean. Dividing by n therefore underestimates the population variance on average.
Dividing by n - 1 (Bessel's correction) makes the estimator unbiased. The intuition is degrees of freedom: once you know the mean and n - 1 of the deviations, the last deviation is fixed, because they must sum to zero.
In Python, statistics.variance and statistics.stdev use n - 1; pvariance and pstdev divide by n and are for a full population. The difference only matters for small samples. Note that the n - 1 standard deviation is still slightly biased, because the square root is not linear.
Likely follow-up: Is the sample standard deviation unbiased? · When would you deliberately use a biased estimator?
Mutually exclusive means the events cannot happen together: P(A and B) = 0. Rolling a 1 and rolling a 6 on one die are mutually exclusive. For them, P(A or B) = P(A) + P(B).
Independent means knowing one happened does not change the probability of the other: P(A and B) = P(A) * P(B), or equivalently P(A | B) = P(A). Two separate coin flips are independent.
They are almost opposites. If A and B are mutually exclusive and both have positive probability, then learning that A happened tells you B certainly did not, so they are strongly dependent. The only way to be both is for at least one event to have probability zero.
In general, use the addition rule P(A or B) = P(A) + P(B) - P(A and B) and do not multiply probabilities unless you have a reason to believe independence.
Likely follow-up: What is conditional independence? · Give an example where multiplying probabilities gives a badly wrong answer.
A p-value is the probability of seeing a result at least as extreme as the one we observed if there were really no effect (the null hypothesis). It measures how surprising the data is under the assumption that nothing changed.
For a product manager: "If the new checkout did nothing at all, random noise alone would produce a lift this big or bigger about 3% of the time." A small p-value means the no-effect story fits the data poorly.
What it is not: it is not the probability that the null is true, not the probability the result is a fluke, and not a measure of how big or important the effect is. A tiny, useless lift can have a tiny p-value with enough users. That is why I always report the effect size with a confidence interval next to the p-value, and decide the threshold (usually 0.05) before looking at data.
Likely follow-up: Why is 0.05 the usual threshold? · What happens to p-values when sample size grows very large?
A type I error (false positive) is rejecting a true null: you conclude the new checkout helps when it does not. Its rate is alpha, usually 0.05. A type II error (false negative) is failing to reject a false null: the checkout really helps but your test misses it. Its rate is beta, and power is 1 - beta, typically targeted at 0.8.
Which is worse depends on costs. Shipping a neutral checkout (type I) costs engineering maintenance and can hide a small harm. Missing a real win (type II) costs the revenue you never collect. For a risky change, such as payments, I guard against type I strictly. For cheap, reversible UI tweaks, a missed win may be the bigger loss.
The two trade off at a fixed sample size: lowering alpha raises beta. The only way to reduce both is more data or less noise.
Likely follow-up: How would you choose alpha for a medical screening test? · What is a type S (sign) error?
That is a correlation, and there are several reasons it may not be causal.
To find out, I would run an experiment: randomly prompt half of eligible users toward X and compare retention for the whole randomised groups (intention to treat), not just the users who clicked. If an experiment is impossible, I would compare similar users (matching on tenure and prior activity), look for a natural experiment such as a staggered rollout, and be explicit that the result is weaker evidence. Pushing everyone to X based on the correlation could produce no lift at all.
Likely follow-up: What is a confounder versus a collider? · How would difference-in-differences help here?
For a normal distribution, about 68.3% of values fall within one standard deviation of the mean, 95.4% within two and 99.7% within three. The exact two-sided 95% cut-off is 1.96 standard deviations, which is where the familiar 1.96 in confidence intervals comes from.
A z-score says how many standard deviations a value is from the mean: z = (x - mean) / sd. It puts different scales on a common footing. If exam scores have mean 70 and sd 10, a score of 90 has z = 2, so it is higher than roughly 97.7% of scores.
The caveat I always add: the percentages only hold for roughly normal data. For skewed data like revenue or latency, "three standard deviations" can be far from rare, and percentiles are a better description. In Python, statistics.NormalDist().cdf(2) returns about 0.977.
Likely follow-up: How would you check whether data is normal? · What is Chebyshev's inequality and when would you use it?
Sampling bias means the sample differs systematically from the population you want to describe, and a bigger sample does not fix it.
The cure is to define the target population first, sample randomly from it where possible, compare the sample's known attributes (platform, country, tenure) against the population, and reweight when they differ.
Likely follow-up: How does stratified sampling help? · What is the difference between bias and variance of an estimate?
A 95% confidence interval comes from a procedure that, if you repeated the study many times, would produce intervals containing the true value 95% of the time. The 95% describes the method, not one particular interval.
Strictly, once you have computed a specific interval such as 2.1% to 3.4%, the true value is either in it or not; in the frequentist framework it is not a random quantity, so "95% chance it is inside" is not correct. A Bayesian credible interval does have that interpretation, given a prior.
In practice I explain it as "a range of values compatible with the data at the 95% level". The width matters as much as the centre: a lift of +2% with an interval from -0.5% to +4.5% is inconclusive, while +2% from +1.6% to +2.4% is precise. An interval that excludes zero corresponds to a two-sided p-value below 0.05.
Likely follow-up: How does the width change if you quadruple the sample? · When would you use a bootstrap interval?
An A/B test randomly assigns units (usually users) to a control and one or more variants, runs them at the same time, and compares a pre-chosen metric. Randomisation makes the groups alike on everything, measured or not, so the only systematic difference is the change itself. That is what lets us read the difference causally.
A before-and-after comparison mixes the change with everything else that moved in the same period: seasonality, marketing campaigns, a competitor's outage, a bug fix elsewhere, a news event. If conversion rose 4% the week we launched and also the week a holiday started, we cannot separate the two.
The essentials of a good test are: a hypothesis and a primary metric chosen in advance, the right randomisation unit, a sample size from a power calculation, a fixed duration covering full weekly cycles, guardrail metrics, and a check that the split ratio came out as designed.
Likely follow-up: What would you randomise on for a search ranking change? · When is an A/B test not possible?
About 16%, not 95%. Bayes' theorem: P(D | +) = P(+ | D) P(D) / P(+).
The easiest way to show it is with counts. Take 10,000 people. 100 have the disease and 95 of them test positive. 9,900 do not, and 5% of them, 495, also test positive. So 95 + 495 = 590 people test positive and only 95 are sick: 95 / 590 = 0.161.
The low answer comes from the base rate. When the condition is rare, even a small false-positive rate applied to the huge healthy group swamps the true positives. This is base-rate neglect, and it is the reason screening programmes confirm positives with a second, independent test: after one positive the prior is 16%, and a second positive raises the posterior to about 78%.
The same logic applies to fraud alerts and anomaly detectors on rare events.
Likely follow-up: What specificity would you need for the posterior to be above 50%? · How does this relate to precision in classification?
The CLT says that for independent draws from a distribution with finite variance, the distribution of the sample mean approaches a normal distribution as n grows, with mean equal to the population mean and standard deviation sigma / sqrt(n), the standard error. The data itself does not become normal; the mean of many draws does.
That is why z-tests and t-tests work on conversion rates and revenue even though individual values are 0/1 or wildly skewed: we test differences in means, and with thousands of users per arm the means are close to normal.
The caveats: "large n" depends on skew. Revenue with a few whales converges slowly, so the normal approximation can be poor at a few hundred users and p-values become unreliable. Fixes include more data, capping (winsorising) extreme values, a log transform with a clearly changed estimand, or bootstrap intervals. The CLT also fails for distributions without finite variance, such as Cauchy.
Likely follow-up: How would you check if n is large enough? · Does the CLT apply to the median?
Binomial(n, p) counts successes in a fixed number n of independent trials, each with the same probability p: how many of 1,000 visitors convert. Mean np, variance np(1 - p).
Poisson(lambda) counts events in a fixed interval of time or space when events happen independently at a constant average rate: support tickets per hour, errors per million requests. Mean and variance are both lambda, and P(X = k) = e^-lambda lambda^k / k!.
They are linked: a binomial with large n, small p and np = lambda is close to Poisson(lambda). That is why rare events among many trials are modelled as Poisson.
In practice, check the variance. Real counts are often overdispersed (variance greater than the mean) because rates differ between users or hours; then a negative binomial model fits better, and Poisson-based confidence intervals are too narrow.
Likely follow-up: If a server averages 3 errors an hour, what is the chance of none in an hour? · How would you test for overdispersion?
The exponential distribution models the waiting time between events in a Poisson process with rate lambda; its mean is 1 / lambda. Memoryless means P(T > s + t | T > s) = P(T > t): if you have already waited s minutes, the chance of waiting another t is the same as if you had just arrived. Past waiting tells you nothing.
Example: requests arrive at 2 per second, so the gap averages 0.5 seconds, and P(gap > 1 s) = e^-2, about 0.135, whether or not the last request was a long time ago.
It shows up in queueing models (M/M/1), time between failures for components with a constant hazard, and survival analysis as the simplest baseline. It is a strong assumption: human behaviour such as time to churn or time between purchases usually has a hazard that changes over time, so a Weibull or empirical survival curve fits better.
Likely follow-up: Which discrete distribution is also memoryless? · What is a hazard function?
Power is the probability that the test rejects the null when a real effect of a given size exists, 1 - beta. A test with 50% power misses a true effect half the time, which is a coin flip after weeks of traffic.
It is set by four linked quantities; fix three and the fourth follows:
sqrt(n), so halving the MDE needs four times the users.Underpowered tests have a second problem: when they do come out significant, the observed effect is likely exaggerated (the winner's curse), because only lucky overestimates crossed the line.
Likely follow-up: Why is post-hoc power from the observed effect not useful? · How does CUPED raise power?
It depends on the data type and what you know about the variance.
Likely follow-up: When would you use Mann-Whitney U instead? · What is the minimum expected count rule for chi-square?
About 14,750 per arm. For two proportions the per-arm sample size is roughly n = (z_alpha/2 + z_beta)^2 * (p1(1 - p1) + p2(1 - p2)) / (p2 - p1)^2.
With z = 1.96 for a two-sided 0.05 and 0.84 for 80% power, the first factor is 7.85. The variance term is 0.09 + 0.0979 = 0.1879 and the squared difference is 0.0001, so n is 7.85 * 0.1879 / 0.0001, about 14,748. A handy shortcut is 16 * p(1 - p) / delta^2, which gives 14,400.
Then I convert to time: if 3,000 eligible users arrive a day and we split 50/50, we need about 10 days, which I would round up to two full weeks to cover weekly cycles. If that is too long, the levers are a bigger MDE, a less noisy metric, variance reduction, or a higher-traffic surface.
from statistics import NormalDist
z = NormalDist().inv_cdf
p1, p2 = 0.10, 0.11
n = (z(0.975) + z(0.80)) ** 2 * (p1 * (1 - p1) + p2 * (1 - p2)) / (p2 - p1) ** 2
print(round(n)) # 14748Likely follow-up: How does the answer change for a 5% relative lift? · What if you split traffic 90/10?
Each look is another chance for noise to cross the threshold. The 5% false-positive rate of a fixed-horizon test assumes one analysis at a planned sample size. If you check repeatedly and stop at the first significant result, the overall false-positive rate climbs well above 5%: in a simulation of A/A tests with 20 daily looks it was about 24%.
The p-value wanders like a random walk early on, and with enough looks it will dip below 0.05 even when nothing changed.
Options, in order of preference:
Likely follow-up: How does O'Brien-Fleming spend alpha? · Is it fine to peek if you never stop early?
Not yet. With 20 independent tests at alpha 0.05 and no real effects, the chance that at least one comes out significant is 1 - 0.95^20, about 64%. One hit in twenty is what pure noise looks like.
Ways to handle it:
alpha / m, here 0.0025. It controls the family-wise error rate and is simple but conservative.The same issue appears with many variants, many segments ("it worked for iOS users in Brazil") and with re-running analyses in different ways until one works, which is p-hacking. A surprising secondary result is a hypothesis for the next test, not a launch decision.
Likely follow-up: When is FDR a better target than FWER? · How do segment deep-dives create the same problem?
Simpson's paradox is when a trend that holds in every subgroup reverses (or disappears) when the groups are combined, because the groups have different sizes in each comparison and a third variable is related to both.
Example: a new onboarding flow has higher conversion than the old one on both mobile and desktop, yet lower conversion overall. That happens if the new flow got mostly mobile traffic, and mobile converts much worse than desktop. The aggregate mixes a platform effect with the flow effect.
Which answer is right depends on the causal structure. Here platform is a confounder (it affects both which flow users saw and their conversion), so the within-platform comparison is the fair one, or better, a weighted average using the same platform mix for both flows. In a properly randomised experiment the mix is the same in both arms, so the paradox cannot appear between arms; it is a warning sign for observational comparisons and for experiments where randomisation broke.
Likely follow-up: When would the aggregate be the correct answer? · How does a mix shift explain a falling average with rising segments?
A useful mnemonic is LINE plus one.
The most important unstated one is exogeneity: errors are uncorrelated with the features. Omitted confounders break it and bias every coefficient.
Likely follow-up: Does y itself need to be normally distributed? · How do you interpret a coefficient after a log transform?
I work from cheapest explanation to most expensive.
I would report what is confirmed, what is ruled out and the next check, rather than one guess.
Likely follow-up: What if every segment dropped by the same amount? · How would you set up alerting so this is caught automatically?
A funnel is an ordered set of steps, such as visit, product view, add to cart, checkout, purchase, counted per user within a time window. For each step I report the number of users, the step conversion (step n / step n - 1) and the overall conversion from the top. The biggest absolute drop is usually where to look first, but the step with the most fixable friction may be elsewhere.
Pitfalls:
I pair funnels with session recordings or qualitative research to explain the drop-off.
Likely follow-up: How would you compute this funnel in SQL? · Why can improving one step lower the next step's rate?
events(user_id, event_date) table, and explain the steps.midThree steps.
MIN(event_date) grouped by user.DISTINCT makes it one row per user per month, so a user with ten events in a month counts once.FIRST_VALUE(COUNT(*)) OVER (PARTITION BY cohort ORDER BY month_n) fetches that size without a second join.On the sample data, the January cohort has 3 users and retains 33% in month 1 and 67% in month 2, because a user can come back after skipping a month. If you want "still active through month n" (rolling retention) instead, count users whose last activity is at least month n. Always state which definition you are using, and exclude cohorts too young to have reached month n.
WITH firsts AS (
SELECT user_id, MIN(event_date) AS first_date
FROM events GROUP BY user_id
),
activity AS (
SELECT DISTINCT e.user_id,
strftime('%Y-%m', f.first_date) AS cohort,
(CAST(strftime('%Y', e.event_date) AS INT) - CAST(strftime('%Y', f.first_date) AS INT)) * 12
+ CAST(strftime('%m', e.event_date) AS INT) - CAST(strftime('%m', f.first_date) AS INT) AS month_n
FROM events e JOIN firsts f USING (user_id)
)
SELECT cohort, month_n, COUNT(*) AS active_users,
ROUND(1.0 * COUNT(*) / FIRST_VALUE(COUNT(*)) OVER (
PARTITION BY cohort ORDER BY month_n), 2) AS retention
FROM activity
GROUP BY cohort, month_n
ORDER BY cohort, month_n;
-- SQLite 3.53 date functions; in PostgreSQL use date_trunc('month', ...) and age()Likely follow-up: How would you change this to weekly cohorts in PostgreSQL? · How do you handle cohorts that are not old enough yet?
That pattern suggests a novelty effect: existing users click on anything new out of curiosity, and the lift fades as the change becomes familiar. The opposite, change aversion, starts negative and recovers. Either way the first week is not the long-run effect.
To check, I would plot the treatment effect by day and by days since first exposure, and compare new users (who have no old habit) against existing ones. If new users show a stable +1%, that is probably the true long-term effect.
For the decision I would look beyond clicks. Clicks are a proxy; the overall evaluation criterion might be sessions per user or retention. Guardrail metrics such as latency, crash rate, unsubscribes, revenue and support contacts must not degrade beyond a pre-set threshold. A +1% click lift that comes with a drop in time to first meaningful action may not be worth shipping. For long-term effects, a small long-running holdback group is the cleanest measure.
Likely follow-up: How would you pick the overall evaluation criterion? · What guardrails would you set for a payments change?
Probably yes. This is a sample ratio mismatch check, a chi-square goodness-of-fit test against the designed split. The expected count is 50,500 per arm, so chi-square is 2 * 500^2 / 50,500, about 9.9 with 1 degree of freedom, and the p-value is about 0.0017. A 1% imbalance looks small, but at this sample size chance rarely produces it.
SRM means the groups are no longer comparable, so the metric results cannot be trusted until the cause is found. Common causes:
Platforms usually alert at p below 0.001 to avoid false alarms. The fix is to find the cause, not to reweight the arms.
Likely follow-up: Why not just downsample the bigger arm? · How would you debug SRM that only appears on one browser?
User-level randomisation assumes no interference: one user's treatment does not affect another's outcome (SUTVA). Marketplaces break it. If treated riders get cheaper prices, they book more and take drivers away from control riders in the same city. Control looks worse than it would without the test, so the measured lift overstates the real effect; at full launch, everyone competes for the same drivers and the lift may vanish.
Designs that reduce interference:
Social networks have the same issue through friends. I would also validate a user-level result against a smaller cluster-level test before trusting it.
Likely follow-up: How do you analyse a switchback test? · How would you choose the switchback interval?
CUPED (Controlled-experiment Using Pre-Experiment Data) adjusts each user's metric with a covariate measured before the experiment, usually the same metric in the prior weeks. The adjusted metric is Y_adj = Y - theta * (X - mean(X)) with theta = cov(X, Y) / var(X), the same coefficient as a regression of Y on X.
Because X is measured before assignment, it is independent of the treatment, so the adjustment does not bias the treatment effect; it only removes the part of the variance explained by how users behaved already. The variance shrinks by a factor of 1 - rho^2, where rho is the correlation between X and Y. With rho = 0.7, variance drops by 49%, which roughly halves the required sample size.
It works best for metrics with strong week-to-week persistence, like revenue or sessions per user. New users have no pre-period data, so they get a missing indicator or a different covariate. It is equivalent to ANCOVA, and regression adjustment with several covariates generalises it.
Likely follow-up: Why must the covariate not be affected by treatment? · How do you handle users with no history?
The two-proportion test treats each page view as an independent trial. But the randomisation unit is the user, and page views from the same user are correlated: a heavy clicker contributes many correlated views. The true variance of the CTR is larger than the binomial formula assumes, so the standard error is too small and false positives go well above 5%.
The analysis unit must match the randomisation unit. Options:
The same issue appears with revenue per session, latency per request, and cluster-randomised tests.
Likely follow-up: Write down the delta method variance for a ratio of means. · When would you prefer the per-user average?
For HH it is 6; for HT it is 4.
For HH, set up states. Let E0 be the expected flips from scratch and E1 the expected flips after one head. From scratch, flip once: heads moves to state 1, tails stays at state 0, so E0 = 1 + 0.5 E1 + 0.5 E0. From state 1, flip once: heads finishes, tails sends you back to the start, so E1 = 1 + 0.5 * 0 + 0.5 E0. Solving gives E0 = 6 and E1 = 4.
For HT, a tail after a head finishes, and a head after a head keeps you in state 1 rather than resetting, so E1 = 2 and E0 = 4.
The asymmetry is the point of the question: with HH, a failure throws away your progress, while with HT a failed attempt still leaves you one step along. Interviewers want the state equations, a check by simulation, and the intuition about overlap.
import random
rng = random.Random(1)
def flips_until(pattern):
seq = ""
while not seq.endswith(pattern):
seq += rng.choice("HT")
return len(seq)
for p in ("HH", "HT"):
print(p, sum(flips_until(p) for _ in range(200_000)) / 200_000)
# HH 5.984845
# HT 3.993295Likely follow-up: What is the expectation for three heads in a row? · Which pattern wins more often in a race between HH and TH?
No questions match that filter.
Prefer multiple choice? All 20 Statistics & Data Science MCQs with answers →