Ch. 32

Statistics & Data Science interview questions & answers

Data science interviews: probability, distributions, hypothesis testing and p-values, confidence intervals, A/B test design, regression, sampling bias and product metrics.

30 interview questions20 quiz questions8 notes
your progress0%

Notes in this chapter

Filter all notes →

30 Statistics & Data Science interview questions study by subtopic

30 questions
  1. 1.When would you report the median instead of the mean? Give an example where the two tell different stories.easy

    The mean uses every value, so a few extreme values pull it toward them. The median is the middle value, so it only cares about order and is robust to outliers. In a right-skewed distribution such as income, order values or session length, the mean sits above the median.

    Example: nine customers spend 20 dollars each and one spends 2,000. The mean basket is 218 dollars, but the median is 20, which is what a typical customer does. If you report the mean you will tell the product team their users spend ten times more than they do.

    I report the median (often with the 90th percentile) for skewed data and the mean when the total matters, for example revenue per user for forecasting, because the mean times the count gives the total. The mode is useful for categorical data or the most common basket size.

    What interviewers listen for
    • Mean is pulled by outliers, median is not
    • Right skew: mean above median
    • Mean matters when you need totals
    • Report percentiles for skewed metrics

    Likely follow-up: How would you summarise latency for an SLO? · What is a trimmed mean and when is it useful?

  2. 2.Why does the sample variance divide by n - 1 instead of n?easy

    Because the deviations are measured from the sample mean, not the true mean. The sample mean is, by construction, the point that minimises the sum of squared deviations for that sample, so squared distances to it are systematically a little smaller than distances to the population mean. Dividing by n therefore underestimates the population variance on average.

    Dividing by n - 1 (Bessel's correction) makes the estimator unbiased. The intuition is degrees of freedom: once you know the mean and n - 1 of the deviations, the last deviation is fixed, because they must sum to zero.

    In Python, statistics.variance and statistics.stdev use n - 1; pvariance and pstdev divide by n and are for a full population. The difference only matters for small samples. Note that the n - 1 standard deviation is still slightly biased, because the square root is not linear.

    What interviewers listen for
    • Deviations from the sample mean are too small
    • n - 1 makes the variance unbiased
    • Degrees of freedom intuition
    • variance vs pvariance in Python

    Likely follow-up: Is the sample standard deviation unbiased? · When would you deliberately use a biased estimator?

  3. 3.What is the difference between independent events and mutually exclusive events? Can two events be both?easy

    Mutually exclusive means the events cannot happen together: P(A and B) = 0. Rolling a 1 and rolling a 6 on one die are mutually exclusive. For them, P(A or B) = P(A) + P(B).

    Independent means knowing one happened does not change the probability of the other: P(A and B) = P(A) * P(B), or equivalently P(A | B) = P(A). Two separate coin flips are independent.

    They are almost opposites. If A and B are mutually exclusive and both have positive probability, then learning that A happened tells you B certainly did not, so they are strongly dependent. The only way to be both is for at least one event to have probability zero.

    In general, use the addition rule P(A or B) = P(A) + P(B) - P(A and B) and do not multiply probabilities unless you have a reason to believe independence.

    What interviewers listen for
    • Exclusive: P(A and B) = 0
    • Independent: P(A and B) = P(A)P(B)
    • Exclusive events with positive probability are dependent
    • General addition rule subtracts the overlap

    Likely follow-up: What is conditional independence? · Give an example where multiplying probabilities gives a badly wrong answer.

  4. 4.What is a p-value? Explain it to a product manager without using the word "significant".easy

    A p-value is the probability of seeing a result at least as extreme as the one we observed if there were really no effect (the null hypothesis). It measures how surprising the data is under the assumption that nothing changed.

    For a product manager: "If the new checkout did nothing at all, random noise alone would produce a lift this big or bigger about 3% of the time." A small p-value means the no-effect story fits the data poorly.

    What it is not: it is not the probability that the null is true, not the probability the result is a fluke, and not a measure of how big or important the effect is. A tiny, useless lift can have a tiny p-value with enough users. That is why I always report the effect size with a confidence interval next to the p-value, and decide the threshold (usually 0.05) before looking at data.

    What interviewers listen for
    • Probability of data this extreme assuming the null
    • Not P(null is true)
    • Says nothing about effect size
    • Report with a confidence interval

    Likely follow-up: Why is 0.05 the usual threshold? · What happens to p-values when sample size grows very large?

  5. 5.What are type I and type II errors? Which one is worse in an A/B test for a new checkout flow?easy

    A type I error (false positive) is rejecting a true null: you conclude the new checkout helps when it does not. Its rate is alpha, usually 0.05. A type II error (false negative) is failing to reject a false null: the checkout really helps but your test misses it. Its rate is beta, and power is 1 - beta, typically targeted at 0.8.

    Which is worse depends on costs. Shipping a neutral checkout (type I) costs engineering maintenance and can hide a small harm. Missing a real win (type II) costs the revenue you never collect. For a risky change, such as payments, I guard against type I strictly. For cheap, reversible UI tweaks, a missed win may be the bigger loss.

    The two trade off at a fixed sample size: lowering alpha raises beta. The only way to reduce both is more data or less noise.

    What interviewers listen for
    • Type I: false positive, rate alpha
    • Type II: false negative, rate beta
    • Power = 1 - beta
    • Trade off at fixed n; more data reduces both

    Likely follow-up: How would you choose alpha for a medical screening test? · What is a type S (sign) error?

  6. 6.Users who use feature X retain 30% better. The PM wants to push everyone to use it. What do you say?easy

    That is a correlation, and there are several reasons it may not be causal.

    • Selection / confounding: engaged users both discover feature X and retain better; engagement causes both.
    • Reverse causation: users who already plan to stay explore more features.
    • Survivorship: you can only use X if you stayed long enough to find it, so the comparison group includes people who left on day one.

    To find out, I would run an experiment: randomly prompt half of eligible users toward X and compare retention for the whole randomised groups (intention to treat), not just the users who clicked. If an experiment is impossible, I would compare similar users (matching on tenure and prior activity), look for a natural experiment such as a staggered rollout, and be explicit that the result is weaker evidence. Pushing everyone to X based on the correlation could produce no lift at all.

    What interviewers listen for
    • Confounding by engagement
    • Reverse causation and survivorship
    • Randomised experiment, intention to treat
    • Observational methods as weaker fallback

    Likely follow-up: What is a confounder versus a collider? · How would difference-in-differences help here?

  7. 7.What is the 68-95-99.7 rule, and what is a z-score?easy

    For a normal distribution, about 68.3% of values fall within one standard deviation of the mean, 95.4% within two and 99.7% within three. The exact two-sided 95% cut-off is 1.96 standard deviations, which is where the familiar 1.96 in confidence intervals comes from.

    A z-score says how many standard deviations a value is from the mean: z = (x - mean) / sd. It puts different scales on a common footing. If exam scores have mean 70 and sd 10, a score of 90 has z = 2, so it is higher than roughly 97.7% of scores.

    The caveat I always add: the percentages only hold for roughly normal data. For skewed data like revenue or latency, "three standard deviations" can be far from rare, and percentiles are a better description. In Python, statistics.NormalDist().cdf(2) returns about 0.977.

    What interviewers listen for
    • 68 / 95 / 99.7 within 1, 2, 3 sd
    • 1.96 for exact 95%
    • z = (x - mean) / sd
    • Only valid for roughly normal data

    Likely follow-up: How would you check whether data is normal? · What is Chebyshev's inequality and when would you use it?

  8. 8.Name the common types of sampling bias and give a product analytics example of each.easy

    Sampling bias means the sample differs systematically from the population you want to describe, and a bigger sample does not fix it.

    • Selection bias: an in-app survey only reaches active users, so satisfaction looks higher than it is across all customers.
    • Non-response bias: only angry users answer the support survey.
    • Survivorship bias: studying only accounts that are still active to learn "what successful customers do" ignores those that churned doing the same things.
    • Undercoverage: a mobile-only study misses desktop users.
    • Voluntary response: app store reviews skew to the extremes.
    • Time-window bias: measuring a week that contains a holiday or a launch.

    The cure is to define the target population first, sample randomly from it where possible, compare the sample's known attributes (platform, country, tenure) against the population, and reweight when they differ.

    What interviewers listen for
    • Bias is systematic; more data does not fix it
    • Selection, non-response, survivorship, undercoverage
    • Define the target population first
    • Check and reweight against known attributes

    Likely follow-up: How does stratified sampling help? · What is the difference between bias and variance of an estimate?

  9. 9.What does a 95% confidence interval mean? Is it correct to say there is a 95% chance the true value is inside it?easy

    A 95% confidence interval comes from a procedure that, if you repeated the study many times, would produce intervals containing the true value 95% of the time. The 95% describes the method, not one particular interval.

    Strictly, once you have computed a specific interval such as 2.1% to 3.4%, the true value is either in it or not; in the frequentist framework it is not a random quantity, so "95% chance it is inside" is not correct. A Bayesian credible interval does have that interpretation, given a prior.

    In practice I explain it as "a range of values compatible with the data at the 95% level". The width matters as much as the centre: a lift of +2% with an interval from -0.5% to +4.5% is inconclusive, while +2% from +1.6% to +2.4% is precise. An interval that excludes zero corresponds to a two-sided p-value below 0.05.

    What interviewers listen for
    • 95% is a property of the procedure
    • A computed interval either contains the value or not
    • Credible intervals are the Bayesian alternative
    • Width shows precision; excluding 0 matches p below 0.05

    Likely follow-up: How does the width change if you quadruple the sample? · When would you use a bootstrap interval?

  10. 10.What is an A/B test, and why do we randomise instead of comparing users before and after a launch?easy

    An A/B test randomly assigns units (usually users) to a control and one or more variants, runs them at the same time, and compares a pre-chosen metric. Randomisation makes the groups alike on everything, measured or not, so the only systematic difference is the change itself. That is what lets us read the difference causally.

    A before-and-after comparison mixes the change with everything else that moved in the same period: seasonality, marketing campaigns, a competitor's outage, a bug fix elsewhere, a news event. If conversion rose 4% the week we launched and also the week a holiday started, we cannot separate the two.

    The essentials of a good test are: a hypothesis and a primary metric chosen in advance, the right randomisation unit, a sample size from a power calculation, a fixed duration covering full weekly cycles, guardrail metrics, and a check that the split ratio came out as designed.

    What interviewers listen for
    • Random assignment balances unknown factors
    • Concurrent groups remove time effects
    • Pre-registered metric and sample size
    • Guardrails and split-ratio check

    Likely follow-up: What would you randomise on for a search ranking change? · When is an A/B test not possible?

  11. 11.A disease affects 1% of people. A test has 95% sensitivity and 95% specificity. If someone tests positive, what is the chance they have the disease?mid

    About 16%, not 95%. Bayes' theorem: P(D | +) = P(+ | D) P(D) / P(+).

    The easiest way to show it is with counts. Take 10,000 people. 100 have the disease and 95 of them test positive. 9,900 do not, and 5% of them, 495, also test positive. So 95 + 495 = 590 people test positive and only 95 are sick: 95 / 590 = 0.161.

    The low answer comes from the base rate. When the condition is rare, even a small false-positive rate applied to the huge healthy group swamps the true positives. This is base-rate neglect, and it is the reason screening programmes confirm positives with a second, independent test: after one positive the prior is 16%, and a second positive raises the posterior to about 78%.

    The same logic applies to fraud alerts and anomaly detectors on rare events.

    What interviewers listen for
    • Answer is about 16%
    • Natural frequencies make it clear
    • Base rate dominates when prevalence is low
    • Second test updates the new prior

    Likely follow-up: What specificity would you need for the posterior to be above 50%? · How does this relate to precision in classification?

  12. 12.State the central limit theorem. Why does it matter for A/B testing when revenue per user is heavily skewed?mid

    The CLT says that for independent draws from a distribution with finite variance, the distribution of the sample mean approaches a normal distribution as n grows, with mean equal to the population mean and standard deviation sigma / sqrt(n), the standard error. The data itself does not become normal; the mean of many draws does.

    That is why z-tests and t-tests work on conversion rates and revenue even though individual values are 0/1 or wildly skewed: we test differences in means, and with thousands of users per arm the means are close to normal.

    The caveats: "large n" depends on skew. Revenue with a few whales converges slowly, so the normal approximation can be poor at a few hundred users and p-values become unreliable. Fixes include more data, capping (winsorising) extreme values, a log transform with a clearly changed estimand, or bootstrap intervals. The CLT also fails for distributions without finite variance, such as Cauchy.

    What interviewers listen for
    • Sample mean becomes normal, not the data
    • Standard error = sigma / sqrt(n)
    • Justifies z- and t-tests on means
    • Heavy skew needs larger n or winsorising

    Likely follow-up: How would you check if n is large enough? · Does the CLT apply to the median?

  13. 13.When would you model a count with a binomial distribution and when with a Poisson? How are they related?mid

    Binomial(n, p) counts successes in a fixed number n of independent trials, each with the same probability p: how many of 1,000 visitors convert. Mean np, variance np(1 - p).

    Poisson(lambda) counts events in a fixed interval of time or space when events happen independently at a constant average rate: support tickets per hour, errors per million requests. Mean and variance are both lambda, and P(X = k) = e^-lambda lambda^k / k!.

    They are linked: a binomial with large n, small p and np = lambda is close to Poisson(lambda). That is why rare events among many trials are modelled as Poisson.

    In practice, check the variance. Real counts are often overdispersed (variance greater than the mean) because rates differ between users or hours; then a negative binomial model fits better, and Poisson-based confidence intervals are too narrow.

    What interviewers listen for
    • Binomial: fixed n trials, probability p
    • Poisson: events per interval, mean = variance
    • Poisson approximates binomial for large n, small p
    • Overdispersion suggests negative binomial

    Likely follow-up: If a server averages 3 errors an hour, what is the chance of none in an hour? · How would you test for overdispersion?

  14. 14.What does it mean that the exponential distribution is memoryless? Where does it come up?mid

    The exponential distribution models the waiting time between events in a Poisson process with rate lambda; its mean is 1 / lambda. Memoryless means P(T > s + t | T > s) = P(T > t): if you have already waited s minutes, the chance of waiting another t is the same as if you had just arrived. Past waiting tells you nothing.

    Example: requests arrive at 2 per second, so the gap averages 0.5 seconds, and P(gap > 1 s) = e^-2, about 0.135, whether or not the last request was a long time ago.

    It shows up in queueing models (M/M/1), time between failures for components with a constant hazard, and survival analysis as the simplest baseline. It is a strong assumption: human behaviour such as time to churn or time between purchases usually has a hazard that changes over time, so a Weibull or empirical survival curve fits better.

    What interviewers listen for
    • Waiting time between Poisson events
    • Mean 1 / lambda
    • Memoryless: past waiting irrelevant
    • Constant hazard is often unrealistic for users

    Likely follow-up: Which discrete distribution is also memoryless? · What is a hazard function?

  15. 15.What is statistical power, and what four things determine it?mid

    Power is the probability that the test rejects the null when a real effect of a given size exists, 1 - beta. A test with 50% power misses a true effect half the time, which is a coin flip after weeks of traffic.

    It is set by four linked quantities; fix three and the fourth follows:

    • Effect size: the minimum detectable effect you care about. Bigger effects are easier to see.
    • Sample size: power grows with n, but the effect you can detect shrinks only with sqrt(n), so halving the MDE needs four times the users.
    • Noise: the metric's variance. Lower-variance metrics, CUPED or a more sensitive proxy raise power.
    • Alpha: a stricter threshold lowers power.

    Underpowered tests have a second problem: when they do come out significant, the observed effect is likely exaggerated (the winner's curse), because only lucky overestimates crossed the line.

    What interviewers listen for
    • Power = P(reject | effect exists)
    • Effect size, n, variance, alpha
    • Half the MDE needs 4x the sample
    • Underpowered wins overstate the effect

    Likely follow-up: Why is post-hoc power from the observed effect not useful? · How does CUPED raise power?

  16. 16.When do you use a z-test, a t-test and a chi-square test?mid

    It depends on the data type and what you know about the variance.

    • z-test: comparing means or proportions when the standard error is known or estimated precisely because n is large. Two-proportion z-tests are the standard for conversion rates in A/B tests.
    • t-test: comparing means of a continuous metric when the variance is estimated from the sample. The t distribution has heavier tails for small samples and converges to the normal as degrees of freedom grow; by a few hundred per group the two give nearly the same answer. Use Welch's t-test by default because it does not assume equal variances, and a paired t-test for before-and-after on the same units.
    • chi-square test: categorical counts. The test of independence asks whether two categorical variables are related (variant by outcome table), and goodness of fit compares counts to expected proportions, such as a sample ratio mismatch check. A 2x2 chi-square test without continuity correction is equivalent to the two-sided two-proportion z-test: chi-square equals z squared.
    What interviewers listen for
    • z: known or large-n standard error, proportions
    • t: estimated variance, Welch by default
    • Paired t for repeated measures
    • Chi-square for counts; 2x2 equals z squared

    Likely follow-up: When would you use Mann-Whitney U instead? · What is the minimum expected count rule for chi-square?

  17. 17.Baseline conversion is 10%. How many users per arm do you need to detect a lift to 11% with alpha 0.05 and 80% power?mid

    About 14,750 per arm. For two proportions the per-arm sample size is roughly n = (z_alpha/2 + z_beta)^2 * (p1(1 - p1) + p2(1 - p2)) / (p2 - p1)^2.

    With z = 1.96 for a two-sided 0.05 and 0.84 for 80% power, the first factor is 7.85. The variance term is 0.09 + 0.0979 = 0.1879 and the squared difference is 0.0001, so n is 7.85 * 0.1879 / 0.0001, about 14,748. A handy shortcut is 16 * p(1 - p) / delta^2, which gives 14,400.

    Then I convert to time: if 3,000 eligible users arrive a day and we split 50/50, we need about 10 days, which I would round up to two full weeks to cover weekly cycles. If that is too long, the levers are a bigger MDE, a less noisy metric, variance reduction, or a higher-traffic surface.

    from statistics import NormalDist
    z = NormalDist().inv_cdf
    p1, p2 = 0.10, 0.11
    n = (z(0.975) + z(0.80)) ** 2 * (p1 * (1 - p1) + p2 * (1 - p2)) / (p2 - p1) ** 2
    print(round(n))  # 14748
    What interviewers listen for
    • Formula with z values and variance
    • Roughly 14,750 per arm
    • 16 p(1 - p) / delta squared shortcut
    • Convert to days and round to full weeks

    Likely follow-up: How does the answer change for a 5% relative lift? · What if you split traffic 90/10?

  18. 18.A stakeholder checks the A/B dashboard every day and wants to stop as soon as p drops below 0.05. What is wrong with that?mid

    Each look is another chance for noise to cross the threshold. The 5% false-positive rate of a fixed-horizon test assumes one analysis at a planned sample size. If you check repeatedly and stop at the first significant result, the overall false-positive rate climbs well above 5%: in a simulation of A/A tests with 20 daily looks it was about 24%.

    The p-value wanders like a random walk early on, and with enough looks it will dip below 0.05 even when nothing changed.

    Options, in order of preference:

    • Fix the sample size in advance and only make the decision at the end; dashboards can show data but nobody stops early.
    • Use a sequential design built for peeking: group-sequential boundaries (O'Brien-Fleming spends little alpha early) or always-valid p-values (mSPRT), which many experiment platforms offer.
    • Allow early stopping only for harm on guardrail metrics, with a strict threshold.
    What interviewers listen for
    • Repeated looks inflate false positives
    • Roughly 24% with 20 looks in an A/A simulation
    • Fixed horizon or sequential methods
    • Early stop only for guardrail harm

    Likely follow-up: How does O'Brien-Fleming spend alpha? · Is it fine to peek if you never stop early?

  19. 19.You test 20 metrics in one experiment and one shows p = 0.03. Should you celebrate?mid

    Not yet. With 20 independent tests at alpha 0.05 and no real effects, the chance that at least one comes out significant is 1 - 0.95^20, about 64%. One hit in twenty is what pure noise looks like.

    Ways to handle it:

    • Decide one primary metric in advance; the others are secondary or guardrails and are read as exploratory.
    • Bonferroni: test each at alpha / m, here 0.0025. It controls the family-wise error rate and is simple but conservative.
    • Holm: a step-down version of Bonferroni that is uniformly more powerful.
    • Benjamini-Hochberg: controls the false discovery rate, the expected share of false positives among the discoveries, which suits exploratory scans across many metrics or segments.

    The same issue appears with many variants, many segments ("it worked for iOS users in Brazil") and with re-running analyses in different ways until one works, which is p-hacking. A surprising secondary result is a hypothesis for the next test, not a launch decision.

    What interviewers listen for
    • 1 - 0.95^20 is about 64%
    • One pre-registered primary metric
    • Bonferroni / Holm control FWER
    • Benjamini-Hochberg controls FDR

    Likely follow-up: When is FDR a better target than FWER? · How do segment deep-dives create the same problem?

  20. 20.What is Simpson's paradox? Give an example from product analytics.mid

    Simpson's paradox is when a trend that holds in every subgroup reverses (or disappears) when the groups are combined, because the groups have different sizes in each comparison and a third variable is related to both.

    Example: a new onboarding flow has higher conversion than the old one on both mobile and desktop, yet lower conversion overall. That happens if the new flow got mostly mobile traffic, and mobile converts much worse than desktop. The aggregate mixes a platform effect with the flow effect.

    Which answer is right depends on the causal structure. Here platform is a confounder (it affects both which flow users saw and their conversion), so the within-platform comparison is the fair one, or better, a weighted average using the same platform mix for both flows. In a properly randomised experiment the mix is the same in both arms, so the paradox cannot appear between arms; it is a warning sign for observational comparisons and for experiments where randomisation broke.

    What interviewers listen for
    • Subgroup trend reverses in aggregate
    • Unequal mix plus a related third variable
    • Causal structure decides which view to trust
    • Randomisation equalises the mix

    Likely follow-up: When would the aggregate be the correct answer? · How does a mix shift explain a falling average with rising segments?

  21. 21.What are the assumptions of ordinary least squares linear regression, and what happens when each is violated?mid

    A useful mnemonic is LINE plus one.

    • Linearity: the mean of y is linear in the features. If not, coefficients are biased; check residuals against fitted values for curves and add transforms or interactions.
    • Independence of errors: violated by repeated measures or time series. Coefficients stay unbiased but standard errors are too small, so p-values are overconfident. Use clustered errors or time-series models.
    • Normality of errors: only matters for exact small-sample inference; with large n the CLT covers it. Check a Q-Q plot.
    • Equal variance (homoscedasticity): if the spread grows with the fitted value, estimates are still unbiased but standard errors are wrong. Use robust (HC) standard errors or a log transform.
    • No perfect multicollinearity: highly correlated features make individual coefficients unstable and hard to interpret, measured by VIF.

    The most important unstated one is exogeneity: errors are uncorrelated with the features. Omitted confounders break it and bias every coefficient.

    What interviewers listen for
    • Linearity, independence, normality, equal variance
    • Which violations bias coefficients vs standard errors
    • Residual plots and Q-Q plots
    • Exogeneity and omitted variables

    Likely follow-up: Does y itself need to be normally distributed? · How do you interpret a coefficient after a log transform?

  22. 22.Daily active users dropped 10% yesterday. Walk me through how you investigate.mid

    I work from cheapest explanation to most expensive.

    • Is it real? Check the data pipeline: late or partial loads, a changed event definition, a tracking SDK release, bot filtering, a timezone change. Compare against an independent source such as server logs.
    • How unusual is it? Compare with the same weekday last week and last year; holidays and weekly cycles explain many drops.
    • Where is it? Segment by platform, app version, country, acquisition channel, new versus returning users. A drop concentrated in Android on the latest version points to a release; one concentrated in one country points to an outage, a holiday or a payment provider.
    • Which step? Decompose: DAU = new users + returning users, and look at the funnel (app opens, logins, key action) to find the step that broke.
    • What changed? Releases, experiments ramped, marketing campaigns ended, competitor events, outages.

    I would report what is confirmed, what is ruled out and the next check, rather than one guess.

    What interviewers listen for
    • Rule out data and logging issues first
    • Compare against seasonality
    • Segment to localise the drop
    • Decompose the metric and the funnel
    • Correlate with releases and external events

    Likely follow-up: What if every segment dropped by the same amount? · How would you set up alerting so this is caught automatically?

  23. 23.How do you build and analyse a conversion funnel, and what pitfalls make funnels misleading?mid

    A funnel is an ordered set of steps, such as visit, product view, add to cart, checkout, purchase, counted per user within a time window. For each step I report the number of users, the step conversion (step n / step n - 1) and the overall conversion from the top. The biggest absolute drop is usually where to look first, but the step with the most fixable friction may be elsewhere.

    Pitfalls:

    • Inconsistent units: counting events at one step and users at another inflates rates.
    • Ordering: requiring strict order hides users who skip steps (direct links to checkout).
    • Window choice: a 1-hour window and a 7-day window give different answers; state it.
    • Mix shifts: a falling overall rate can come from more low-intent traffic at the top, not a broken step, so segment by channel and device.
    • Causal claims: a funnel shows where users leave, not why; confirm a fix with an experiment.

    I pair funnels with session recordings or qualitative research to explain the drop-off.

    What interviewers listen for
    • Per-user counts within a window
    • Step and overall conversion
    • Consistent units and ordering rules
    • Segment for mix shifts

    Likely follow-up: How would you compute this funnel in SQL? · Why can improving one step lower the next step's rate?

  24. 24.Write SQL that computes monthly cohort retention from an events(user_id, event_date) table, and explain the steps.mid

    Three steps.

    • Assign cohorts: each user's cohort is the month of their first event, MIN(event_date) grouped by user.
    • Compute the age of each activity: join events back to the cohort and calculate months since the first month. DISTINCT makes it one row per user per month, so a user with ten events in a month counts once.
    • Count and normalise: group by cohort and month number, then divide each count by the cohort's month-0 size. FIRST_VALUE(COUNT(*)) OVER (PARTITION BY cohort ORDER BY month_n) fetches that size without a second join.

    On the sample data, the January cohort has 3 users and retains 33% in month 1 and 67% in month 2, because a user can come back after skipping a month. If you want "still active through month n" (rolling retention) instead, count users whose last activity is at least month n. Always state which definition you are using, and exclude cohorts too young to have reached month n.

    WITH firsts AS (
      SELECT user_id, MIN(event_date) AS first_date
      FROM events GROUP BY user_id
    ),
    activity AS (
      SELECT DISTINCT e.user_id,
             strftime('%Y-%m', f.first_date) AS cohort,
             (CAST(strftime('%Y', e.event_date) AS INT) - CAST(strftime('%Y', f.first_date) AS INT)) * 12
             + CAST(strftime('%m', e.event_date) AS INT) - CAST(strftime('%m', f.first_date) AS INT) AS month_n
      FROM events e JOIN firsts f USING (user_id)
    )
    SELECT cohort, month_n, COUNT(*) AS active_users,
           ROUND(1.0 * COUNT(*) / FIRST_VALUE(COUNT(*)) OVER (
             PARTITION BY cohort ORDER BY month_n), 2) AS retention
    FROM activity
    GROUP BY cohort, month_n
    ORDER BY cohort, month_n;
    -- SQLite 3.53 date functions; in PostgreSQL use date_trunc('month', ...) and age()
    What interviewers listen for
    • Cohort = first activity month
    • One row per user per period
    • Divide by cohort size via window function
    • Classic vs rolling retention definitions

    Likely follow-up: How would you change this to weekly cohorts in PostgreSQL? · How do you handle cohorts that are not old enough yet?

  25. 25.A redesigned home feed shows +6% clicks in week one and +1% in week three. What is happening, and how should the decision be made?hard

    That pattern suggests a novelty effect: existing users click on anything new out of curiosity, and the lift fades as the change becomes familiar. The opposite, change aversion, starts negative and recovers. Either way the first week is not the long-run effect.

    To check, I would plot the treatment effect by day and by days since first exposure, and compare new users (who have no old habit) against existing ones. If new users show a stable +1%, that is probably the true long-term effect.

    For the decision I would look beyond clicks. Clicks are a proxy; the overall evaluation criterion might be sessions per user or retention. Guardrail metrics such as latency, crash rate, unsubscribes, revenue and support contacts must not degrade beyond a pre-set threshold. A +1% click lift that comes with a drop in time to first meaningful action may not be worth shipping. For long-term effects, a small long-running holdback group is the cleanest measure.

    What interviewers listen for
    • Novelty effect fades; change aversion recovers
    • Plot effect over exposure time
    • Compare new vs existing users
    • Guardrail metrics and a long-term holdback

    Likely follow-up: How would you pick the overall evaluation criterion? · What guardrails would you set for a payments change?

  26. 26.A 50/50 experiment ends with 50,000 users in control and 51,000 in treatment. Is that a problem?hard

    Probably yes. This is a sample ratio mismatch check, a chi-square goodness-of-fit test against the designed split. The expected count is 50,500 per arm, so chi-square is 2 * 500^2 / 50,500, about 9.9 with 1 degree of freedom, and the p-value is about 0.0017. A 1% imbalance looks small, but at this sample size chance rarely produces it.

    SRM means the groups are no longer comparable, so the metric results cannot be trusted until the cause is found. Common causes:

    • the treatment crashes or loads slowly, so fewer (or more) users fire the logging event that defines exposure;
    • bots or redirects that affect one arm;
    • assignment depending on something that changed mid-test;
    • filtering in the analysis that interacts with the treatment.

    Platforms usually alert at p below 0.001 to avoid false alarms. The fix is to find the cause, not to reweight the arms.

    What interviewers listen for
    • Chi-square goodness of fit vs designed split
    • Chi-square about 9.9, p about 0.0017
    • Results invalid until explained
    • Exposure logging and crashes are common causes

    Likely follow-up: Why not just downsample the bigger arm? · How would you debug SRM that only appears on one browser?

  27. 27.You are testing a new pricing algorithm for drivers in a ride-sharing marketplace. Why might a user-level A/B test give the wrong answer?hard

    User-level randomisation assumes no interference: one user's treatment does not affect another's outcome (SUTVA). Marketplaces break it. If treated riders get cheaper prices, they book more and take drivers away from control riders in the same city. Control looks worse than it would without the test, so the measured lift overstates the real effect; at full launch, everyone competes for the same drivers and the lift may vanish.

    Designs that reduce interference:

    • Cluster randomisation: randomise whole cities or geographic regions, accepting far fewer units and lower power.
    • Switchback tests: alternate the whole market between treatment and control over time slots (for example 30-minute blocks), with care for carryover between slots.
    • Two-sided randomisation or budget-split designs in ads marketplaces.

    Social networks have the same issue through friends. I would also validate a user-level result against a smaller cluster-level test before trusting it.

    What interviewers listen for
    • SUTVA / no interference assumption
    • Shared supply biases user-level lift
    • Cluster and switchback designs
    • Fewer units means lower power

    Likely follow-up: How do you analyse a switchback test? · How would you choose the switchback interval?

  28. 28.What is CUPED, and why does it let an experiment reach the same power with fewer users?hard

    CUPED (Controlled-experiment Using Pre-Experiment Data) adjusts each user's metric with a covariate measured before the experiment, usually the same metric in the prior weeks. The adjusted metric is Y_adj = Y - theta * (X - mean(X)) with theta = cov(X, Y) / var(X), the same coefficient as a regression of Y on X.

    Because X is measured before assignment, it is independent of the treatment, so the adjustment does not bias the treatment effect; it only removes the part of the variance explained by how users behaved already. The variance shrinks by a factor of 1 - rho^2, where rho is the correlation between X and Y. With rho = 0.7, variance drops by 49%, which roughly halves the required sample size.

    It works best for metrics with strong week-to-week persistence, like revenue or sessions per user. New users have no pre-period data, so they get a missing indicator or a different covariate. It is equivalent to ANCOVA, and regression adjustment with several covariates generalises it.

    What interviewers listen for
    • Pre-experiment covariate, usually same metric
    • theta = cov(X, Y) / var(X)
    • Variance falls by 1 - rho squared
    • Unbiased because X is pre-treatment

    Likely follow-up: Why must the covariate not be affected by treatment? · How do you handle users with no history?

  29. 29.You randomise by user but measure click-through rate as total clicks over total page views. Why is a standard two-proportion test wrong here?hard

    The two-proportion test treats each page view as an independent trial. But the randomisation unit is the user, and page views from the same user are correlated: a heavy clicker contributes many correlated views. The true variance of the CTR is larger than the binomial formula assumes, so the standard error is too small and false positives go well above 5%.

    The analysis unit must match the randomisation unit. Options:

    • Delta method: treat CTR as a ratio of two user-level means, mean clicks over mean views, and approximate its variance from the variances and covariance of those per-user sums. This is the standard approach in experimentation platforms.
    • Bootstrap by user: resample users, not views, and recompute the ratio.
    • Per-user metric: average each user's CTR, which changes the estimand to give every user equal weight.

    The same issue appears with revenue per session, latency per request, and cluster-randomised tests.

    What interviewers listen for
    • Views within a user are correlated
    • Binomial standard error is too small
    • Analysis unit must match randomisation unit
    • Delta method or user-level bootstrap

    Likely follow-up: Write down the delta method variance for a ratio of means. · When would you prefer the per-user average?

  30. 30.What is the expected number of fair coin flips until you see two heads in a row? Is it the same for heads followed by tails?hard

    For HH it is 6; for HT it is 4.

    For HH, set up states. Let E0 be the expected flips from scratch and E1 the expected flips after one head. From scratch, flip once: heads moves to state 1, tails stays at state 0, so E0 = 1 + 0.5 E1 + 0.5 E0. From state 1, flip once: heads finishes, tails sends you back to the start, so E1 = 1 + 0.5 * 0 + 0.5 E0. Solving gives E0 = 6 and E1 = 4.

    For HT, a tail after a head finishes, and a head after a head keeps you in state 1 rather than resetting, so E1 = 2 and E0 = 4.

    The asymmetry is the point of the question: with HH, a failure throws away your progress, while with HT a failed attempt still leaves you one step along. Interviewers want the state equations, a check by simulation, and the intuition about overlap.

    import random
    rng = random.Random(1)
    
    def flips_until(pattern):
        seq = ""
        while not seq.endswith(pattern):
            seq += rng.choice("HT")
        return len(seq)
    
    for p in ("HH", "HT"):
        print(p, sum(flips_until(p) for _ in range(200_000)) / 200_000)
    # HH 5.984845
    # HT 3.993295
    What interviewers listen for
    • Define states by progress
    • Write expectation equations
    • HH = 6, HT = 4
    • Overlap explains the asymmetry

    Likely follow-up: What is the expectation for three heads in a row? · Which pattern wins more often in a race between HH and TH?

Prefer multiple choice? All 20 Statistics & Data Science MCQs with answers →

esc