The statistics a data science or product analytics round actually tests, with the numbers worth knowing by heart. Code uses the Python 3.14 standard library (statistics, math, random).
Descriptive statistics
| Measure | Use when | Watch out |
|---|---|---|
| Mean | you need totals (revenue per user × users = revenue) | pulled by outliers; right skew puts it above the median |
| Median, p90, p99 | skewed data: latency, basket size, income | medians don’t add up across segments |
| Mode | categorical data, most common value | may not be unique |
| Variance / sd | spread around the mean | sample version divides by n - 1 (Bessel) |
| IQR (p75 - p25) | robust spread, outlier fences at 1.5 × IQR | quantile methods differ between tools |
from statistics import mean, median, stdev, pstdev, quantiles
data = [2, 4, 4, 4, 5, 5, 7, 9]
mean(data), median(data) # 5, 4.5
pstdev(data), round(stdev(data), 3) # 2.0 (divide by n), 2.138 (n - 1)
quantiles(range(1, 11), n=4) # [2.75, 5.5, 8.25] default method="exclusive"Probability rules
- Addition:
P(A or B) = P(A) + P(B) - P(A and B). Mutually exclusive: overlap is 0. - Multiplication:
P(A and B) = P(A) P(B | A). Independent:P(A and B) = P(A) P(B). - Exclusive events with positive probability are dependent, not independent.
- Complement for “at least one”:
1 - P(none). At least one six in 4 rolls:1 - (5/6)^4 = 0.518. - Expectation is linear even for dependent variables:
E[X + Y] = E[X] + E[Y]. Variances add only when independent (or uncorrelated). - Zero correlation does not mean independence:
y = x^2on symmetricxhas correlation 0.
Bayes and base rates
P(H | E) = P(E | H) P(H) / P(E), with P(E) = P(E | H) P(H) + P(E | not H) P(not H).
| Prevalence 1%, sensitivity 95%, specificity 95% | Count per 10,000 |
|---|---|
| Sick and positive (true positives) | 95 |
| Healthy and positive (false positives) | 495 |
| P(sick given positive) | 95 / 590 = 16.1% |
| After a second independent positive | 78.5% |
Odds form: posterior odds = prior odds × likelihood ratio, where LR+ = sensitivity / (1 - specificity). Here 1/99 × 19 = 19/99, so 0.161. Same maths as precision of a classifier on a rare class.
Distributions
| Distribution | Models | Mean | Variance |
|---|---|---|---|
| Bernoulli(p) | one yes/no trial (a conversion) | p | p(1 - p) |
| Binomial(n, p) | successes in n trials | np | np(1 - p) |
| Poisson(λ) | events per interval at a constant rate | λ | λ (variance > mean means overdispersion) |
| Geometric(p) | trials until the first success | 1/p | (1 - p)/p² |
| Exponential(λ) | waiting time between Poisson events, memoryless | 1/λ | 1/λ² |
| Normal(μ, σ) | sums and means of many small effects | μ | σ² |
| Uniform(a, b) | equally likely values; null p-values are Uniform(0, 1) | (a + b)/2 | (b - a)²/12 |
Normal: 68.3% / 95.4% / 99.7% within 1 / 2 / 3 sd. NormalDist().cdf(2) is 0.977.
Critical values to memorise
| Quantity | z |
|---|---|
| 90% two-sided, or 95% one-sided | 1.645 |
| 95% two-sided | 1.960 |
| 99% two-sided | 2.576 |
| 80% power | 0.842 |
| 90% power | 1.282 |
| Chi-square 1 df at 0.05 | 3.841 (= 1.96²) |
| t at 0.975, df = 9 / 30 | 2.262 / 2.042 (converges to 1.96) |
Sampling and the CLT
- Standard error of a mean:
σ / sqrt(n); of a proportion:sqrt(p(1 - p) / n). - CLT: the sample mean of independent draws with finite variance is approximately normal for large n. The data does not become normal.
- Skewed metrics (revenue with whales) need much larger n; the skewness of the mean shrinks like
skew / sqrt(n). Cauchy-like tails (no finite variance) never converge. - Four times the data halves the standard error and the CI width.
- Sampling bias (selection, non-response, survivorship, undercoverage) does not shrink with n.
Hypothesis testing
- p-value: probability of a result at least this extreme if the null were true. Not P(null is true), not effect size.
- Type I (false positive) rate α, usually 0.05. Type II (false negative) rate β. Power = 1 - β, usually 0.8.
- Power rises with effect size, sample size and α, and falls with variance.
- 20 independent tests at α = 0.05 with no real effect:
1 - 0.95^20 = 64%chance of at least one false positive. - Corrections: Bonferroni α/m and Holm control FWER; Benjamini-Hochberg controls FDR.
- A 95% CI that excludes 0 matches a two-sided p below 0.05.
Which test?
| Data | Question | Test |
|---|---|---|
| Two proportions, large n | conversion A vs B | two-proportion z-test |
| Continuous, two independent groups | mean A vs B | Welch’s t-test (default, unequal variances) |
| Continuous, same units twice | before vs after | paired t-test (t-test on the differences) |
| Counts in categories | is variant related to outcome? | chi-square test of independence |
| Counts vs expected shares | did the 50/50 split come out 50/50? | chi-square goodness of fit (SRM check) |
| Skewed or ordinal, small n | shift in distribution | Mann-Whitney U, or bootstrap |
| 3+ group means | any difference? | ANOVA, then post-hoc with correction |
A 2x2 chi-square (no continuity correction) equals the two-sided two-proportion z-test: χ² = z².
Confidence intervals
| Estimate | 95% interval |
|---|---|
| Mean, large n | x̄ ± 1.96 s / sqrt(n) (use t critical value for small n) |
| Proportion (Wald) | p̂ ± 1.96 sqrt(p̂(1 - p̂)/n); fails near 0 or 1 and for small n |
| Proportion (Wilson) | better coverage; use it for small n or rare events |
| Zero events in n trials | upper bound about 3 / n (rule of three) |
| Difference of proportions | (p̂₂ - p̂₁) ± 1.96 sqrt(p̂₁(1 - p̂₁)/n₁ + p̂₂(1 - p̂₂)/n₂) |
| Median, ratio, anything awkward | bootstrap: resample units, take the 2.5th and 97.5th percentiles |
“95%” describes the procedure: 95% of intervals built this way contain the true value.
A/B testing
- Per-arm sample size for proportions:
n = (z₁₋α/₂ + z₁₋β)² (p₁(1 - p₁) + p₂(1 - p₂)) / (p₂ - p₁)². Shortcut:16 p(1 - p) / δ². - 10% to 11% at α 0.05, power 0.8: 14,748 per arm. Halving the MDE needs 4× the users.
- Run whole weeks; fix the duration up front. Daily peeking and stopping at the first p < 0.05 pushes false positives to about 20 to 25% over 20 looks.
- Use sequential designs (O’Brien-Fleming, Pocock, always-valid p-values) if you must look early.
- SRM: chi-square on assignment counts; 50,000 vs 51,000 gives p ≈ 0.0017. Fix the cause, never reweight.
- CUPED:
Y - θ(X - mean(X)),θ = cov(X, Y)/var(X); variance falls by1 - ρ². - Ratio metrics (clicks / views) randomised by user need the delta method or a user-level bootstrap.
- Interference (marketplaces, social graphs) breaks SUTVA: cluster or switchback designs.
- Novelty effects fade; check the effect by days since exposure and keep a long-term holdback.
Correlation, causation, Simpson
- Confounder: affects both treatment and outcome; adjust for it. Mediator: on the causal path; adjusting hides the effect. Collider: caused by both; adjusting creates fake correlation.
- Simpson’s paradox: every segment improves but the total falls, because the segment mix differs between groups. Compare at a common mix (standardisation).
- Randomisation equalises the mix of everything, measured or not. Without it: matching, regression adjustment, difference-in-differences, natural experiments.
Linear regression (OLS)
| Assumption | Violation costs | Check / fix |
|---|---|---|
| Linearity | biased coefficients | residuals vs fitted; transforms, splines |
| Independent errors | standard errors too small | clustered SEs, time-series models |
| Constant variance | standard errors wrong (coefficients fine) | residual funnel; robust HC SEs, log y |
| Normal errors | small-n inference only | Q-Q plot; large n relies on the CLT |
| No perfect multicollinearity | unstable, uninterpretable coefficients | VIF = 1 / (1 - R²ⱼ), above 5 to 10 is a flag |
| Exogeneity (no omitted confounder) | every coefficient biased | domain knowledge, experiments, instruments |
from statistics import linear_regression, correlation
slope, intercept = linear_regression([1, 2, 3, 4, 5], [2, 4, 5, 4, 5])
round(slope, 2), round(intercept, 2) # 0.6, 2.2: slope = cov(x, y) / var(x)Anscombe’s quartet: four datasets with the same slope (0.5), intercept (3.0) and correlation (0.82) and completely different shapes. Always plot residuals.
Practise all of this in the Statistics & Data Science chapter.