Statistics & Data Science · cheat sheet

Statistics & Data Science

Probability rules, Bayes, distributions, CLT, p-values, confidence intervals, test selection, A/B sample size, peeking and regression assumptions on one page.

The statistics a data science or product analytics round actually tests, with the numbers worth knowing by heart. Code uses the Python 3.14 standard library (statistics, math, random).

Descriptive statistics

Measure Use when Watch out
Mean you need totals (revenue per user × users = revenue) pulled by outliers; right skew puts it above the median
Median, p90, p99 skewed data: latency, basket size, income medians don’t add up across segments
Mode categorical data, most common value may not be unique
Variance / sd spread around the mean sample version divides by n - 1 (Bessel)
IQR (p75 - p25) robust spread, outlier fences at 1.5 × IQR quantile methods differ between tools
from statistics import mean, median, stdev, pstdev, quantiles
data = [2, 4, 4, 4, 5, 5, 7, 9]
mean(data), median(data)          # 5, 4.5
pstdev(data), round(stdev(data), 3)   # 2.0 (divide by n), 2.138 (n - 1)
quantiles(range(1, 11), n=4)      # [2.75, 5.5, 8.25]  default method="exclusive"
python

Probability rules

  • Addition: P(A or B) = P(A) + P(B) - P(A and B). Mutually exclusive: overlap is 0.
  • Multiplication: P(A and B) = P(A) P(B | A). Independent: P(A and B) = P(A) P(B).
  • Exclusive events with positive probability are dependent, not independent.
  • Complement for “at least one”: 1 - P(none). At least one six in 4 rolls: 1 - (5/6)^4 = 0.518.
  • Expectation is linear even for dependent variables: E[X + Y] = E[X] + E[Y]. Variances add only when independent (or uncorrelated).
  • Zero correlation does not mean independence: y = x^2 on symmetric x has correlation 0.

Bayes and base rates

P(H | E) = P(E | H) P(H) / P(E), with P(E) = P(E | H) P(H) + P(E | not H) P(not H).

Prevalence 1%, sensitivity 95%, specificity 95% Count per 10,000
Sick and positive (true positives) 95
Healthy and positive (false positives) 495
P(sick given positive) 95 / 590 = 16.1%
After a second independent positive 78.5%

Odds form: posterior odds = prior odds × likelihood ratio, where LR+ = sensitivity / (1 - specificity). Here 1/99 × 19 = 19/99, so 0.161. Same maths as precision of a classifier on a rare class.

Distributions

Distribution Models Mean Variance
Bernoulli(p) one yes/no trial (a conversion) p p(1 - p)
Binomial(n, p) successes in n trials np np(1 - p)
Poisson(λ) events per interval at a constant rate λ λ (variance > mean means overdispersion)
Geometric(p) trials until the first success 1/p (1 - p)/p²
Exponential(λ) waiting time between Poisson events, memoryless 1/λ 1/λ²
Normal(μ, σ) sums and means of many small effects μ σ²
Uniform(a, b) equally likely values; null p-values are Uniform(0, 1) (a + b)/2 (b - a)²/12

Normal: 68.3% / 95.4% / 99.7% within 1 / 2 / 3 sd. NormalDist().cdf(2) is 0.977.

Critical values to memorise

Quantity z
90% two-sided, or 95% one-sided 1.645
95% two-sided 1.960
99% two-sided 2.576
80% power 0.842
90% power 1.282
Chi-square 1 df at 0.05 3.841 (= 1.96²)
t at 0.975, df = 9 / 30 2.262 / 2.042 (converges to 1.96)

Sampling and the CLT

  • Standard error of a mean: σ / sqrt(n); of a proportion: sqrt(p(1 - p) / n).
  • CLT: the sample mean of independent draws with finite variance is approximately normal for large n. The data does not become normal.
  • Skewed metrics (revenue with whales) need much larger n; the skewness of the mean shrinks like skew / sqrt(n). Cauchy-like tails (no finite variance) never converge.
  • Four times the data halves the standard error and the CI width.
  • Sampling bias (selection, non-response, survivorship, undercoverage) does not shrink with n.

Hypothesis testing

  • p-value: probability of a result at least this extreme if the null were true. Not P(null is true), not effect size.
  • Type I (false positive) rate α, usually 0.05. Type II (false negative) rate β. Power = 1 - β, usually 0.8.
  • Power rises with effect size, sample size and α, and falls with variance.
  • 20 independent tests at α = 0.05 with no real effect: 1 - 0.95^20 = 64% chance of at least one false positive.
  • Corrections: Bonferroni α/m and Holm control FWER; Benjamini-Hochberg controls FDR.
  • A 95% CI that excludes 0 matches a two-sided p below 0.05.

Which test?

Data Question Test
Two proportions, large n conversion A vs B two-proportion z-test
Continuous, two independent groups mean A vs B Welch’s t-test (default, unequal variances)
Continuous, same units twice before vs after paired t-test (t-test on the differences)
Counts in categories is variant related to outcome? chi-square test of independence
Counts vs expected shares did the 50/50 split come out 50/50? chi-square goodness of fit (SRM check)
Skewed or ordinal, small n shift in distribution Mann-Whitney U, or bootstrap
3+ group means any difference? ANOVA, then post-hoc with correction

A 2x2 chi-square (no continuity correction) equals the two-sided two-proportion z-test: χ² = z².

Confidence intervals

Estimate 95% interval
Mean, large n x̄ ± 1.96 s / sqrt(n) (use t critical value for small n)
Proportion (Wald) p̂ ± 1.96 sqrt(p̂(1 - p̂)/n); fails near 0 or 1 and for small n
Proportion (Wilson) better coverage; use it for small n or rare events
Zero events in n trials upper bound about 3 / n (rule of three)
Difference of proportions (p̂₂ - p̂₁) ± 1.96 sqrt(p̂₁(1 - p̂₁)/n₁ + p̂₂(1 - p̂₂)/n₂)
Median, ratio, anything awkward bootstrap: resample units, take the 2.5th and 97.5th percentiles

“95%” describes the procedure: 95% of intervals built this way contain the true value.

A/B testing

  • Per-arm sample size for proportions: n = (z₁₋α/₂ + z₁₋β)² (p₁(1 - p₁) + p₂(1 - p₂)) / (p₂ - p₁)². Shortcut: 16 p(1 - p) / δ².
  • 10% to 11% at α 0.05, power 0.8: 14,748 per arm. Halving the MDE needs 4× the users.
  • Run whole weeks; fix the duration up front. Daily peeking and stopping at the first p < 0.05 pushes false positives to about 20 to 25% over 20 looks.
  • Use sequential designs (O’Brien-Fleming, Pocock, always-valid p-values) if you must look early.
  • SRM: chi-square on assignment counts; 50,000 vs 51,000 gives p ≈ 0.0017. Fix the cause, never reweight.
  • CUPED: Y - θ(X - mean(X)), θ = cov(X, Y)/var(X); variance falls by 1 - ρ².
  • Ratio metrics (clicks / views) randomised by user need the delta method or a user-level bootstrap.
  • Interference (marketplaces, social graphs) breaks SUTVA: cluster or switchback designs.
  • Novelty effects fade; check the effect by days since exposure and keep a long-term holdback.

Correlation, causation, Simpson

  • Confounder: affects both treatment and outcome; adjust for it. Mediator: on the causal path; adjusting hides the effect. Collider: caused by both; adjusting creates fake correlation.
  • Simpson’s paradox: every segment improves but the total falls, because the segment mix differs between groups. Compare at a common mix (standardisation).
  • Randomisation equalises the mix of everything, measured or not. Without it: matching, regression adjustment, difference-in-differences, natural experiments.

Linear regression (OLS)

Assumption Violation costs Check / fix
Linearity biased coefficients residuals vs fitted; transforms, splines
Independent errors standard errors too small clustered SEs, time-series models
Constant variance standard errors wrong (coefficients fine) residual funnel; robust HC SEs, log y
Normal errors small-n inference only Q-Q plot; large n relies on the CLT
No perfect multicollinearity unstable, uninterpretable coefficients VIF = 1 / (1 - R²ⱼ), above 5 to 10 is a flag
Exogeneity (no omitted confounder) every coefficient biased domain knowledge, experiments, instruments
from statistics import linear_regression, correlation
slope, intercept = linear_regression([1, 2, 3, 4, 5], [2, 4, 5, 4, 5])
round(slope, 2), round(intercept, 2)   # 0.6, 2.2: slope = cov(x, y) / var(x)
python

Anscombe’s quartet: four datasets with the same slope (0.5), intercept (3.0) and correlation (0.82) and completely different shapes. Always plot residuals.

Practise all of this in the Statistics & Data Science chapter.

esc