“When would you use a t-test instead of a z-test? And where does chi-square fit?” is a staple of data science screens, often phrased as a scenario: “We changed the pricing page. Here are conversion counts, average order values and plan choices by variant. Which tests do you run?” Interviewers ask it because choosing the wrong test is one of the most common silent errors in analytics: the code runs, a p-value comes out, and nobody notices that the assumptions behind it do not hold. A good answer is not a memorised table; it is a short chain of reasoning from the data type and the experimental design to the test.
Before you start
You should know what a p-value and a standard error are, and the difference between a proportion (share of users who converted), a continuous metric (minutes, dollars) and categorical counts (how many users chose each plan). The code uses the Python 3.14 standard library. It has no t or chi-square distribution, so a short numerical integration supplies t-test p-values; in day-to-day work you would call scipy.stats.ttest_ind(a, b, equal_var=False) or chi2_contingency. Outputs in comments come from real runs, with the blocks run in order.
The short answer
Choose by data type and design. For two proportions with reasonable sample sizes, use a two-proportion z-test. For a continuous metric in two independent groups, use Welch’s t-test, which estimates the variance from the data and does not assume equal variances; for the same units measured twice, use a paired t-test on the differences. For categorical counts, use a chi-square test: of independence for a variant-by-category table, or goodness of fit against expected shares, such as checking the traffic split. With large samples the t and z distributions coincide, and a 2x2 chi-square is exactly the z-test squared.
How it works
Every one of these tests computes a ratio of signal to noise and asks how surprising it is under the null:
| Test | Statistic | Reference distribution |
|---|---|---|
| Two-proportion z | difference in rates / pooled standard error | standard normal |
| Welch t | difference in means / unpooled standard error | t with Welch-Satterthwaite df |
| Paired t | mean difference / (sd of differences / sqrt(n)) | t with n - 1 df |
| Chi-square | sum of (observed - expected)² / expected | chi-square with (rows - 1)(cols - 1) df |
The t distribution exists because a standard error estimated from a small sample is itself uncertain. It has heavier tails than the normal, and it converges to the normal as degrees of freedom grow: the 0.975 critical value is 2.262 at 9 df, 2.042 at 30 and 1.962 at 1,000. That is why “z-test versus t-test” stops mattering at A/B-test scale; what keeps mattering is choosing the right standard error.
Step-by-step walkthrough
Step 1: Identify the data type and the design
Before any code, answer two questions. What is one observation: a yes/no outcome, a number or a category? And are the groups independent (different users in each arm) or paired (the same users, servers or stores measured twice)? Those two answers pick the row of the table.
def choose_test(outcome, paired=False, groups=2):
if outcome == "binary" and groups == 2 and not paired:
return "two-proportion z-test (or 2x2 chi-square)"
if outcome == "continuous":
if paired:
return "paired t-test on the differences"
return "Welch t-test" if groups == 2 else "Welch ANOVA"
if outcome == "categorical":
return "chi-square test of independence"
return "think again: McNemar for paired binary, exact tests for tiny counts"
print(choose_test("continuous", paired=True)) # paired t-test on the differencesThen check the randomisation unit matches the analysis unit; a test on page views when you randomised users needs a different variance formula, whichever test you choose.
Step 2: Compare two proportions with a z-test, and see the chi-square link
Control converts 100 of 1,000 and treatment 130 of 1,000.
import math, random
from statistics import fmean, stdev, variance, NormalDist
c1, n1, c2, n2 = 100, 1000, 130, 1000
p = (c1 + c2) / (n1 + n2)
z = (c2 / n2 - c1 / n1) / math.sqrt(p * (1 - p) * (1 / n1 + 1 / n2))
p_value = 2 * (1 - NormalDist().cdf(abs(z)))
observed = [c1, n1 - c1, c2, n2 - c2]
expected = [n1 * p, n1 * (1 - p), n2 * p, n2 * (1 - p)]
chi2 = sum((o - e) ** 2 / e for o, e in zip(observed, expected))
print(round(z, 3), round(p_value, 4), round(chi2, 3), round(z * z, 3)) # 2.103 0.0355 4.422 4.422The Pearson chi-square on the 2x2 table equals z², and the chi-square p-value with 1 degree of freedom equals the two-sided z p-value. They are the same test in two notations. (Some libraries apply Yates’ continuity correction to 2x2 tables by default, which makes the chi-square slightly more conservative; know which one your tool runs.)
Step 3: Compare two means with Welch’s t-test
Minutes of session time for a small pilot, with visibly different spreads:
def t_pdf(x, df):
c = math.exp(math.lgamma((df + 1) / 2) - math.lgamma(df / 2)) / math.sqrt(df * math.pi)
return c * (1 + x * x / df) ** (-(df + 1) / 2)
def t_two_sided_p(t, df, steps=2000):
# Simpson's rule for the area between 0 and |t|; both tails are what is left
h = abs(t) / steps
s = t_pdf(0, df) + t_pdf(abs(t), df)
s += sum((4 if i % 2 else 2) * t_pdf(i * h, df) for i in range(1, steps))
return 1 - 2 * s * h / 3
def welch(a, b):
va, vb = variance(a) / len(a), variance(b) / len(b)
t = (fmean(b) - fmean(a)) / math.sqrt(va + vb)
df = (va + vb) ** 2 / (va ** 2 / (len(a) - 1) + vb ** 2 / (len(b) - 1))
return t, df, t_two_sided_p(t, df)
control = [4.1, 6.3, 3.8, 5.2, 7.9, 4.6, 5.5, 3.2, 6.8, 4.9, 5.1, 6.0]
treatment = [5.9, 8.4, 4.2, 7.7, 10.3, 6.1, 5.0, 9.6, 6.6, 8.8, 4.7, 7.2, 11.4, 6.9]
print(round(stdev(control), 3), round(stdev(treatment), 3)) # 1.326 2.156
t, df, p = welch(control, treatment)
print(f"t = {t:.3f}, df = {df:.1f}, p = {p:.4f}") # t = 2.978, df = 22.0, p = 0.0070Welch uses each group’s own variance and adjusts the degrees of freedom (22.0 here rather than 24) to reflect the uncertainty in those two estimates. When the variances happen to be equal it loses almost nothing compared with Student’s pooled test, which is why it is the default in R’s t.test and a sensible default everywhere.
Step 4: Test categorical counts with chi-square
Plan choice by variant on a pricing page is a 2x3 table. The test of independence compares each cell with the count expected if plan choice did not depend on variant: row total × column total / grand total.
table = {"control": [620, 280, 100], "treatment": [560, 300, 140]} # free, basic, pro
rows = list(table.values())
col_totals = [sum(col) for col in zip(*rows)]
grand = sum(col_totals)
chi2 = sum((o - sum(r) * ct / grand) ** 2 / (sum(r) * ct / grand)
for r in rows for o, ct in zip(r, col_totals))
print(round(chi2, 3), round(math.exp(-chi2 / 2), 4)) # 10.407 0.0055The degrees of freedom are (2 - 1) × (3 - 1) = 2, and for 2 df the chi-square tail probability has the closed form exp(-x / 2). The test says plan mix differs between variants (p = 0.0055) but not how; look at the cells, where treatment moved users from free to pro. Keep expected counts above about 5 per cell, or merge categories.
The goodness-of-fit version compares counts with fixed expected shares. Its most useful job in experimentation is the sample ratio mismatch (SRM) check on a planned 50/50 split:
counts = [50_000, 51_000]
exp_each = sum(counts) / 2
srm = sum((o - exp_each) ** 2 / exp_each for o in counts)
print(round(srm, 3), round(2 * (1 - NormalDist().cdf(math.sqrt(srm))), 4)) # 9.901 0.0017A 1% imbalance is very unlikely by chance at this size, so the assignment or logging is broken and the experiment’s results cannot be trusted until it is explained.
Worked scenario
An engineer optimises a serialisation library and measures p95 latency on the same 12 endpoints before and after. They run an independent two-sample t-test, get p = 0.76 and conclude the optimisation did nothing.
before = [212, 340, 158, 505, 276, 189, 422, 230, 367, 298, 145, 610]
after = [198, 321, 151, 470, 262, 180, 401, 214, 350, 281, 139, 572]
t, df, p = welch(before, after)
print(f"unpaired: t = {t:.3f}, p = {p:.4f}") # unpaired: t = -0.313, p = 0.7573
diffs = [a - b for a, b in zip(after, before)]
t_paired = fmean(diffs) / (stdev(diffs) / math.sqrt(len(diffs)))
print(fmean(diffs), f"paired: t = {t_paired:.3f}, p = {t_two_sided_p(t_paired, len(diffs) - 1):.6f}")
# -17.75 paired: t = -6.199, p = 0.000067The unpaired test compares the spread between endpoints (145 ms to 610 ms) with an improvement of about 18 ms, so the signal drowns. But every single endpoint got faster. The paired test works on the 12 differences, removing the endpoint-to-endpoint variation entirely, and the evidence is overwhelming. The fix is to match the test to the design: same units twice means paired. The same applies to before-and-after on the same users, matched stores, or A/B tests with pre-period covariates (where CUPED plays the role of pairing).
Common mistake
- “Use a z-test when n > 30, a t-test otherwise.” The real distinction is whether the standard error is known or estimated; at large n both give the same answer, and the more important choice is the right standard error.
- Student’s pooled t-test as the default. When the smaller group has the larger variance it can produce false positives far above 5%, as the simulation below shows.
- “The t-test needs normally distributed data.” It needs approximately normal sample means; with a few hundred users per arm that is usually satisfied, though heavy skew needs more.
- Chi-square on percentages or on non-independent rows, such as page views from the same users. Chi-square needs raw counts of independent units.
- Running a t-test on 0/1 conversion data and a z-test on the same data and treating them as two pieces of evidence. They are the same evidence.
Verify the behavior
Two checks: the t helper reproduces published critical values, and under a true null with unequal variances and sizes, Welch keeps its 5% promise while Student’s pooled test does not.
print(round(t_two_sided_p(2.262, 9), 4), round(t_two_sided_p(2.228, 10), 4)) # 0.05 0.05
def student_p(a, b):
na, nb = len(a), len(b)
sp = ((na - 1) * variance(a) + (nb - 1) * variance(b)) / (na + nb - 2)
t = (fmean(b) - fmean(a)) / math.sqrt(sp * (1 / na + 1 / nb))
return t_two_sided_p(t, na + nb - 2)
rng = random.Random(6)
sims, fp_welch, fp_student = 4000, 0, 0
for _ in range(sims):
small = [rng.gauss(0, 3) for _ in range(10)] # small group, large spread
large = [rng.gauss(0, 1) for _ in range(40)] # same mean, so the null is true
fp_welch += welch(small, large)[2] < 0.05
fp_student += student_p(small, large) < 0.05
print(fp_welch / sims, fp_student / sims) # 0.054 0.2525Student’s test calls a winner a quarter of the time when there is no difference, because pooling lets the 40 tight observations understate the noise in the 10 wide ones.
Follow-up questions
- When would you use Mann-Whitney U? For ordinal data or small, skewed samples where you care about a shift in distribution rather than the mean. It tests whether one group tends to have larger values, which is not the same hypothesis as equal means.
- What about three or more groups? One-way ANOVA (or Welch’s ANOVA) for means, then pairwise comparisons with a multiple-testing correction; chi-square handles more categories naturally.
- What if expected counts are small? Use Fisher’s exact test for 2x2 tables, or merge sparse categories.
- Is a paired t-test just a one-sample t-test? Yes: it is a one-sample t-test of whether the mean of the differences is zero.
Interview exercise
A pricing experiment reports three things per variant: conversion (purchased or not), revenue per purchaser, and which of four plans buyers chose. Users were randomised individually, about 20,000 per arm. Which test do you use for each, and what would you check first?
Answer and reasoning
First the SRM check: a chi-square goodness-of-fit test on the user counts per arm against the planned split, because any imbalance invalidates everything after it. Conversion is a proportion per randomised user, so a two-proportion z-test (equivalently a 2x2 chi-square). Revenue per purchaser is continuous, so Welch’s t-test, with two cautions: it is conditional on purchasing, so if the treatment changes who purchases, the groups are no longer comparable and revenue per user (including zeros) is the cleaner primary metric; and revenue is skewed, so check the sample is large enough or cap extreme values. Plan choice is a 2x4 table of counts, so a chi-square test of independence with 3 degrees of freedom, followed by looking at which cells drive it. With three tests, I would name the primary metric in advance and treat the others as secondary.
Continue learning
- Practise the theory in the Statistics & Data Science chapter and test yourself with the data science MCQs.
- Learn why sample means behave well enough for these tests in the central limit theorem.
- Size the test before you run it with A/B test sample size and peeking.
- Read the NIST/SEMATECH handbook on the two-sample t-test and the chi-square goodness-of-fit test.