Ch. 32 · Statistics & Data Science

T-Test vs Z-Test vs Chi-Square: Which Test to Use in Interviews

Pick the right test from the data type and design: two-proportion z-tests, Welch and paired t-tests, and chi-square, each run in Python.

~9 min readintermediateupdated Oct 6, 2026

“When would you use a t-test instead of a z-test? And where does chi-square fit?” is a staple of data science screens, often phrased as a scenario: “We changed the pricing page. Here are conversion counts, average order values and plan choices by variant. Which tests do you run?” Interviewers ask it because choosing the wrong test is one of the most common silent errors in analytics: the code runs, a p-value comes out, and nobody notices that the assumptions behind it do not hold. A good answer is not a memorised table; it is a short chain of reasoning from the data type and the experimental design to the test.

Before you start

You should know what a p-value and a standard error are, and the difference between a proportion (share of users who converted), a continuous metric (minutes, dollars) and categorical counts (how many users chose each plan). The code uses the Python 3.14 standard library. It has no t or chi-square distribution, so a short numerical integration supplies t-test p-values; in day-to-day work you would call scipy.stats.ttest_ind(a, b, equal_var=False) or chi2_contingency. Outputs in comments come from real runs, with the blocks run in order.

The short answer

Choose by data type and design. For two proportions with reasonable sample sizes, use a two-proportion z-test. For a continuous metric in two independent groups, use Welch’s t-test, which estimates the variance from the data and does not assume equal variances; for the same units measured twice, use a paired t-test on the differences. For categorical counts, use a chi-square test: of independence for a variant-by-category table, or goodness of fit against expected shares, such as checking the traffic split. With large samples the t and z distributions coincide, and a 2x2 chi-square is exactly the z-test squared.

How it works

Every one of these tests computes a ratio of signal to noise and asks how surprising it is under the null:

Test Statistic Reference distribution
Two-proportion z difference in rates / pooled standard error standard normal
Welch t difference in means / unpooled standard error t with Welch-Satterthwaite df
Paired t mean difference / (sd of differences / sqrt(n)) t with n - 1 df
Chi-square sum of (observed - expected)² / expected chi-square with (rows - 1)(cols - 1) df

The t distribution exists because a standard error estimated from a small sample is itself uncertain. It has heavier tails than the normal, and it converges to the normal as degrees of freedom grow: the 0.975 critical value is 2.262 at 9 df, 2.042 at 30 and 1.962 at 1,000. That is why “z-test versus t-test” stops mattering at A/B-test scale; what keeps mattering is choosing the right standard error.

Step-by-step walkthrough

Step 1: Identify the data type and the design

Before any code, answer two questions. What is one observation: a yes/no outcome, a number or a category? And are the groups independent (different users in each arm) or paired (the same users, servers or stores measured twice)? Those two answers pick the row of the table.

def choose_test(outcome, paired=False, groups=2):
    if outcome == "binary" and groups == 2 and not paired:
        return "two-proportion z-test (or 2x2 chi-square)"
    if outcome == "continuous":
        if paired:
            return "paired t-test on the differences"
        return "Welch t-test" if groups == 2 else "Welch ANOVA"
    if outcome == "categorical":
        return "chi-square test of independence"
    return "think again: McNemar for paired binary, exact tests for tiny counts"

print(choose_test("continuous", paired=True))   # paired t-test on the differences
python

Then check the randomisation unit matches the analysis unit; a test on page views when you randomised users needs a different variance formula, whichever test you choose.

Control converts 100 of 1,000 and treatment 130 of 1,000.

import math, random
from statistics import fmean, stdev, variance, NormalDist

c1, n1, c2, n2 = 100, 1000, 130, 1000
p = (c1 + c2) / (n1 + n2)
z = (c2 / n2 - c1 / n1) / math.sqrt(p * (1 - p) * (1 / n1 + 1 / n2))
p_value = 2 * (1 - NormalDist().cdf(abs(z)))

observed = [c1, n1 - c1, c2, n2 - c2]
expected = [n1 * p, n1 * (1 - p), n2 * p, n2 * (1 - p)]
chi2 = sum((o - e) ** 2 / e for o, e in zip(observed, expected))
print(round(z, 3), round(p_value, 4), round(chi2, 3), round(z * z, 3))   # 2.103 0.0355 4.422 4.422
python

The Pearson chi-square on the 2x2 table equals z², and the chi-square p-value with 1 degree of freedom equals the two-sided z p-value. They are the same test in two notations. (Some libraries apply Yates’ continuity correction to 2x2 tables by default, which makes the chi-square slightly more conservative; know which one your tool runs.)

Step 3: Compare two means with Welch’s t-test

Minutes of session time for a small pilot, with visibly different spreads:

def t_pdf(x, df):
    c = math.exp(math.lgamma((df + 1) / 2) - math.lgamma(df / 2)) / math.sqrt(df * math.pi)
    return c * (1 + x * x / df) ** (-(df + 1) / 2)

def t_two_sided_p(t, df, steps=2000):
    # Simpson's rule for the area between 0 and |t|; both tails are what is left
    h = abs(t) / steps
    s = t_pdf(0, df) + t_pdf(abs(t), df)
    s += sum((4 if i % 2 else 2) * t_pdf(i * h, df) for i in range(1, steps))
    return 1 - 2 * s * h / 3

def welch(a, b):
    va, vb = variance(a) / len(a), variance(b) / len(b)
    t = (fmean(b) - fmean(a)) / math.sqrt(va + vb)
    df = (va + vb) ** 2 / (va ** 2 / (len(a) - 1) + vb ** 2 / (len(b) - 1))
    return t, df, t_two_sided_p(t, df)

control = [4.1, 6.3, 3.8, 5.2, 7.9, 4.6, 5.5, 3.2, 6.8, 4.9, 5.1, 6.0]
treatment = [5.9, 8.4, 4.2, 7.7, 10.3, 6.1, 5.0, 9.6, 6.6, 8.8, 4.7, 7.2, 11.4, 6.9]
print(round(stdev(control), 3), round(stdev(treatment), 3))   # 1.326 2.156
t, df, p = welch(control, treatment)
print(f"t = {t:.3f}, df = {df:.1f}, p = {p:.4f}")   # t = 2.978, df = 22.0, p = 0.0070
python

Welch uses each group’s own variance and adjusts the degrees of freedom (22.0 here rather than 24) to reflect the uncertainty in those two estimates. When the variances happen to be equal it loses almost nothing compared with Student’s pooled test, which is why it is the default in R’s t.test and a sensible default everywhere.

Step 4: Test categorical counts with chi-square

Plan choice by variant on a pricing page is a 2x3 table. The test of independence compares each cell with the count expected if plan choice did not depend on variant: row total × column total / grand total.

table = {"control": [620, 280, 100], "treatment": [560, 300, 140]}   # free, basic, pro
rows = list(table.values())
col_totals = [sum(col) for col in zip(*rows)]
grand = sum(col_totals)
chi2 = sum((o - sum(r) * ct / grand) ** 2 / (sum(r) * ct / grand)
           for r in rows for o, ct in zip(r, col_totals))
print(round(chi2, 3), round(math.exp(-chi2 / 2), 4))   # 10.407 0.0055
python

The degrees of freedom are (2 - 1) × (3 - 1) = 2, and for 2 df the chi-square tail probability has the closed form exp(-x / 2). The test says plan mix differs between variants (p = 0.0055) but not how; look at the cells, where treatment moved users from free to pro. Keep expected counts above about 5 per cell, or merge categories.

The goodness-of-fit version compares counts with fixed expected shares. Its most useful job in experimentation is the sample ratio mismatch (SRM) check on a planned 50/50 split:

counts = [50_000, 51_000]
exp_each = sum(counts) / 2
srm = sum((o - exp_each) ** 2 / exp_each for o in counts)
print(round(srm, 3), round(2 * (1 - NormalDist().cdf(math.sqrt(srm))), 4))   # 9.901 0.0017
python

A 1% imbalance is very unlikely by chance at this size, so the assignment or logging is broken and the experiment’s results cannot be trusted until it is explained.

Worked scenario

An engineer optimises a serialisation library and measures p95 latency on the same 12 endpoints before and after. They run an independent two-sample t-test, get p = 0.76 and conclude the optimisation did nothing.

before = [212, 340, 158, 505, 276, 189, 422, 230, 367, 298, 145, 610]
after = [198, 321, 151, 470, 262, 180, 401, 214, 350, 281, 139, 572]
t, df, p = welch(before, after)
print(f"unpaired: t = {t:.3f}, p = {p:.4f}")   # unpaired: t = -0.313, p = 0.7573

diffs = [a - b for a, b in zip(after, before)]
t_paired = fmean(diffs) / (stdev(diffs) / math.sqrt(len(diffs)))
print(fmean(diffs), f"paired: t = {t_paired:.3f}, p = {t_two_sided_p(t_paired, len(diffs) - 1):.6f}")
# -17.75 paired: t = -6.199, p = 0.000067
python

The unpaired test compares the spread between endpoints (145 ms to 610 ms) with an improvement of about 18 ms, so the signal drowns. But every single endpoint got faster. The paired test works on the 12 differences, removing the endpoint-to-endpoint variation entirely, and the evidence is overwhelming. The fix is to match the test to the design: same units twice means paired. The same applies to before-and-after on the same users, matched stores, or A/B tests with pre-period covariates (where CUPED plays the role of pairing).

Common mistake

  • “Use a z-test when n > 30, a t-test otherwise.” The real distinction is whether the standard error is known or estimated; at large n both give the same answer, and the more important choice is the right standard error.
  • Student’s pooled t-test as the default. When the smaller group has the larger variance it can produce false positives far above 5%, as the simulation below shows.
  • “The t-test needs normally distributed data.” It needs approximately normal sample means; with a few hundred users per arm that is usually satisfied, though heavy skew needs more.
  • Chi-square on percentages or on non-independent rows, such as page views from the same users. Chi-square needs raw counts of independent units.
  • Running a t-test on 0/1 conversion data and a z-test on the same data and treating them as two pieces of evidence. They are the same evidence.

Verify the behavior

Two checks: the t helper reproduces published critical values, and under a true null with unequal variances and sizes, Welch keeps its 5% promise while Student’s pooled test does not.

print(round(t_two_sided_p(2.262, 9), 4), round(t_two_sided_p(2.228, 10), 4))   # 0.05 0.05

def student_p(a, b):
    na, nb = len(a), len(b)
    sp = ((na - 1) * variance(a) + (nb - 1) * variance(b)) / (na + nb - 2)
    t = (fmean(b) - fmean(a)) / math.sqrt(sp * (1 / na + 1 / nb))
    return t_two_sided_p(t, na + nb - 2)

rng = random.Random(6)
sims, fp_welch, fp_student = 4000, 0, 0
for _ in range(sims):
    small = [rng.gauss(0, 3) for _ in range(10)]   # small group, large spread
    large = [rng.gauss(0, 1) for _ in range(40)]   # same mean, so the null is true
    fp_welch += welch(small, large)[2] < 0.05
    fp_student += student_p(small, large) < 0.05
print(fp_welch / sims, fp_student / sims)   # 0.054 0.2525
python

Student’s test calls a winner a quarter of the time when there is no difference, because pooling lets the 40 tight observations understate the noise in the 10 wide ones.

Follow-up questions

  • When would you use Mann-Whitney U? For ordinal data or small, skewed samples where you care about a shift in distribution rather than the mean. It tests whether one group tends to have larger values, which is not the same hypothesis as equal means.
  • What about three or more groups? One-way ANOVA (or Welch’s ANOVA) for means, then pairwise comparisons with a multiple-testing correction; chi-square handles more categories naturally.
  • What if expected counts are small? Use Fisher’s exact test for 2x2 tables, or merge sparse categories.
  • Is a paired t-test just a one-sample t-test? Yes: it is a one-sample t-test of whether the mean of the differences is zero.

Interview exercise

A pricing experiment reports three things per variant: conversion (purchased or not), revenue per purchaser, and which of four plans buyers chose. Users were randomised individually, about 20,000 per arm. Which test do you use for each, and what would you check first?

Answer and reasoning

First the SRM check: a chi-square goodness-of-fit test on the user counts per arm against the planned split, because any imbalance invalidates everything after it. Conversion is a proportion per randomised user, so a two-proportion z-test (equivalently a 2x2 chi-square). Revenue per purchaser is continuous, so Welch’s t-test, with two cautions: it is conditional on purchasing, so if the treatment changes who purchases, the groups are no longer comparable and revenue per user (including zeros) is the cleaner primary metric; and revenue is skewed, so check the sample is large enough or cap extreme values. Plan choice is a 2x4 table of counts, so a chi-square test of independence with 3 degrees of freedom, followed by looking at which cells drive it. With three tests, I would name the primary metric in advance and treat the others as secondary.

Continue learning

More in Statistics & Data Science

esc