Ch. 32 · Statistics & Data Science

P-Values and Hypothesis Testing: Data Science Interview Guide

What a p-value actually measures, how to run a two-proportion test by hand in Python, and the misreadings that sink data science candidates.

~8 min readbeginnerupdated Oct 6, 2026

“What is a p-value?” sounds like a warm-up, but it is one of the most reliable filters in a data science interview. Most candidates can recite a definition; far fewer can say what a p-value is not, compute one from raw counts, or explain to a product manager why p = 0.04 on a tiny lift is not a reason to ship. Interviewers ask it because every A/B test readout, every “is this metric change real?” question and every regression table you will produce depends on reading p-values correctly, and misreading them is how teams ship changes that do nothing.

Before you start

You need the idea of a sample versus a population, a mean and a proportion, and a rough feel for the normal distribution (about 95% of values within two standard deviations). The code uses only the Python 3.14 standard library: math, random and statistics.NormalDist. Every printed value in the comments came from running the code with the seeds shown. No calculus is needed.

The short answer

A p-value is the probability of seeing a result at least as extreme as the one observed, assuming the null hypothesis is true. A small p-value says the data would be surprising if there were no effect, so the no-effect explanation fits poorly. It is not the probability that the null is true, not the probability that the result is a fluke, and it says nothing about how large or valuable the effect is, so I always report it next to the effect size and a confidence interval.

How it works

Every hypothesis test has the same skeleton. You state a null hypothesis (usually “no difference”), choose a test statistic that grows as the data moves away from the null, work out how that statistic is distributed if the null were true, and then ask where your observed value falls in that distribution. The p-value is the area in the tail beyond your observed value. “Two-sided” means you count both tails, because a large drop would be as surprising as a large lift.

For conversion rates, the standard statistic is the two-proportion z: the difference in rates divided by its standard error under the null. Under the null both arms share one true rate, so the standard error uses the pooled rate.

import math
from statistics import NormalDist

def two_prop_z(c1, n1, c2, n2):
    p1, p2 = c1 / n1, c2 / n2
    pooled = (c1 + c2) / (n1 + n2)
    se = math.sqrt(pooled * (1 - pooled) * (1 / n1 + 1 / n2))
    z = (p2 - p1) / se
    p_value = 2 * (1 - NormalDist().cdf(abs(z)))
    return z, p_value

z, p = two_prop_z(1000, 10_000, 1090, 10_000)
print(f"z = {z:.3f}, p = {p:.4f}")   # z = 2.080, p = 0.0375
python

Control converted 1,000 of 10,000 users (10.0%) and treatment 1,090 of 10,000 (10.9%). The observed difference is about 2.08 standard errors from zero, and the probability of landing at least that far out in either direction, if the two flows were truly identical, is 3.75%.

The key mental shift is that the p-value is computed in an imaginary world where the null is true. It conditions on the null. That is exactly why it cannot tell you the probability of the null: that would require flipping the conditional, which needs a prior, which is Bayes’ theorem territory.

Step-by-step walkthrough

Step 1: State the hypotheses and the threshold before you look

Write the null and alternative down, choose alpha (usually 0.05) and decide one-sided or two-sided before seeing results. Choosing after seeing the data turns a 5% false-positive rate into something larger, because you are picking whichever framing makes the result look strongest.

# H0: p_treatment == p_control      H1: p_treatment != p_control (two-sided)
ALPHA = 0.05
c1, n1 = 1000, 10_000   # control conversions, users
c2, n2 = 1090, 10_000   # treatment conversions, users
print(c1 / n1, c2 / n2)   # 0.1 0.109
python

Two-sided is the honest default for product tests, because a change can hurt as well as help and you want to detect both.

Step 2: Compute the test statistic

The z statistic measures how many standard errors separate the two rates. A difference of 0.9 percentage points means nothing on its own; with 100 users per arm it is noise, with 10,000 per arm it is borderline, with a million per arm it is overwhelming. Dividing by the standard error turns the raw difference into a scale-free measure of surprise.

z, p = two_prop_z(c1, n1, c2, n2)
print(round(z, 3))   # 2.08
python

Step 3: Turn the statistic into a p-value, and check it by simulation

The normal approximation gives 0.0375. To see what the p-value really means, simulate the null world directly: give both arms the same pooled conversion rate, rerun the experiment many times, and count how often noise alone produces a gap at least as large as the observed 0.9 points.

import random

rng = random.Random(7)
pooled = 2090 / 20_000
observed = 1090 / 10_000 - 1000 / 10_000
sims = 20_000
extreme = 0
for _ in range(sims):
    a = rng.binomialvariate(10_000, pooled)
    b = rng.binomialvariate(10_000, pooled)
    if abs(b - a) / 10_000 >= observed - 1e-12:
        extreme += 1
print(f"simulated p = {extreme / sims:.4f}")   # simulated p = 0.0360
python

The simulated answer, 3.6%, agrees with the formula within simulation noise. This is the definition made concrete: in a world with no effect, about 1 run in 27 produces a gap this big. (The 1e-12 guards against floating-point rounding when a simulated gap equals the observed one exactly.)

Step 4: Report the effect size with a confidence interval

The p-value answers “is this distinguishable from zero?” The decision needs “how big is it, and how sure are we?” The 95% confidence interval for the difference uses the unpooled standard error, because it no longer assumes the null.

p1, p2 = 0.10, 0.109
se = math.sqrt(p1 * (1 - p1) / 10_000 + p2 * (1 - p2) / 10_000)
print(f"diff = {p2 - p1:.4f}, 95% CI = ({p2 - p1 - 1.96 * se:.4f}, {p2 - p1 + 1.96 * se:.4f})")
# diff = 0.0090, 95% CI = (0.0005, 0.0175)
python

The lift is plausibly anywhere from 0.05 to 1.75 percentage points. That is a much more useful sentence for a product manager than “p = 0.0375”: the change probably helps, but the data is compatible with an effect too small to matter.

Worked scenario

A team tests a new checkout button. The readout says treatment converted 1,080 of 10,000 against control’s 1,000 of 10,000, p = 0.064. The analyst writes “no significant difference, the new button has no effect” and the change is abandoned.

z, p = two_prop_z(1000, 10_000, 1080, 10_000)
print(round(z, 3), round(p, 4))   # 1.853 0.0639
se = math.sqrt(0.10 * 0.90 / 10_000 + 0.108 * 0.892 / 10_000)
print(round(0.008 - 1.96 * se, 4), round(0.008 + 1.96 * se, 4))   # -0.0005 0.0165
python

The broken reasoning treats “not significant” as “no effect”. The interval runs from a 0.05-point loss to a 1.65-point gain: the data cannot rule out a large win. The fixed readout says: “The observed lift is 0.8 points (8% relative). The test cannot distinguish it from zero; the 95% interval is -0.05 to +1.65 points. The test was sized to detect only a 1.5-point lift, so it was underpowered for effects of this size.” The right next step is a properly powered follow-up, not abandonment. Equally, if the p had come out at 0.049, the same evidence would not suddenly have become strong: 0.049 and 0.064 are nearly the same amount of evidence.

Common mistake

  • “p = 0.03 means a 3% chance the null is true.” The p-value is P(data this extreme | null), not P(null | data). Turning one into the other needs a prior.
  • “p = 0.03 means a 97% chance the result will replicate.” Replication probability depends on power and the true effect, and is usually much lower.
  • “Not significant means no effect.” Absence of evidence is not evidence of absence; read the confidence interval.
  • “A smaller p-value means a bigger effect.” With enough users, a 0.01-point lift gets p below 0.001. Effect size and p-value are different things.
  • Choosing one-sided after seeing the direction, or testing many metrics and reporting the one that cleared 0.05. Both inflate false positives without changing the reported alpha.

Verify the behavior

If the null is true and the test is valid, p-values should be uniformly distributed: about 5% below 0.05, about 10% in every decile. Simulating A/A tests (identical arms) checks both your understanding and your test code.

rng = random.Random(42)
pvals = []
for _ in range(10_000):
    a = rng.binomialvariate(10_000, 0.10)
    b = rng.binomialvariate(10_000, 0.10)
    pvals.append(two_prop_z(a, 10_000, b, 10_000)[1])
print("share below 0.05:", sum(p < 0.05 for p in pvals) / len(pvals))
# share below 0.05: 0.0473
print([round(sum(i / 10 <= p < (i + 1) / 10 for p in pvals) / len(pvals), 3) for i in range(10)])
# [0.096, 0.1, 0.105, 0.102, 0.102, 0.102, 0.098, 0.106, 0.087, 0.093]
python

The false-positive rate is close to 5% (slightly under, because counts are discrete), and the deciles are flat within noise. Running the same check against your experimentation pipeline is a cheap way to catch a broken variance formula: a pipeline whose A/A false-positive rate is 12% has a bug, often an analysis unit that does not match the randomisation unit.

Follow-up questions

  • Why 0.05? Convention, popularised by Fisher, not a law. Choose alpha from the cost of a false positive: lower for irreversible or risky launches, and much lower (such as 5 × 10⁻⁸ in genome-wide studies) when you run millions of tests.
  • What happens to p-values as the sample grows? For any non-zero true effect, however tiny, the p-value heads towards zero. That is why large-sample tests need a practical-significance threshold, not just a p-value.
  • How do p-values relate to confidence intervals? A two-sided p below 0.05 corresponds to a 95% interval that excludes the null value, when both come from the same model.
  • What is p-hacking? Trying analyses (metrics, segments, outlier rules, stopping times) until one gives p below 0.05. Pre-registering the primary metric and analysis plan prevents it.

Interview exercise

A ranking change runs with 2,000,000 users per arm. Control converts 10.00%, treatment 10.10%. The p-value is 0.0009. The change adds 40 ms of latency to every search. The PM says “p is tiny, ship it.” What do you say?

Answer and reasoning

First quantify the effect: the lift is 0.1 percentage points, a 1% relative increase, and the 95% interval for the difference is about 0.04 to 0.16 points (two_prop_z(200_000, 2_000_000, 202_000, 2_000_000) gives z = 3.326, p = 0.00088). The tiny p-value only says the lift is very unlikely to be exactly zero; with four million users, even small effects become statistically detectable. The decision is about value versus cost: translate 0.04 to 0.16 points into revenue and compare it with the cost of 40 ms latency, which may hurt engagement in ways this test’s primary metric does not capture. I would check the latency guardrails in the same test, and recommend shipping only if the lower end of the interval still pays for the latency cost, or if the latency can be optimised first.

Continue learning

More in Statistics & Data Science

esc