“What does a 95% confidence interval mean?” is a favourite interview question because the natural answer, “there is a 95% chance the true value is inside”, is subtly wrong, and the follow-ups expose whether you can actually build one. Interviewers then push into practice: compute an interval for a conversion rate with 3 successes out of 50, explain why the interval for zero errors is not zero, put a range on median latency. Every one of those appears in real work, and an analyst who reports a point estimate without an honest interval invites bad decisions.
Before you start
You should know the mean, the standard deviation and the standard error (sd / sqrt(n), the spread of the sample mean). Knowing that about 95% of a normal distribution lies within 1.96 standard deviations is enough background for the critical values. The code uses the Python 3.14 standard library; outputs in comments are from real runs with the seeds shown, and the blocks are meant to run in order as one script.
The short answer
A 95% confidence interval comes from a procedure that, over many repetitions of the study, captures the true value 95% of the time. The 95% describes the method, not any single interval: once computed, a specific interval either contains the truth or does not. Most intervals are estimate ± critical value × standard error, with 1.96 for large samples, a t critical value for small samples of means, the Wilson formula for proportions, and the bootstrap when no formula fits. I report the width as well as the centre, because the width is the honest statement of uncertainty.
How it works
The logic runs backwards from the sampling distribution. If an estimate is approximately normal around the true value θ with standard error SE, then 95% of the time the estimate lands within 1.96 SE of θ. Turn that sentence around: 95% of the time, θ lies within 1.96 SE of the estimate. The interval moves from sample to sample; the true value stays put.
Three things set the width: the critical value (more confidence means wider), the spread of the data, and the sample size, which enters through sqrt(n). Quadrupling the data halves the width, which is why precision gets expensive quickly.
Step-by-step walkthrough
Step 1: Compute the estimate and its standard error
Ten page loads from a new build, in milliseconds:
import math, random
from statistics import fmean, stdev, median, NormalDist
load_ms = [812, 1040, 955, 1210, 880, 1475, 990, 1130, 920, 1060]
n = len(load_ms)
m = fmean(load_ms)
se = stdev(load_ms) / math.sqrt(n)
print(m, round(stdev(load_ms), 1), round(se, 1)) # 1047.2 190.9 60.4stdev divides by n - 1, the unbiased sample variance, which is the right input for a standard error.
Step 2: Pick the right critical value
With ten observations we estimated the standard deviation from the same data, and that estimate is itself noisy. Student’s t distribution accounts for it with heavier tails, so its critical value is larger than 1.96 for small samples. The standard library has no t distribution; take the value from a t table (or scipy.stats.t.ppf(0.975, 9) if SciPy is installed).
z = NormalDist().inv_cdf(0.975)
t = 2.262 # t critical value for 0.975 with n - 1 = 9 degrees of freedom
print(round(z, 3)) # 1.96
print("z:", round(m - z * se, 1), round(m + z * se, 1)) # z: 928.9 1165.5
print("t:", round(m - t * se, 1), round(m + t * se, 1)) # t: 910.7 1183.7The t interval is about 15% wider. The t critical value is 2.042 at 30 degrees of freedom and 1.962 at 1,000, so the distinction only matters for small samples; with thousands of users per arm, z and t give the same answer.
Step 3: Build a proportion interval that does not lie at the edges
The textbook (Wald) interval for a proportion, p ± 1.96 sqrt(p(1 - p)/n), behaves badly when n is small or p is near 0 or 1. The Wilson interval inverts the score test instead and is the better default.
def wald(k, n, z=1.96):
p = k / n
h = z * math.sqrt(p * (1 - p) / n)
return max(0.0, p - h), min(1.0, p + h)
def wilson(k, n, z=1.96):
p = k / n
d = 1 + z * z / n
centre = (p + z * z / (2 * n)) / d
h = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / d
return centre - h, centre + h
for k, total in ((3, 50), (0, 50), (120, 1000)):
print(k, total, [round(v, 4) for v in wald(k, total)], [round(v, 4) for v in wilson(k, total)])
# 3 50 [0.0, 0.1258] [0.0206, 0.1622]
# 0 50 [0.0, 0.0] [0.0, 0.0714]
# 120 1000 [0.0999, 0.1401] [0.1013, 0.1416]At 3 of 50 the Wald interval runs below zero and has to be clipped. At 0 of 50 it collapses to a single point, claiming certainty that the rate is exactly zero. At 120 of 1,000 the two methods agree closely. Wilson’s centre is pulled slightly towards 0.5, which is exactly what fixes the edge cases.
Step 4: Use the bootstrap when there is no formula
Medians, percentiles, ratios and differences in medians have messy or no closed-form standard errors. The percentile bootstrap resamples the data with replacement many times, recomputes the statistic each time, and takes the middle 95% of the results.
rng = random.Random(2)
latency = [rng.lognormvariate(5.5, 0.6) for _ in range(400)]
boot = sorted(median(rng.choices(latency, k=len(latency))) for _ in range(5000))
print(round(median(latency), 1), round(boot[124], 1), round(boot[4874], 1))
# 245.8 232.2 263.9The 2.5th and 97.5th percentiles of 5,000 bootstrap medians are positions 124 and 4874 when counting from zero. Resample the unit you randomised or sampled: if you have many requests per user, resample users, not requests, or the interval will be too narrow.
Worked scenario
A canary release serves 50 requests with zero errors. The deploy checklist asks for the error rate with a 95% interval, and the engineer reports “0.0% (95% CI 0.0% to 0.0%)” using the Wald formula, then promotes the release to all traffic.
The interval is nonsense: zero events in 50 trials is perfectly compatible with a true error rate of a few percent. The fixed report uses Wilson, or the rule of three: when you observe zero events in n trials, the 95% upper bound is about 3 / n.
print([round(v, 4) for v in wilson(0, 50)]) # [0.0, 0.0714]
print(3 / 50, round(1 - 0.05 ** (1 / 50), 4)) # 0.06 0.0582The exact one-sided bound solves (1 - p)^50 = 0.05, giving 5.8%, and 3 / n approximates it well. The honest statement is “no errors observed; the error rate is plausibly anywhere up to about 6 to 7%.” If the SLO allows 1%, the canary needs about 300 clean requests (3 / 0.01) before it says anything useful.
Common mistake
- “There is a 95% probability the true value is in this interval.” In the frequentist framework the true value is fixed; the 95% is the long-run hit rate of the procedure. A Bayesian credible interval does carry that interpretation, given a prior.
- “95% of the data falls inside the interval.” That is a prediction or reference interval, which is far wider. A confidence interval is about a parameter such as the mean.
- Comparing two groups by checking whether their intervals overlap. Two 95% intervals can overlap while the difference is still significant; build an interval for the difference directly.
- Using Wald for small samples or rare events, and resampling the wrong unit in a bootstrap.
- Thinking a wider confidence level is “more accurate”. It is more likely to contain the truth and less precise.
Verify the behavior
The definition is testable: simulate many studies where you know the truth and count how often the interval captures it. This checks the t interval on small normal samples, then Wald against Wilson on a rare event.
vrng = random.Random(9)
def mean_coverage(crit, n=10, reps=20_000):
hits = 0
for _ in range(reps):
s = [vrng.gauss(100, 15) for _ in range(n)]
half = crit * stdev(s) / math.sqrt(n)
hits += fmean(s) - half <= 100 <= fmean(s) + half
return hits / reps
print(mean_coverage(1.96), mean_coverage(2.262)) # 0.91965 0.9502
def prop_coverage(method, p, n, reps=20_000):
hits = 0
for _ in range(reps):
lo, hi = method(vrng.binomialvariate(n, p), n)
hits += lo <= p <= hi
return hits / reps
print(prop_coverage(wald, 0.03, 50), prop_coverage(wilson, 0.03, 50)) # 0.7819 0.9376With ten observations, the z interval only covers 92% of the time while the t interval delivers its 95%. For a 3% rate with 50 trials, the Wald interval covers 78% of the time, far from its label, while Wilson is close to 94%. Discrete counts make exact 95% coverage impossible, so small wobbles around the target are expected.
Follow-up questions
- How does the width change if you quadruple the sample? It halves, because the standard error scales with
1 / sqrt(n). - How do intervals relate to p-values? A two-sided p-value below 0.05 corresponds to a 95% interval that excludes the null value, when both come from the same model.
- When would you prefer a bootstrap? For medians, percentiles, ratio metrics and anything without a reliable formula, as long as the sample is large enough to represent the tails.
- What is a credible interval? The Bayesian counterpart: given a prior and the data, the posterior probability that the parameter lies inside it is 95%.
Interview exercise
An experiment shows a conversion lift of +0.6 percentage points with a 95% confidence interval of -0.2 to +1.4 points. The PM says: “The interval includes zero, so the feature does nothing.” A second team ran the same test on another market and got +0.5 points with an interval of +0.1 to +0.9. How do you summarise the evidence?
Answer and reasoning
The first interval includes zero, so that test alone cannot rule out no effect, but it equally cannot rule out a lift of more than one point: “inconclusive” is the right word, not “does nothing.” The two estimates (+0.6 and +0.5) are very consistent with each other; the first interval is wider, most likely because that test had less traffic. Taken together, the evidence points to a small positive lift around half a point. A sensible next step is to pool the two results (a fixed-effect meta-analysis weights each estimate by the inverse of its variance), check that the markets are comparable, and decide whether a lift of that size is worth shipping.
Continue learning
- Practise the theory in the Statistics & Data Science chapter and test yourself with the data science MCQs.
- Learn why the normal approximation works, and when it does not, in the central limit theorem.
- See how intervals and p-values fit together in p-values and hypothesis testing.
- Read the NIST/SEMATECH handbook on confidence limits for the mean and on confidence intervals for a proportion.