Ch. 32 · Statistics & Data Science

Central Limit Theorem Explained: Data Science Interview Guide

What the central limit theorem really promises, shown by simulation on skewed revenue data, and where it quietly breaks in A/B tests.

~9 min readintermediateupdated Oct 6, 2026

“Explain the central limit theorem” appears in nearly every data science and quant interview loop, usually followed by a sharper question: “Revenue per user is wildly skewed. Why can we still use a t-test on it?” Interviewers ask because the CLT is the reason almost every A/B test readout works, and candidates who understand it can also say when it stops working. Candidates who memorised “n of 30 makes everything normal” give answers that are wrong for exactly the metrics companies care most about.

Before you start

You should know what a mean, a standard deviation and a histogram are, and that the normal distribution has about 95% of its mass within 1.96 standard deviations of its centre. The code uses the Python 3.14 standard library only (random, math, statistics); every output in the comments came from a real run of the code in order, with the seed shown. Run the blocks as one script, since later blocks reuse the random generator and the simulated population.

The short answer

The central limit theorem says that if you take independent samples from any distribution with a finite variance, the distribution of the sample mean gets closer to a normal distribution as the sample size grows, centred on the true mean with a standard deviation of sigma / sqrt(n), the standard error. The individual data stays as skewed as ever; only averages become normal. How large n must be depends on the skew: a symmetric metric is fine at a few dozen, a revenue metric with big spenders can need thousands.

How it works

Think of a sample mean as a sum of many independent pieces, each divided by n. Extreme values push the sum around, but in a large sum, unusually high draws and unusually low draws tend to offset, and the leftover wobble has a predictable bell shape. Two quantities control the result:

  • The centre: the mean of the sample means equals the population mean, for any n. The sample mean is unbiased.
  • The spread: the standard deviation of sample means, the standard error, is sigma / sqrt(n). Four times the data halves it.

The shape is the part that takes time. A useful rule of thumb comes from the maths of sums: the skewness of the sample mean is the population skewness divided by sqrt(n). A population with skewness 14 needs n around 2,000 before its mean has skewness near 0.3, which is close enough to normal for most purposes. Here is such a population: revenue per user where 90% of users spend nothing and buyers spend a lognormal amount.

import math, random
from statistics import fmean, stdev, median

def revenue(rng):
    # 90% of users buy nothing; buyers spend a lognormal amount (median about 33)
    return rng.lognormvariate(3.5, 1.0) if rng.random() < 0.10 else 0.0

def skew(xs):
    m, s = fmean(xs), stdev(xs)
    return fmean(((x - m) / s) ** 3 for x in xs)

rng = random.Random(1)
population = [revenue(rng) for _ in range(1_000_000)]
mu, sigma = fmean(population), stdev(population)
print(f"mean {mu:.2f}  sd {sigma:.2f}  median {median(population):.1f}  skew {skew(population):.1f}")
# mean 5.45  sd 27.90  median 0.0  skew 13.7
python

Nothing about this looks normal: the median user spends zero, the standard deviation is five times the mean, and the skewness is 13.7 (a normal distribution has 0).

Step-by-step walkthrough

Step 1: Draw many samples and record each mean

The sampling distribution is what you would see if you could rerun the same study thousands of times. We can, because we own the population: draw a sample of size n, record its mean, repeat.

for n in (10, 100, 1000):
    means = [fmean(rng.choices(population, k=n)) for _ in range(4000)]
    print(f"n={n:<5} sd of means {stdev(means):.3f}  sigma/sqrt(n) {sigma / math.sqrt(n):.3f}  skew {skew(means):.2f}")
# n=10    sd of means 8.432  sigma/sqrt(n) 8.823  skew 3.34
# n=100   sd of means 2.865  sigma/sqrt(n) 2.790  skew 1.31
# n=1000  sd of means 0.888  sigma/sqrt(n) 0.882  skew 0.45
python

Step 2: Check the spread against the standard error formula

The second and third columns match: the spread of sample means is sigma / sqrt(n) at every sample size, even n = 10 where the shape is still badly skewed. The formula also tells you what precision costs:

for n in (250, 1000, 4000):
    print(n, round(sigma / math.sqrt(n), 3))
# 250 1.765
# 1000 0.882
# 4000 0.441
python

Each fourfold increase in users halves the standard error. The formula for the standard error does not need the CLT at all; it only needs independent draws with finite variance. This is a common confusion. What the CLT adds is the shape, which you need to turn a standard error into a p-value or a 95% interval with the number 1.96.

Step 3: Watch the skew shrink, and measure what it costs

The skewness of the means fell from 3.34 to 1.31 to 0.45, close to the predicted 13.7 / sqrt(n) (4.3, 1.4 and 0.43). To see what residual skew costs in practice, build the textbook 95% interval, mean ± 1.96 × standard error, on thousands of samples and count how often it contains the true mean.

def coverage(pop, target, n, reps=4000):
    hits = 0
    for _ in range(reps):
        s = rng.choices(pop, k=n)
        m, se = fmean(s), stdev(s) / math.sqrt(n)
        hits += m - 1.96 * se <= target <= m + 1.96 * se
    return hits / reps

for n in (30, 300, 3000):
    print(n, coverage(population, mu, n))
# 30 0.71725
# 300 0.89375
# 3000 0.94375
python

At n = 30 the “95%” interval contains the truth only 72% of the time. At 300 it reaches 89%, and only at 3,000 is it close to its label. The misses are lopsided: a small sample usually contains no big spenders, so its mean and its standard deviation are both too low together, and the interval sits entirely below the truth.

Step 4: Know where the theorem does not apply

The finite-variance condition is not a technicality. For a Cauchy distribution, which has tails so heavy that neither mean nor variance exists, the average of a million draws is no more stable than the average of a hundred.

crng = random.Random(3)
def cauchy():
    return math.tan(math.pi * (crng.random() - 0.5))

for n in (100, 10_000, 1_000_000):
    print(n, [round(fmean(cauchy() for _ in range(n)), 2) for _ in range(4)])
# 100 [3.45, -0.86, -2.28, -0.48]
# 10000 [-0.19, 0.02, -0.29, -0.34]
# 1000000 [0.71, 1.54, -8.21, 0.06]
python

Power-law data such as followers per account can behave like this at realistic sample sizes. Independence matters too: sessions from the same user are correlated, so a CLT applied to sessions when you randomised users gives standard errors that are too small.

Worked scenario

Finance runs a pilot with 300 users to estimate average revenue per user and asks for a 95% interval. The analyst computes mean ± 1.96 × sd / sqrt(300) and reports it. From Step 3 we know that procedure covers the truth only about 89% of the time on this kind of data, and when it misses, it almost always underestimates. The “95%” label is wrong, and the forecast built on it is biased low.

Two fixes, each with a cost:

cap = sorted(population)[int(0.99 * len(population))]   # 99th percentile
capped = [min(x, cap) for x in population]
print(round(cap, 1), round(fmean(capped), 2))      # 119.2 4.52
print(coverage(capped, fmean(capped), 300))        # 0.93
python
  • More data: the honest fix. Coverage reaches about 94% at 3,000 users.
  • Capping (winsorising) at the 99th percentile lifts coverage to 93% at n = 300, but it changes the question: the capped mean is 4.52, not 5.45. That is acceptable for an A/B test, where you want a sensitive comparison, and unacceptable for a revenue forecast, where you need the real total.

A/B tests are more forgiving than single-arm estimates. When both arms have the same skewed shape, the skew largely cancels in the difference of means:

def diff_coverage(n, reps=4000):
    hits = 0
    for _ in range(reps):
        a, b = rng.choices(population, k=n), rng.choices(population, k=n)
        d = fmean(b) - fmean(a)
        se = math.sqrt(stdev(a) ** 2 / n + stdev(b) ** 2 / n)
        hits += d - 1.96 * se <= 0 <= d + 1.96 * se
    return hits / reps

print(diff_coverage(300))   # 0.95675
python

That is why A/A false-positive rates on revenue look fine at moderate n even when a single-arm interval would not. It does not rescue precision, which is where capping and variance reduction earn their keep.

Common mistake

  • “The CLT makes the data normal.” It makes the sampling distribution of the mean approximately normal. The data keeps its shape.
  • “n = 30 is enough.” That rule of thumb came from mildly skewed textbook examples. With skewness 13.7, n = 30 gives 72% coverage on a nominal 95% interval.
  • “The CLT gives the standard error.” The sigma / sqrt(n) formula holds at any n for independent draws; the CLT adds the normal shape that turns it into p-values.
  • Applying it to the median or the maximum the same way. Medians have their own asymptotic theory with a different standard error, and maxima follow extreme-value distributions, not the normal.
  • Ignoring dependence: per-session or per-pageview analyses of user-randomised tests violate independence.

Verify the behavior

Turn the theory into assertions: the standard error matches its formula at every n, and coverage approaches 95% as n grows on this skewed population.

vrng = random.Random(11)
for n in (50, 2000):
    means = [fmean(vrng.choices(population, k=n)) for _ in range(3000)]
    ratio = stdev(means) / (sigma / math.sqrt(n))
    assert 0.9 < ratio < 1.1, ratio
assert coverage(population, mu, 40) < 0.85 < 0.93 < coverage(population, mu, 4000)
print("standard error and coverage checks passed")
# standard error and coverage checks passed
python

Follow-up questions

  • How do you know if n is large enough? Check the skewness of the metric and divide by sqrt(n), or simulate: resample your own data at the planned n and look at the distribution of means or the A/A false-positive rate.
  • Does the CLT apply to proportions? Yes: a proportion is the mean of 0/1 values. The usual rule is at least about 10 successes and 10 failures before the normal approximation is trusted.
  • What is the difference between the CLT and the law of large numbers? The law of large numbers says the sample mean converges to the true mean. The CLT describes the shape and size of the error around it on the way.
  • Why does a t-test work on skewed revenue? It tests a difference in means, and with enough users per arm the means are close to normal; with few users, use more data, capping or a permutation test.

Interview exercise

Session length on your app has mean 4 minutes and standard deviation 6 minutes, with a long right tail (skewness about 6). A test has 400 users per arm, and you analyse per-user average session length. A colleague says the t-test is invalid because session length is not normal. Are they right? What standard error do you expect for one arm’s mean, and what would make you worried?

Answer and reasoning

The t-test does not need normal data; it needs approximately normal sample means. The standard error of one arm’s mean is 6 / sqrt(400) = 0.3 minutes, which holds regardless of shape. The skewness of the mean is roughly 6 / sqrt(400) = 0.3, which is mild, and the difference between two similarly skewed arms is closer to symmetric still, so the test is reasonable. I would get worried if a few users had sessions of many hours (bots or tabs left open), which inflate the skew; I would check the maximum values, consider capping at a high percentile for the test metric, and validate by running A/A simulations on historical data at n = 400 to confirm the false-positive rate is near 5%.

Continue learning

More in Statistics & Data Science

esc