Ch. 32 · Statistics & Data Science

Bayes' Theorem and the Base Rate Fallacy: Interview Guide

Solve the classic 1% disease test question three ways, see why rare events wreck precision, and avoid the base rate traps interviewers set.

~8 min readintermediateupdated Oct 6, 2026

“A disease affects 1% of people. A test is 95% sensitive and 95% specific. You test positive. What is the chance you have the disease?” Variations of this question appear in data science, quant and machine learning interviews constantly, because most people answer 95% and the right answer is about 16%. Interviewers are not testing whether you remember the formula. They want to see whether you notice the base rate, whether you can explain the result to a non-statistician, and whether you connect it to the work: fraud alerts, anomaly detectors and any classifier for a rare class suffer from exactly the same arithmetic.

Before you start

You need conditional probability: P(A | B) means the probability of A given that B happened. Know the vocabulary of a diagnostic test: sensitivity is P(positive | condition), the true-positive rate, and specificity is P(negative | no condition), so 1 - specificity is the false-positive rate. The code uses the Python 3.14 standard library; outputs come from real runs with the seeds shown, and the blocks run in order as one script.

The short answer

Bayes’ theorem reverses a conditional probability: P(D | +) = P(+ | D) × P(D) / P(+). With 1% prevalence and 95% sensitivity and specificity, the answer is about 16%: of 10,000 people, 95 sick people test positive but so do 495 healthy people, so only 95 of 590 positives are real. The result is dominated by the base rate: when the condition is rare, a small false-positive rate applied to the large healthy group swamps the true positives. Ignoring the prior is called base rate neglect.

How it works

The test’s accuracy figures are conditioned on the truth: given that you are sick, how often does it say so. The question asks the reverse: given what the test said, how likely is the truth. Bayes’ theorem connects the two through the prior P(D) and the overall positive rate P(+), which itself has two sources: sick people testing positive and healthy people testing positive.

import random

def posterior(prior, sens, spec):
    p_pos = sens * prior + (1 - spec) * (1 - prior)   # law of total probability
    return sens * prior / p_pos

print(round(posterior(0.01, 0.95, 0.95), 4))   # 0.161
python

The denominator is where intuition fails. 0.95 × 0.01 = 0.0095 of the population are true positives, but 0.05 × 0.99 = 0.0495 are false positives, five times as many.

Step-by-step walkthrough

Step 1: Turn probabilities into counts

The clearest way to answer in an interview, and to explain it to a product manager, is natural frequencies: imagine 10,000 people and count.

Test positive Test negative Total
Sick (1%) 95 5 100
Healthy (99%) 495 9,405 9,900
Total 590 9,410 10,000
population = 10_000
sick = population * 0.01
true_pos = sick * 0.95
false_pos = (population - sick) * 0.05
print(round(true_pos), round(false_pos), round(true_pos / (true_pos + false_pos), 3))   # 95 495 0.161
python

Reading down the “Test positive” column gives the answer without any formula: 95 of 590.

Step 2: Write it as Bayes’ theorem and see what drives it

The function from earlier lets you vary the prior. The test never changes; only the population it is used on does.

for prevalence in (0.001, 0.01, 0.1, 0.5):
    print(prevalence, round(posterior(prevalence, 0.95, 0.95), 3))
# 0.001 0.019
# 0.01 0.161
# 0.1 0.679
# 0.5 0.95
python

The same test is nearly useless as a screen when 1 in 1,000 people is sick (2% of positives are real) and very informative in a clinic where half the patients have the condition. When the prior is 50% and sensitivity equals specificity, the posterior equals the sensitivity, which is the only situation where the naive answer of 95% is correct.

Step 3: Use the odds form for quick mental arithmetic

Bayes’ theorem is simplest in odds: posterior odds = prior odds × likelihood ratio. The positive likelihood ratio is sensitivity / (1 - specificity), how much more likely a positive is for a sick person than for a healthy one.

prior_odds = 0.01 / 0.99
lr_positive = 0.95 / 0.05          # 19
post_odds = prior_odds * lr_positive
print(round(post_odds, 4), round(post_odds / (1 + post_odds), 4))   # 0.1919 0.161
python

Prior odds of 1 to 99 times 19 gives 19 to 99, which is 19 / 118 = 0.161. In an interview you can do this in your head, and it makes the question “what specificity would you need for a positive to mean more likely than not?” easy: you need the likelihood ratio to exceed 99, so with 95% sensitivity the false-positive rate must be below 0.95 / 99, a specificity above 99.04%.

needed_spec = 1 - 0.95 * 0.01 / 0.99
print(round(needed_spec, 4), round(posterior(0.01, 0.95, needed_spec), 3))   # 0.9904 0.5
python

Step 4: Update again, but check independence

After one positive, the 16% posterior becomes the prior for the next piece of evidence. A second, independent positive test gives:

first = posterior(0.01, 0.95, 0.95)
print(round(posterior(first, 0.95, 0.95), 4))   # 0.7848
python

In odds, that is 1/99 × 19 × 19, about 3.6 to 1, or 78%. This is why screening programmes confirm positives with a second test. But the multiplication assumes the second test’s errors are independent of the first. If 3% of healthy people have a trait that makes both tests react, and the remaining false positives are independent, the overall false-positive rate is still 5%, yet two positives prove much less:

sick_both = 0.01 * 0.95 ** 2
healthy_both = 0.99 * (0.03 + 0.97 * 0.0206 ** 2)   # 0.03 + 0.97 * 0.0206 is still about 5%
print(round(sick_both / (sick_both + healthy_both), 4))   # 0.2306
python

Repeating the same test on the same sample is the extreme case: it mostly re-measures the same error. The confirmatory test should use a different mechanism.

Worked scenario

A payments team ships a fraud model and announces it is “99% accurate”: 99% of fraud is flagged and 99% of legitimate transactions pass. Fraud is 0.1% of transactions. Within a week, the review team is overwhelmed and complains that almost every alert is legitimate.

print(round(posterior(0.001, 0.99, 0.99), 4))    # 0.0902
print(round(posterior(0.001, 0.90, 0.999), 4))   # 0.4739
python

The broken reasoning took accuracy as the probability an alert is fraud. Out of 1,000,000 transactions, 1,000 are fraud and 990 are flagged, but 1% of the 999,000 legitimate ones, 9,990, are also flagged. Precision is 9%: eleven alerts per real case. “Accuracy” was 99% only because the classes are so unbalanced that it is dominated by the easy negatives.

The fix is to choose the operating threshold by precision and recall instead of accuracy. Raising the threshold so the false-positive rate drops to 0.1%, accepting that recall falls to 90%, lifts precision to 47%, a workable review queue. Further gains come from features that separate the classes better (raising the likelihood ratio) or a two-stage design where a cheap high-recall model feeds a slower, more specific one.

Common mistake

  • Confusing P(+ | D) with P(D | +), sometimes called the prosecutor’s fallacy when it appears in court: the probability of the evidence given innocence is not the probability of innocence given the evidence.
  • Ignoring the prior because it “seems subjective”. Prevalence is usually a measured fact, and leaving it out is itself an assumption of 50%.
  • Multiplying evidence that is not independent, such as two alerts driven by the same feature or two tests with the same failure mode.
  • Quoting accuracy for rare classes. A model that predicts “not fraud” for everything is 99.9% accurate here. Report precision, recall and the base rate.
  • Forgetting that the posterior depends on who gets tested. A test used on symptomatic patients has a higher prior than the same test used on everyone.

Verify the behavior

Simulate a million people and count, rather than trusting the algebra.

rng = random.Random(4)
tp = fp = 0
for _ in range(1_000_000):
    is_sick = rng.random() < 0.01
    positive = rng.random() < (0.95 if is_sick else 0.05)
    if positive:
        tp += is_sick
        fp += not is_sick
print(tp, fp, round(tp / (tp + fp), 4))   # 9532 49706 0.1609
assert abs(tp / (tp + fp) - posterior(0.01, 0.95, 0.95)) < 0.005
python

Follow-up questions

  • How does this relate to precision in classification? Precision is exactly P(positive class | predicted positive), a Bayesian posterior; recall is the sensitivity and the false-positive rate is 1 - specificity.
  • What is the difference between a frequentist and a Bayesian here? Both agree on this calculation because the prior is a known frequency. They diverge when the “prior” is a belief about a parameter, such as a conversion rate.
  • What is naive Bayes and why is it naive? A classifier that multiplies per-feature likelihood ratios, assuming features are conditionally independent given the class. Like the correlated tests above, that assumption is usually false, which makes its probabilities overconfident even when its rankings are good.
  • How do you pick a prior for an A/B test in a Bayesian analysis? From the distribution of past experiment results, which typically shows most lifts are small, so it shrinks surprising estimates towards zero.

Interview exercise

A jar holds 1,000 coins. One has heads on both sides; the rest are fair. You pick a coin at random, flip it 10 times and get 10 heads. What is the probability you picked the double-headed coin? What is the probability the next flip is heads?

Answer and reasoning

Prior odds of the double-headed coin are 1 to 999. The likelihood of 10 heads is 1 for that coin and (1/2)^10 = 1/1024 for a fair coin, so the likelihood ratio is 1,024. Posterior odds are 1,024 to 999, a probability of 1024 / 2023, about 0.506. Even ten straight heads only just tips the balance, because the prior was so small. For the next flip, average over both possibilities: 0.506 × 1 + 0.494 × 0.5, about 0.753. The interviewer is checking that you set up prior, likelihood and posterior explicitly and that you use the posterior, not the prior, for the prediction.

Continue learning

More in Statistics & Data Science

esc