Statistics & Data Science MCQs multiple-choice questions with answers & explanations
All 20 Statistics & Data Science quiz questions on one page. Pick an answer in your head, then open Show answer to check it and read why. Want a score and a timer? Take them as a quiz instead.
Official reference: NIST Engineering Statistics Handbook
- 1.easy
Nine customers spend 20 and one spends 2,000. What does this print?
from statistics import mean, median spend = [20] * 9 + [2000] print(mean(spend), median(spend))- A
218 20.0 - B
218 218 - C
20 20.0 - D
218.0 1010.0
Show answer
Answer: A (
218 20.0)The mean is 2,180 / 10 = 218 and stays an
intbecause every input is anint. The median of an even-length list averages the two middle values, 20 and 20, so it returns the float20.0. One outlier moves the mean tenfold and leaves the median untouched. - A
- 2.easy
What does this print?
from statistics import pvariance, variance data = [2, 4, 4, 4, 5, 5, 7, 9] print(pvariance(data), round(variance(data), 3))- A
4 4 - B
4 4.571 - C
4.571 4 - D
2 2.138
Show answer
Answer: B (
4 4.571)The squared deviations from the mean of 5 sum to 32.
pvariancedivides by n = 8, giving 4;varianceis the sample variance and divides by n - 1 = 7, giving 4.571. Then - 1version corrects for measuring deviations from the sample mean. - A
- 3.easy
Scores are normal with mean 100 and standard deviation 15. What does this print?
from statistics import NormalDist iq = NormalDist(mu=100, sigma=15) print(round(iq.cdf(130), 3), round(iq.inv_cdf(0.975), 1))- A
0.95 130.0 - B
0.977 129.4 - C
0.997 145.0 - D
0.841 115.0
Show answer
Answer: B (
0.977 129.4)130 is two standard deviations above the mean, and about 97.7% of a normal distribution lies below z = 2. The 97.5th percentile is 1.96 standard deviations up, 100 + 1.96 * 15 = 129.4, not exactly two.
- A
- 4.mid
A condition has 1% prevalence; the test has 95% sensitivity and 95% specificity. What does this print?
prior, sens, spec = 0.01, 0.95, 0.95 p_pos = sens * prior + (1 - spec) * (1 - prior) print(round(sens * prior / p_pos, 3))- A
0.95 - B
0.5 - C
0.161 - D
0.01
Show answer
Answer: C (
0.161)Of 10,000 people, 95 sick people test positive and 495 healthy people also test positive, so only 95 / 590 = 0.161 of positives are real. Answering 0.95 confuses the sensitivity P(+ | sick) with the posterior P(sick | +) and ignores the base rate.
- A
- 5.mid
An experiment tracks 20 independent metrics, none of which is truly affected. At alpha 0.05, what does this print?
print(round(1 - 0.95 ** 20, 2))- A
0.05 - B
0.36 - C
0.64 - D
1.0
Show answer
Answer: C (
0.64)The chance that every one of 20 null tests stays non-significant is 0.95^20, about 0.36, so the chance of at least one false positive is about 0.64. That is why experiments name one primary metric in advance or correct for multiple comparisons.
- A
- 6.mid
A service averages 3 errors per hour (Poisson). What does this print?
import math lam = 3 print(round(math.exp(-lam), 3), round(1 - math.exp(-lam) * sum(lam**k / math.factorial(k) for k in range(3)), 3))- A
0.05 0.577 - B
0.0 0.5 - C
0.05 0.423 - D
0.333 0.667
Show answer
Answer: A (
0.05 0.577)For a Poisson with rate 3, P(0) = e^-3, about 0.05. The second value is P(X >= 3) = 1 - P(0) - P(1) - P(2) = 1 - e^-3 (1 + 3 + 4.5), about 0.577. 0.423 is P(X <= 2), the complement.
- A
- 7.easy
What is the probability of at least one six in four rolls of a fair die?
print(round(1 - (5 / 6) ** 4, 3))- A
0.667 - B
0.518 - C
0.482 - D
0.167
Show answer
Answer: B (
0.518)Use the complement: the chance of no six in four independent rolls is (5/6)^4, about 0.482, so at least one six is 0.518. Adding 1/6 four times (0.667) double-counts outcomes with several sixes.
- A
- 8.mid
What does this birthday-problem calculation print for 23 people?
p_unique = 1.0 for i in range(23): p_unique *= (365 - i) / 365 print(round(1 - p_unique, 3))- A
0.063 - B
0.23 - C
0.507 - D
0.999
Show answer
Answer: C (
0.507)The loop multiplies the chances that each new person avoids every earlier birthday, and 23 people already give 253 pairs, so a shared birthday is slightly more likely than not. 0.063 is the chance that someone shares a birthday with one specific person, a different question.
- A
- 9.mid
In Monty Hall, the host always opens a goat door you did not pick. What win rate for switching does this print?
import random rng = random.Random(42) wins = 0 for _ in range(100_000): car, pick = rng.randrange(3), rng.randrange(3) wins += pick != car # switching wins exactly when the first pick was wrong print(round(wins / 100_000, 2))- A
0.33 - B
0.5 - C
0.67 - D
1.0
Show answer
Answer: C (
0.67)Your first pick is wrong two times in three. Because the host knowingly removes the other goat, switching turns every wrong first pick into a win, so it wins about 2/3 of the time. The 0.5 intuition treats the two remaining doors as symmetric, but the host's choice depended on where the car was.
- A
- 10.mid
Draws come from a skewed exponential distribution with mean 1 and standard deviation 1. What does this print?
import random from statistics import mean, stdev rng = random.Random(0) means = [mean(rng.expovariate(1.0) for _ in range(30)) for _ in range(2000)] print(round(mean(means), 2), round(stdev(means), 2))- A
1.0 1.0 - B
1.0 0.18 - C
0.69 0.18 - D
1.0 0.03
Show answer
Answer: B (
1.0 0.18)The sample means centre on the population mean, 1.0, and their spread is the standard error sigma / sqrt(n) = 1 / sqrt(30), about 0.18. The spread of the raw data (1.0) is not the spread of means, and 0.69 is the median of the exponential, not the mean.
- A
- 11.mid
What does this print?
from statistics import correlation x = [-3, -2, -1, 0, 1, 2, 3] y = [v * v for v in x] print(correlation(x, y))- A
1.0 - B
0.0 - C
-1.0 - D
0.5
Show answer
Answer: B (
0.0)Pearson correlation measures linear association only.
yis completely determined byx, but the relationship is a symmetric parabola, so the positive and negative halves cancel and the correlation is exactly 0. Zero correlation does not mean independence. - A
- 12.easy
What does this print?
from statistics import linear_regression slope, intercept = linear_regression([1, 2, 3, 4, 5], [2, 4, 5, 4, 5]) print(round(slope, 2), round(intercept, 2))- A
1.0 1.0 - B
0.6 2.2 - C
0.75 1.5 - D
0.6 4.0
Show answer
Answer: B (
0.6 2.2)The least-squares slope is cov(x, y) / var(x) = 1.5 / 2.5 = 0.6, and the line passes through the means (3, 4), so the intercept is 4 - 0.6 * 3 = 2.2.
statistics.linear_regressionreturns a named tuple you can unpack. - A
- 13.mid
Two onboarding flows got different device mixes. What does the last value on each line show?
data = { # (conversions, users) "old": {"mobile": (20, 400), "desktop": (90, 600)}, "new": {"mobile": (60, 1000), "desktop": (19, 100)}, } for flow, seg in data.items(): rates = {k: c / n for k, (c, n) in seg.items()} total = sum(c for c, _ in seg.values()) / sum(n for _, n in seg.values()) print(flow, {k: round(v, 3) for k, v in rates.items()}, round(total, 3)) # old {'mobile': 0.05, 'desktop': 0.15} 0.11 # new {'mobile': 0.06, 'desktop': 0.19} 0.072- AThe new flow wins overall, matching each segment
- BThe new flow wins in both segments but loses overall
- CThe new flow loses in both segments and overall
- DThe overall rate is the average of the two segment rates
Show answer
Answer: B (The new flow wins in both segments but loses overall)
The new flow converts better on mobile (6% vs 5%) and desktop (19% vs 15%), but 91% of its users are on mobile, which converts poorly, so its pooled rate is lower. That reversal is Simpson's paradox; the overall rate is a mix-weighted average, not the simple average of segments.
- 14.hard
Control converts 100 of 1,000 and treatment 130 of 1,000. What does this print?
import math c1, n1, c2, n2 = 100, 1000, 130, 1000 p = (c1 + c2) / (n1 + n2) z = (c2 / n2 - c1 / n1) / math.sqrt(p * (1 - p) * (1 / n1 + 1 / n2)) obs = [c1, n1 - c1, c2, n2 - c2] exp = [n1 * p, n1 * (1 - p), n2 * p, n2 * (1 - p)] chi2 = sum((o - e) ** 2 / e for o, e in zip(obs, exp)) print(round(z, 3), round(z * z, 3), round(chi2, 3))- A
2.103 4.422 4.422 - B
2.103 4.422 2.103 - C
1.96 3.841 3.841 - D
2.103 4.422 8.844
Show answer
Answer: A (
2.103 4.422 4.422)The pooled two-proportion z statistic is 2.103, and the Pearson chi-square statistic on the same 2x2 table (without continuity correction) is exactly z squared, 4.422. Both give the same two-sided p-value, about 0.035.
- A
- 15.mid
Waiting times are exponential with rate 0.5 per minute. What does this print?
import math rate = 0.5 p_gt = lambda t: math.exp(-rate * t) print(round(p_gt(3) / p_gt(1), 3), round(p_gt(2), 3))- A
0.223 0.368 - B
0.368 0.368 - C
0.607 0.368 - D
0.368 0.135
Show answer
Answer: B (
0.368 0.368)The first value is P(T > 3 | T > 1), and it equals P(T > 2) = e^-1, about 0.368. That is the memoryless property: having already waited one minute does not change the distribution of the remaining wait.
- A
- 16.hard
What does this print?
from statistics import quantiles print(quantiles(range(1, 11), n=4))- A
[3, 5.5, 8] - B
[2.75, 5.5, 8.25] - C
[3.25, 5.5, 7.75] - D
[2.5, 5, 7.5]
Show answer
Answer: B (
[2.75, 5.5, 8.25])statistics.quantilesdefaults tomethod="exclusive", which places cut points at positions (n + 1)p and gives 2.75 and 8.25 for the quartiles.method="inclusive"(the convention NumPy uses by default) gives 3.25 and 7.75. Different tools disagree on small samples, so state the method. - A
- 17.mid
User 1 has orders of 50, 50 and 20. For the 20 order, what are
rn,rnkanddrnk?SELECT user_id, amount, ROW_NUMBER() OVER (PARTITION BY user_id ORDER BY amount DESC) AS rn, RANK() OVER (PARTITION BY user_id ORDER BY amount DESC) AS rnk, DENSE_RANK() OVER (PARTITION BY user_id ORDER BY amount DESC) AS drnk FROM orders;- A
3, 3, 3 - B
3, 2, 2 - C
3, 3, 2 - D
2, 3, 2
Show answer
Answer: C (
3, 3, 2)The two 50s tie.
ROW_NUMBERstill numbers rows 1, 2, 3;RANKgives both ties 1 and skips to 3;DENSE_RANKgives both ties 1 and continues with 2. PickROW_NUMBERwhen you need exactly one row per user, such as the first order. - A
- 18.mid
An A/A test (no real difference) is checked daily for 20 days and stopped the first time p is below 0.05. Roughly what is the false-positive rate?
- A5%, because alpha is 0.05
- BAbout 2.5%, because only one direction can win
- CAround 20 to 25%
- D100%, given enough days
Show answer
Answer: C (Around 20 to 25%)
Every look gives noise another chance to cross the threshold, and stopping at the first crossing locks in the false positive. A simulation with 20 looks gives about 24%, not 5%. It approaches 100% only with unlimited looks; sequential methods fix the problem by adjusting the boundaries.
- 19.easy
A 95% confidence interval for a conversion rate is plus or minus 2 points with 1,000 users. About how wide is it with 4,000 users?
- APlus or minus 0.5 points
- BPlus or minus 1 point
- CPlus or minus 2 points
- DPlus or minus 4 points
Show answer
Answer: B (Plus or minus 1 point)
The margin of error is proportional to the standard error, which shrinks with 1 / sqrt(n). Four times the users halves the margin, so it is about 1 point. Expecting it to shrink by four confuses n with sqrt(n).
- 20.mid
Which change increases the power of an A/B test, everything else fixed?
- ALowering alpha from 0.05 to 0.01
- BChoosing a smaller minimum detectable effect
- CUsing CUPED to reduce the metric's variance
- DSplitting traffic 90/10 instead of 50/50
Show answer
Answer: C (Using CUPED to reduce the metric's variance)
Power rises with lower variance, larger effects, more data and looser alpha. CUPED removes pre-existing variation, so the same effect stands out more clearly. A stricter alpha, a smaller target effect and an unbalanced split all reduce power for a given total sample.
No questions match these filters. Try a different subtopic or clear the filters.