Ch. 32 · Statistics & Data Science

Simpson's Paradox and Correlation vs Causation: Interview Guide

Why a variant can win in every segment yet lose overall, how to standardise for mix, and how to tell confounders from causes in product data.

~8 min readintermediateupdated Oct 6, 2026

Two classic interview questions turn out to be the same question. The first: “A new onboarding flow converts better on mobile and on desktop, but worse overall. How is that possible, and which number do you believe?” The second: “Users who adopt feature X retain much better. The PM wants to push everyone into it. What do you say?” Both test whether you can see a third variable hiding behind a comparison. Interviewers ask them because dashboards are full of pooled averages and observational correlations, and an analyst who takes them at face value will recommend rollbacks of good launches and investments in features that change nothing.

Before you start

You need conversion rates, weighted averages, and the idea of a randomised experiment. A confounder is a variable that influences both the thing you are comparing (which flow, which feature) and the outcome. The code uses the Python 3.14 standard library; outputs in comments are from real runs with the seeds shown.

The short answer

Simpson’s paradox is when a relationship that holds inside every subgroup reverses or vanishes when the subgroups are pooled. It happens when the groups being compared have very different subgroup mixes and the subgroups have very different base rates, so the pooled number mostly measures the mix. To decide which view is right, reason about cause: if the subgrouping variable is a confounder, compare within segments or standardise to a common mix; if it is caused by the treatment, the pooled number is the right one. The same thinking answers correlation versus causation, and the clean fix for both is randomisation.

How it works

A pooled rate is a weighted average of segment rates, with weights equal to each group’s own mix. If two groups use different weights, their pooled rates can differ for reasons that have nothing to do with the rates themselves.

data = {  # flow -> segment -> (conversions, users)
    "old": {"mobile": (20, 400), "desktop": (90, 600)},
    "new": {"mobile": (60, 1000), "desktop": (19, 100)},
}
for flow, seg in data.items():
    conv = sum(c for c, _ in seg.values())
    users = sum(n for _, n in seg.values())
    print(flow, conv, users, round(conv / users, 3))
# old 110 1000 0.11
# new 79 1100 0.072
python

The new flow looks clearly worse: 7.2% against 11%. Hold that thought.

Step-by-step walkthrough

Step 1: Break the comparison down by segment

Pick segments that plausibly affect the outcome strongly: platform, country, acquisition channel, new versus returning. Then compare the rates within each one.

for flow, seg in data.items():
    print(flow, {k: round(c / n, 3) for k, (c, n) in seg.items()})
# old {'mobile': 0.05, 'desktop': 0.15}
# new {'mobile': 0.06, 'desktop': 0.19}
python

The new flow wins on mobile (6% against 5%) and on desktop (19% against 15%). Every segment improved, yet the total fell.

Step 2: Look at the mix

The explanation is always in the weights.

for flow, seg in data.items():
    total = sum(n for _, n in seg.values())
    print(flow, {k: round(n / total, 2) for k, (_, n) in seg.items()})
# old {'mobile': 0.4, 'desktop': 0.6}
# new {'mobile': 0.91, 'desktop': 0.09}
python

91% of the new flow’s users came from mobile, which converts at a third of the desktop rate. The old flow’s traffic was mostly desktop. The pooled comparison is mostly a comparison of device mixes. That pattern appears when a launch is rolled out on one platform first, when a marketing campaign drives a particular kind of traffic to one experience, or when an “experiment” was really a before-and-after comparison.

Step 3: Standardise to a common mix

To compare like with like, apply both flows’ segment rates to the same mix. Any reasonable reference mix works (the combined traffic, last quarter’s traffic, the mix you expect after launch); state which one you used.

mix = {"mobile": 1400 / 2100, "desktop": 700 / 2100}   # combined traffic
for flow, seg in data.items():
    print(flow, round(sum(mix[k] * c / n for k, (c, n) in seg.items()), 4))
# old 0.0833
# new 0.1033
python

At the same mix, the new flow converts about 10.3% against 8.3%. This is called direct standardisation; regression with segment as a covariate does the same job when there are many segments.

Step 4: Decide using the causal structure

Numbers alone cannot tell you which view is right. Draw the arrows. Here, device affects which flow a user saw (the rollout) and affects conversion, so device is a confounder and the within-segment comparison is fair. Now change the story: suppose the new flow itself pushed users from desktop to the mobile app, where they convert worse. Then device is a mediator, part of the flow’s effect, and the pooled number is the honest total effect of shipping it.

The feature-X question has the same shape. In the simulation below, engaged users are both more likely to use feature X and more likely to retain, and X has no effect at all.

import random
from statistics import fmean

rng = random.Random(5)
users = []
for _ in range(200_000):
    engaged = rng.random() < 0.3
    uses_x = rng.random() < (0.6 if engaged else 0.1)
    retained = rng.random() < (0.7 if engaged else 0.3)   # X plays no part
    users.append((engaged, uses_x, retained))

def ret(rows):
    return fmean(r for *_, r in rows)

with_x = [u for u in users if u[1]]
without_x = [u for u in users if not u[1]]
print(round(ret(with_x), 3), round(ret(without_x), 3))   # 0.586 0.365
for e in (True, False):
    print(e, round(ret([u for u in with_x if u[0] == e]), 3),
          round(ret([u for u in without_x if u[0] == e]), 3))
# True 0.698 0.7
# False 0.301 0.3
python

The naive comparison shows a 22-point retention gap. Inside each engagement level it disappears. Real data is harder: you rarely measure “engagement” cleanly, and there may be confounders you never thought of, which is why stratification is weaker evidence than an experiment.

Worked scenario

The weekly conversion alarm fires the week after a checkout redesign: conversion fell from 4.40% to 3.34%. The first instinct in the incident channel is to roll back. The same week, marketing launched a paid campaign.

before = {"organic": (400, 8_000), "paid": (40, 2_000)}
after = {"organic": (416, 8_000), "paid": (252, 12_000)}

def summary(week):
    conv = sum(c for c, _ in week.values())
    users = sum(n for _, n in week.values())
    rates = {k: c / n for k, (c, n) in week.items()}
    shares = {k: n / users for k, (_, n) in week.items()}
    return conv / users, rates, shares

r0, rate0, share0 = summary(before)
r1, rate1, share1 = summary(after)
print(round(r0, 4), round(r1, 4), round(r1 - r0, 4))   # 0.044 0.0334 -0.0106
rate_effect = sum((share0[k] + share1[k]) / 2 * (rate1[k] - rate0[k]) for k in before)
mix_effect = sum((rate0[k] + rate1[k]) / 2 * (share1[k] - share0[k]) for k in before)
print(round(rate_effect, 4), round(mix_effect, 4))   # 0.0016 -0.0122
python

The shift-share decomposition splits the 1.06-point drop exactly into a rate effect (both channels converted better, +0.16 points) and a mix effect (paid traffic, which converts at about 2%, went from 20% to 60% of visitors, -1.22 points). The redesign is not the problem; the campaign brought lower-intent traffic. The fixed response is to keep the redesign, report conversion by channel, and judge the redesign with a proper A/B test rather than week-over-week totals. Notice that pooled numbers are not useless here: total orders went up from 440 to 668, which is what the business cares about.

Common mistake

  • “The segmented view is always right.” Not if the segment is caused by the treatment. Conditioning on a mediator hides real effects, and conditioning on a collider (a variable caused by both treatment and outcome, such as “users who left a review”) can create correlations that do not exist.
  • Explaining the paradox as “small samples”. It appears with millions of users; it is about weights, not noise.
  • Thinking an A/B test can show Simpson’s paradox between arms. Randomisation gives both arms the same expected mix, so it cannot happen between properly randomised arms. If you see it, check for a broken split or sample ratio mismatch.
  • Answering correlation questions with “correlation is not causation” and stopping. Name the specific mechanisms (confounding, reverse causation, selection) and say how you would test it.

Verify the behavior

The decisive check for the feature-X claim is an experiment that changes X usage at random. Here a nudge raises usage by 25 points in a random half of users; because X does nothing, retention should not move, even though the observational gap is still there.

vrng = random.Random(5)
def simulate_user(nudged):
    engaged = vrng.random() < 0.3
    p_use = (0.6 if engaged else 0.1) + (0.25 if nudged else 0.0)
    uses_x = vrng.random() < p_use
    retained = vrng.random() < (0.7 if engaged else 0.3)
    return uses_x, retained

control = [simulate_user(False) for _ in range(100_000)]
treated = [simulate_user(True) for _ in range(100_000)]
use_lift = fmean(u for u, _ in treated) - fmean(u for u, _ in control)
ret_lift = fmean(r for _, r in treated) - fmean(r for _, r in control)
naive = fmean(r for u, r in control if u) - fmean(r for u, r in control if not u)
print(round(use_lift, 3), round(ret_lift, 4), round(naive, 3))   # 0.251 -0.0022 0.22
assert abs(ret_lift) < 0.01 and naive > 0.15
python

Comparing the whole randomised groups (intention to treat), not just the users who adopted X, is what keeps the comparison clean.

Follow-up questions

  • When is the pooled number the right answer? When the segment variable is affected by the treatment, or when the decision is about the total under the real future mix.
  • What can you do without an experiment? Stratify or regress on measured confounders, use matching, find a natural experiment (a staggered rollout by country), or use difference-in-differences against a comparable group, and state that unmeasured confounding remains possible.
  • What is reverse causation in the feature-X example? Users who already intend to stay explore more features, so retention intent causes X usage rather than the reverse.
  • Give a famous Simpson’s paradox. The 1973 Berkeley admissions data: men were admitted at a higher overall rate, but most departments admitted women at similar or higher rates, because women applied more to competitive departments.

Interview exercise

A marketplace reports that average order value rose 5% this quarter, but average order value fell in every one of its five countries. The CFO thinks the data is wrong. Explain what could be happening, how you would show it, and what you would report.

Answer and reasoning

This is Simpson’s paradox through a mix shift: the share of orders from high-value countries grew (for example, a launch or campaign in the most expensive market), so the pooled average rose while every country’s average fell. I would show orders and average order value by country for both quarters, then a shift-share decomposition splitting the change into a within-country rate effect (negative) and a mix effect (positive, larger). I would report both pieces: the business mix improved, but something is lowering basket size inside every market, which deserves its own investigation (discounting, a change in assortment, a shift to smaller orders). Saying “AOV is up 5%” alone would hide a real problem.

Continue learning

More in Statistics & Data Science

esc