Ch. 31 · MLOps

Data Drift vs Concept Drift: Detecting Drift with PSI and KS

How data drift, concept drift and prediction drift differ, how PSI and the KS test measure them, and why p-value alerts drown your team in noise.

~8 min readintermediateupdated Oct 6, 2026

“What is the difference between data drift and concept drift, and how would you detect each?” is close to a guaranteed question in an MLOps or ML engineering interview. The definitions take ten seconds. What interviewers actually probe is whether you have run monitoring in production: which statistic you would compute, what threshold you would alert on, what you do when labels arrive months late, and why a drift alert is not the same thing as a broken model.

Before you start

You should know what a trained classifier is, what a feature distribution looks like as a histogram, and roughly what a p-value means. The code uses Python 3.14 with NumPy 2.5, SciPy 1.18 and scikit-learn 1.9; every output shown in a comment came from running it with the seeds given. Two words recur: the reference window is the data you compare against (usually the training set or a recent stable period), and the current window is the production data you are checking, for example yesterday’s requests.

The short answer

Data drift is a change in the input distribution, P(X): the model receives inputs unlike the ones it was trained on. Concept drift is a change in the relationship between inputs and the target, P(y given X): the same inputs now deserve a different answer. Data drift can be measured immediately by comparing feature distributions between a reference and a current window, with the Population Stability Index (PSI) or a two-sample Kolmogorov-Smirnov (KS) test. Concept drift usually needs ground-truth labels, so until they arrive you watch prediction drift and business proxies. Any drift signal starts an investigation; it does not prove the model got worse.

How it works

PSI splits the reference distribution into bins, normally deciles, and compares the share of reference and current rows in each bin:

import numpy as np
from scipy.stats import ks_2samp

def psi(reference, current, bins=10, eps=1e-6):
    edges = np.quantile(reference, np.linspace(0, 1, bins + 1))
    edges[0], edges[-1] = -np.inf, np.inf          # catch values outside the old range
    ref_pct = np.histogram(reference, edges)[0] / len(reference)
    cur_pct = np.histogram(current, edges)[0] / len(current)
    ref_pct = np.clip(ref_pct, eps, None)           # an empty bin would give log(0)
    cur_pct = np.clip(cur_pct, eps, None)
    return float(np.sum((cur_pct - ref_pct) * np.log(cur_pct / ref_pct)))

rng = np.random.default_rng(42)
train = rng.normal(50, 10, 20_000)                  # basket value at training time
for name, cur in [("same", rng.normal(50, 10, 5_000)),
                  ("shift", rng.normal(54, 10, 5_000)),
                  ("big", rng.normal(60, 14, 5_000))]:
    ks = ks_2samp(train, cur)
    print(f"{name:5}  psi={psi(train, cur):.3f}  ks={ks.statistic:.3f}  p={ks.pvalue:.2g}")
# same   psi=0.002  ks=0.016  p=0.23
# shift  psi=0.151  ks=0.163  p=1.5e-93
# big    psi=0.689  ks=0.351  p=0
python

Quantile bins put an equal share of reference rows in every bin, so each bin starts at 10%. The widely used rule of thumb from credit risk scorecards reads PSI below 0.1 as stable, 0.1 to 0.25 as a moderate shift worth a look, and above 0.25 as a significant shift. A mean shift of 0.4 standard deviations scores 0.151. PSI works for categorical features too, with one bin per category plus an “other” bucket, and it is symmetric: swapping reference and current gives the same number.

The two-sample KS test builds the empirical cumulative distribution of each sample and reports the largest vertical gap between them as its statistic, along with a p-value for the hypothesis that both samples come from the same distribution. It only makes sense for continuous numeric features.

Concept drift is different in kind, and a small experiment shows why input monitoring cannot see it. Here the inputs come from exactly the same distribution, but the rule that generates the label moves:

from sklearn.linear_model import LogisticRegression

rng = np.random.default_rng(7)
X_old = rng.normal(0, 1, (5_000, 2)); y_old = (X_old[:, 0] > 0).astype(int)
model = LogisticRegression().fit(X_old, y_old)
X_new = rng.normal(0, 1, (5_000, 2)); y_new = (X_new[:, 0] > 0.5).astype(int)

print("psi x0:", round(psi(X_old[:, 0], X_new[:, 0]), 3))
# psi x0: 0.002
print("acc old:", round(model.score(X_old, y_old), 3), "acc new:", round(model.score(X_new, y_new), 3))
# acc old: 0.999 acc new: 0.814
python

Feature PSI is 0.002 and so is the PSI of the model’s predicted probabilities, because the model sees the same inputs and so produces the same outputs. Accuracy has still fallen from 0.999 to 0.814. Only labels reveal it: the true positive rate dropped to about 30% while the model still predicts positive for about 49% of rows.

Step-by-step walkthrough

Step 1: Log what the model saw and said

Monitoring is impossible without data. Every prediction should be logged with a prediction id, timestamp, model version, the feature values actually used and the output. The prediction id is the join key for labels that arrive later, and the logged features are what you compute drift on.

Step 2: Choose the reference window deliberately

Comparing against the training set answers “is the model working outside what it learned?”. Comparing against last week answers “did something just change?”. Both are useful, and many teams compute both. Seasonal businesses should compare like with like, or every Monday will look like drift.

Step 3: Compute per-feature drift and prediction drift

Run PSI (or KS statistics) per feature on a schedule, and run the same on the model’s score distribution. Prediction drift is the cheapest single signal: if the average predicted fraud risk doubles overnight with no policy change, something upstream moved. Rank features by importance: drift on a top feature pages someone, drift on a minor one goes to a dashboard.

Step 4: Alert on effect size and persistence

Use thresholds on PSI or the KS statistic, not on p-values, and require the breach to persist for two or three windows unless it is extreme. Add plain data quality checks beside drift: null rate, share of default values, out-of-range values and row count. Many “drift” incidents are really broken pipelines.

Step 5: Close the loop with labels

When labels arrive, compute the real metric (AUC, precision at the operating threshold, calibration) by prediction date and by segment. If labels take months, find an earlier proxy, such as a missed first payment for a 12-month default model. This is the only layer that catches pure concept drift.

Worked scenario

A payments team ran a nightly job that applied a KS test to all 60 features of a fraud model and opened a ticket for every feature with p below 0.05. Each window held about 50,000 transactions, and within a week nobody read the alerts. Reproducing the setup with features that only wobble by a harmless 0.02 standard deviations shows the cause:

rng = np.random.default_rng(1)
flags_p = flags_psi = 0
for f in range(60):
    ref = rng.normal(0, 1, 50_000)
    cur = rng.normal(rng.normal(0, 0.02), 1, 50_000)   # tiny, harmless wobble
    flags_p += ks_2samp(ref, cur).pvalue < 0.05
    flags_psi += psi(ref, cur) > 0.1
print("ks p<0.05 flags:", flags_p, "psi>0.1 flags:", flags_psi)
# ks p<0.05 flags: 39 psi>0.1 flags: 0
python

Thirty-nine of sixty stable features “drifted”. Then a mobile release started sending amounts in cents instead of dollars. Its KS alert was one more line in a channel nobody read, and the model scored inflated amounts for two days.

The fix had three parts. Alerts moved to effect sizes: PSI above 0.25 on a top-ten feature, or above 0.1 for three consecutive days. A unit bug like the cents change is not subtle on that scale:

amount_ref = rng.lognormal(3.5, 0.6, 50_000)
amount_bug = rng.lognormal(3.5, 0.6, 50_000) * 100
print("cents bug psi:", round(psi(amount_ref, amount_bug), 2), "ks:", round(ks_2samp(amount_ref, amount_bug).statistic, 3))
# cents bug psi: 12.43 ks: 1.0
python

Second, a schema check at the API rejected amounts outside a plausible range. Third, prediction drift on the fraud score was added as a single top-level alert, because a sudden shift in the score distribution is visible whichever feature caused it.

Common mistake

  • Treating p below 0.05 as “drift happened”. With large windows every trivial shift is significant; with tiny windows real shifts are missed. The p-value answers “are these samples different at all?”, not “is the difference big enough to matter?”.
  • Retraining automatically on any drift alert. Drift on a feature the model barely uses may change nothing, and retraining on a broken feed bakes the bug into the next model. Investigate first.
  • Saying data drift always lowers accuracy. If a model learned the right relationship, a shift towards a well-covered region can leave accuracy unchanged or even improve it.
  • Claiming input monitoring catches concept drift. As the experiment showed, inputs and predictions can be perfectly stable while accuracy drops.

Verify the behavior

Two quick checks prove the properties discussed above. PSI is symmetric, and an unsmoothed empty bin produces infinity:

f = lambda a, b: np.sum((b - a) * np.log(b / a))
ref, cur = np.array([0.5, 0.3, 0.2]), np.array([0.4, 0.4, 0.2])
print(round(f(ref, cur), 4), round(f(cur, ref), 4))
# 0.0511 0.0511
with np.errstate(all="ignore"):
    print(f(np.array([0.5, 0.5, 0.0]), cur))
# inf
python

To see the sample-size effect, compare two samples whose means differ by 0.01 standard deviations: with 500 rows each KS gives p of 0.56 in our run, and with 1,000,000 rows each p falls to about 1e-10 even though the statistic is only about 0.005.

Follow-up questions

  • What is label shift? A change in P(y), the class balance, such as fraud rising from 0.5% to 2% of transactions. It changes precision at a fixed threshold even if P(X given y) is stable.
  • How do you monitor high-cardinality categoricals such as merchant id? Bucket by frequency (top N plus “other”), track the share of unseen categories, and use PSI or chi-square on the buckets.
  • How would you decide whether drift justifies retraining? Evaluate the current model on recent labelled data, or retrain on recent data and compare on a fresh holdout. Retrain when the gap is real, not when the alert fires.

Interview exercise

A grocery delivery company predicts order lateness. After a city launch, PSI on distance_km is 0.31 and on courier_count is 0.02, the predicted lateness rate rose from 12% to 19%, and actual lateness labels arrive within two hours. Precision and recall measured on yesterday’s orders are unchanged. Should you retrain? What do you do next?

Answer and reasoning

Not yet. Distance drifted, as expected after launching a city with longer routes, and the model responded by predicting more lateness. The important evidence is that labels arrive quickly and measured precision and recall have not moved, so the model is handling the new region correctly: this is data drift without performance loss. I would slice the metrics by city to confirm the new city is not hidden inside a healthy average, keep watching it daily, and update the reference window for the distance feature once the launch is accepted as the new normal so the alert stops firing. Retraining becomes worthwhile if the new-city slice degrades or the business wants the model to learn city-specific patterns.

Continue learning

More in MLOps

esc