Interviewers ask “how does cross-validation work?” expecting a two-sentence answer, then follow it with the question that actually matters: “your model scores 0.95 in cross-validation and 0.70 in production, what happened?” Almost always the answer is leakage, and leakage almost always comes from a validation procedure that did not mirror how the model will be used. This guide connects the two topics, because in practice they are one topic: cross-validation is only as honest as the split and the preprocessing inside it.
Before you start
You should know what a train and test split is and have trained at least one scikit-learn model. The examples use Python with NumPy 2.5 and scikit-learn 1.9; outputs in comments are from real runs with fixed seeds. Two terms come up repeatedly: a fold is one chunk of the data used for validation, and a pipeline is a chain of preprocessing steps and a model that is fitted as one object.
The short answer
K-fold cross-validation splits the training data into k folds, trains k models, each on k minus 1 folds, and scores each on the fold it did not see; you report the mean and standard deviation. Use StratifiedKFold for classification, GroupKFold when several rows belong to the same user or patient, and TimeSeriesSplit when the future must not inform the past. Data leakage is any path by which information unavailable at prediction time reaches training or evaluation: target-derived features, preprocessing fitted on all rows, or a split that puts related rows on both sides. It inflates offline scores, and the standard defence is a Pipeline evaluated with the right splitter.
How it works
A single train and validation split gives one noisy number. Cross-validation reuses the data so every row is validated exactly once. Writing the loop by hand shows what cross_val_score does:
import numpy as np
from sklearn.base import clone
from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import StratifiedKFold, cross_val_score
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
X, y = make_classification(n_samples=1000, n_features=10, n_informative=4,
n_redundant=2, random_state=3)
model = make_pipeline(StandardScaler(), LogisticRegression())
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=0)
scores = []
for train_idx, val_idx in cv.split(X, y):
m = clone(model).fit(X[train_idx], y[train_idx]) # fresh, unfitted copy per fold
scores.append(m.score(X[val_idx], y[val_idx]))
print(np.round(scores, 3), round(np.mean(scores), 3), round(np.std(scores), 3))
# [0.91 0.88 0.89 0.93 0.915] 0.905 0.018
print(cross_val_score(model, X, y, cv=cv).round(3))
# [0.91 0.88 0.89 0.93 0.915]The essential detail is clone: each fold gets a brand new model, including a brand new scaler, fitted only on that fold’s training rows. The spread across folds (0.88 to 0.93) is as informative as the mean: a model that is 0.905 plus or minus 0.018 is not meaningfully better than one at 0.900.
Leakage is what happens when that isolation breaks. It has three common shapes. Target leakage is a feature that encodes the answer, such as account_closed_date in a churn model or chargeback_filed in a fraud model; it is only known after the outcome. Preprocessing contamination is fitting a scaler, imputer, feature selector, PCA, target encoder or oversampler on all rows and then splitting, so statistics from the validation rows flow into training. Split leakage is choosing folds that do not respect how rows are related: several rows per patient, near-duplicates, or random splits of time-ordered data.
The contamination case can manufacture signal from nothing. Here the labels are random and all 5,000 features are noise, so the honest accuracy is 50%:
from sklearn.feature_selection import SelectKBest, f_classif
rng = np.random.default_rng(0)
X = rng.normal(size=(100, 5000))
y = rng.integers(0, 2, size=100)
X_sel = SelectKBest(f_classif, k=20).fit_transform(X, y) # sees every label
print(round(cross_val_score(LogisticRegression(), X_sel, y, cv=5).mean(), 2)) # 0.85
pipe = make_pipeline(SelectKBest(f_classif, k=20), LogisticRegression())
print(round(cross_val_score(pipe, X, y, cv=5).mean(), 2)) # 0.48With 5,000 noise columns and 100 rows, some columns correlate with the labels by chance. Selecting them on all 100 rows picks the ones that happen to match the validation labels too, so every fold looks good. Inside the pipeline, selection is redone on each training fold and the chance correlations do not carry over.
Step-by-step walkthrough
Step 1: Hold out a test set before anything else
Split off a final test set first and do not look at it while you iterate. Cross-validation on the remaining data is for choosing models and hyperparameters. The test set answers one question at the end: how well does the chosen procedure do on data it never influenced? If you tune against the test set, it becomes a second validation set and its score becomes optimistic.
Step 2: Choose the splitter that copies production
Ask how the model will meet new data, and make the folds do the same:
- New rows from the same population, balanced or not:
StratifiedKFold(shuffle=True), which keeps class ratios equal across folds. An integercv=5with a classifier already meansStratifiedKFold, without shuffling. - New users, patients, stores or devices that the model has never seen:
GroupKFoldwith the entity id asgroups. - Future periods:
TimeSeriesSplit, which always trains on earlier rows and validates on later ones, optionally with agapso features computed over a trailing window cannot overlap the validation period.
from sklearn.model_selection import TimeSeriesSplit
for train, test in TimeSeriesSplit(n_splits=3).split(list(range(6))):
print(train.tolist(), test.tolist())
# [0, 1, 2] [3]
# [0, 1, 2, 3] [4]
# [0, 1, 2, 3, 4] [5]Step 3: Put every fitted step in a pipeline
Anything with a fit method that learns from data belongs inside the pipeline: imputers, scalers, encoders, feature selection, dimensionality reduction and resampling (imbalanced-learn provides a pipeline that accepts samplers). Then cross_val_score, GridSearchCV and the final fit all refit those steps on training rows only. Stateless transforms, such as taking a log of a column, can safely happen before the split.
Step 4: Audit features for when they become known
For each feature, ask: at the moment this prediction is made in production, does this value exist, and does it have this value? Fields that are updated after the outcome, aggregates computed over the whole dataset (such as “average spend of this customer” that includes future purchases), and identifiers that correlate with collection time are classic leaks. A feature importance chart dominated by one surprising feature is often the first visible symptom.
Worked scenario
A team classifies medical scans. They have 120 patients with 8 scans each, and the diagnosis is per patient. Each patient’s scans share a recognisable “look” (scanner settings, anatomy) that has nothing to do with the diagnosis. They use shuffled k-fold:
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import GroupKFold, KFold
rng = np.random.default_rng(0)
n_patients, scans_each = 120, 8
patient = np.repeat(np.arange(n_patients), scans_each)
signature = rng.normal(size=(n_patients, 6))
label_per_patient = rng.integers(0, 2, n_patients)
X = signature[patient] + rng.normal(scale=0.3, size=(patient.size, 6))
y = label_per_patient[patient]
rf = RandomForestClassifier(n_estimators=200, random_state=0)
print(cross_val_score(rf, X, y, cv=KFold(5, shuffle=True, random_state=0)).mean().round(3))
# 0.849An 85% accurate diagnostic model would be exciting, but the features carry no diagnostic information at all. With random folds, almost every validation scan has sibling scans from the same patient in training, so the forest recognises the patient and copies their label. In production the model meets new patients and that trick fails. Grouping by patient gives the honest answer:
print(cross_val_score(rf, X, y, cv=GroupKFold(5), groups=patient).mean().round(3))
# 0.51The same pattern appears with multiple sessions per user, multiple transactions per card, several photos of one product and augmented copies of one image. Deduplicate near-identical rows before splitting as well, because duplicates across folds leak the same way.
Common mistake
- “I scaled the data first, but scaling cannot leak.” It leaks a little (validation means and variances), and the habit leaks a lot once the step is a selector, an encoder or an oversampler.
- Shuffling time series. A random split lets the model learn from next month to predict last month, and seasonal or trend features look far more predictive than they are.
- Reporting the best fold. Report the mean and spread; the best fold is luck.
- Reusing the cross-validation score after tuning as the final estimate. The score of the best of 200 configurations is biased upward; use the held-out test set or nested cross-validation.
- Oversampling before splitting. SMOTE or duplicated positives end up on both sides of a fold and the model is validated on copies of its training rows.
Verify the behavior
Make the grouping guarantee an assertion so a refactor cannot silently break it:
def assert_no_group_overlap(cv, X, y, groups):
for i, (tr, va) in enumerate(cv.split(X, y, groups)):
shared = set(groups[tr]) & set(groups[va])
assert not shared, f"fold {i}: {len(shared)} groups in both train and validation"
print("ok:", cv)
assert_no_group_overlap(GroupKFold(5), X, y, patient)
# ok: GroupKFold(n_splits=5, random_state=None, shuffle=False)
assert_no_group_overlap(KFold(5, shuffle=True, random_state=0), X, y, patient)
# AssertionError: fold 0: 98 groups in both train and validationKFold also emits a UserWarning that it ignores the groups argument, which is itself a useful hint. A second cheap test is the noise test from earlier: shuffle the labels, rerun your full pipeline, and assert the score drops to chance. If it does not, something is leaking.
Follow-up questions
- What is nested cross-validation? An inner loop tunes hyperparameters and an outer loop scores the whole tuning procedure, so the reported number is not biased by the search.
- How many folds? Five or ten is standard. More folds mean larger training sets and less bias, but more compute and more correlated estimates; leave-one-out is rarely worth it.
- How does target encoding leak, and how is it fixed? Each row’s encoding includes its own label. Compute encodings out of fold; scikit-learn’s
TargetEncodercross-fits insidefit_transform. - How would you validate a model that will be retrained weekly? Walk-forward validation: train on data up to week t, validate on week t + 1, slide forward, and report the distribution of weekly scores.
Interview exercise
You predict whether a loan will default within 12 months. Features include income, credit score at application, number of late payments (taken from the current customer record), and branch id. Rows span 2019 to 2025. A random 5-fold cross-validation gives ROC-AUC 0.94. List what you would check and how you would re-validate.
Answer and reasoning
The “number of late payments” field is taken from the current record, so for old loans it includes payments missed after the application, which is target leakage: defaulters accumulate late payments. I would rebuild it as of the application date, or drop it. Branch id could be fine, but if branches opened or closed at particular times it can proxy for the period. Then I would replace random folds with a time-based scheme: train on 2019 to 2022, validate on 2023, and so on, with a 12-month gap because a loan’s label is only known a year after application. Repeated applications by the same customer call for grouping by customer as well. I would expect the AUC to drop, and the new, lower number is the one to report.
Continue learning
- Practise in the Machine Learning chapter and the machine learning MCQs.
- Read why the score you validate on matters in precision, recall and F1, and how to resample safely in imbalanced classification.
- Reference: scikit-learn’s cross-validation guide and common pitfalls page, and Google’s guide to dividing datasets.