“Your dataset has 1% positives. How do you build a model for it?” Fraud, churn, defects, rare diseases and click prediction all look like this, so the question appears in almost every applied machine learning interview. Weak answers jump straight to “use SMOTE”. Strong answers start with evaluation, because most imbalance disasters are measurement failures: a model that never predicts the rare class scores 99% accuracy, and a resampling step placed before the split can produce perfect validation scores for a mediocre model. This guide works through both, then compares the training-side fixes on the same data.
Before you start
You should know precision, recall and the confusion matrix (see precision vs recall and F1), and what cross-validation does. The code uses Python 3.14 with scikit-learn 1.9 and NumPy 2.5. The dataset throughout is make_classification(n_samples=50_000, n_features=20, n_informative=5, weights=[0.99], flip_y=0, class_sep=1.0, random_state=3): 500 positives among 50,000 rows. Outputs come from real runs.
The short answer
First fix the evaluation: use a stratified split, report precision and recall at the operating threshold plus PR-AUC (average precision), and make sure the test set holds enough positives for stable numbers. Then train, cheapest option first: fit a model that ranks well and tune the decision threshold on validation data; add class weights if the learner needs a push towards the rare class; resample (undersample negatives, or oversample positives with duplication or SMOTE) only inside the training folds. Reweighting and resampling shift predicted probabilities, so recalibrate if anything downstream uses them as probabilities.
How it works
Imbalance breaks the familiar metrics in two different ways. Accuracy rewards the majority class: predicting “negative” for everyone is right 99% of the time. ROC-AUC is subtler. The ROC curve plots recall against the false positive rate, and with 49,500 negatives the false positive rate barely moves even when false alarms outnumber real cases. Precision counts false positives against the alerts you actually raise, so the precision-recall curve exposes that cost. The two summaries also have different baselines:
import numpy as np
from sklearn.metrics import roc_auc_score, average_precision_score
r = np.random.default_rng(0).random(len(y_te)) # scores with no information
print(round(roc_auc_score(y_te, r), 3), round(average_precision_score(y_te, r), 4))
# 0.51 0.0105A random model scores about 0.5 on ROC-AUC regardless of balance, but its average precision equals the positive rate, here 1%. A PR-AUC of 0.6 is therefore sixty times better than chance, while a ROC-AUC of 0.96 can sit alongside a model whose alerts are mostly wrong.
On the training side, a learner minimising average loss sees 99 negatives for each positive, so its scores lean low and the default 0.5 threshold flags few positives. Class weights multiply each positive’s loss (with "balanced", by n_samples / (2 · n_positives), about 50 here), which pushes scores up. Resampling changes the training rows themselves. Both mostly move where the scores sit relative to 0.5; neither adds information the features did not already contain.
Step-by-step walkthrough
Step 1: Split with stratification and measure the baseline
from sklearn.model_selection import train_test_split
from sklearn.dummy import DummyClassifier
from sklearn.metrics import accuracy_score
X_tmp, X_te, y_tmp, y_te = train_test_split(X, y, test_size=0.3, stratify=y, random_state=0)
X_tr, X_val, y_tr, y_val = train_test_split(X_tmp, y_tmp, test_size=0.25, stratify=y_tmp, random_state=0)
print(y_tr.sum(), y_val.sum(), y_te.sum()) # 263 87 150
dummy = DummyClassifier(strategy="most_frequent").fit(X_tr, y_tr)
print(accuracy_score(y_te, dummy.predict(X_te))) # 0.99Stratification guarantees each split keeps the 1% rate; an unstratified split of a small dataset can leave a validation set with a handful of positives. The dummy’s 0.99 accuracy is the number any real model must be judged against, and it shows why accuracy is not on the scorecard.
Step 2: Train a plain model and read the right metrics
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.metrics import precision_score, recall_score
plain = HistGradientBoostingClassifier(random_state=0).fit(X_tr, y_tr)
p = plain.predict_proba(X_te)[:, 1]
pred = (p >= 0.5).astype(int)
# accuracy 0.9933, precision 0.775, recall 0.46, ROC-AUC 0.965, PR-AUC 0.647Accuracy barely beats the dummy and ROC-AUC looks excellent, yet at the default threshold the model misses more than half of the positives. PR-AUC of 0.647 is the honest summary of ranking quality, and precision 0.775 with recall 0.46 is what the business would actually experience.
Step 3: Compare class weights with threshold tuning
weighted = HistGradientBoostingClassifier(class_weight="balanced", random_state=0).fit(X_tr, y_tr)
# at threshold 0.5: precision 0.348, recall 0.767, ROC-AUC 0.969, PR-AUC 0.621
# mean predicted probability 0.0757 versus an actual positive rate of 0.01The weighted model catches far more positives at 0.5, but its ranking metrics are almost unchanged: ROC-AUC rose by 0.004 and PR-AUC fell by 0.026. The weights mainly slid every score upwards. To prove it, set a recall target of 0.8 and choose each model’s threshold on the validation set with precision_recall_curve:
# plain model: threshold 0.006 -> test precision 0.222, recall 0.827
# weighted model: threshold 0.316 -> test precision 0.248, recall 0.827At the same operating point the two models are close. Threshold tuning on the unweighted model got nearly the same result with no retraining, and it kept probabilities that average 0.0078 against a true rate of 0.01, whereas the weighted model overstates risk about sevenfold. Class weights still earn their place when a learner struggles to separate the minority class at all, or when you need predict() itself to behave sensibly, but they are a threshold shift more than a new model.
Step 4: Undersample for speed, then correct the probabilities
With millions of negatives, keeping a random fraction β of them makes training much cheaper and usually costs little ranking quality. It inflates scores in a predictable way, which can be undone exactly:
beta = 0.1 # kept 10% of negatives in training
p = beta * p_s / (beta * p_s - p_s + 1) # p_s: score from the undersampled modelOn a 200,000-row version of the dataset with logistic regression, the undersampled model’s mean probability was 0.0719 against a true rate of 0.01; the corrected mean was 0.0095 and its log loss (0.0462) matched a model trained on all rows (0.0461). The correction is monotonic, so ROC-AUC and PR-AUC do not change at all.
Worked scenario
A team oversamples fraud cases 20 times by duplication, then runs 5-fold cross-validation with a random forest and reports recall 1.0, precision 0.998 and PR-AUC 1.0. Production recall turns out to be about 0.4.
# oversample BEFORE cross-validation: recall 1.000, precision 0.998, PR-AUC 1.000
# oversample INSIDE each training fold: recall 0.429, precision 0.956, PR-AUC 0.727
# held-out test set: recall 0.413, precision 0.899, PR-AUC 0.733Oversampling before splitting puts copies of the same fraud case into both the training and validation folds. Deep trees memorise the copy they trained on and “recognise” its twin, so validation measures memory, not generalisation. The validation folds also carry the oversampled base rate, which inflates precision further. Resampling inside each training fold, after the split, gives 0.429 recall, which the untouched test set confirms.
The structural fix is to make resampling part of the model so cross-validation cannot get it wrong. imbalanced-learn provides a Pipeline that applies samplers only during fit:
from imblearn.pipeline import make_pipeline
from imblearn.over_sampling import SMOTE
from sklearn.model_selection import cross_validate, StratifiedKFold
model = make_pipeline(SMOTE(random_state=0), HistGradientBoostingClassifier(random_state=0))
scores = cross_validate(model, X_tr, y_tr, scoring=["average_precision", "recall"],
cv=StratifiedKFold(5, shuffle=True, random_state=0))SMOTE creates synthetic positives by interpolating between neighbouring positives rather than copying them, which reduces exact-duplicate memorisation, but it has the same leakage problem if run before the split.
Common mistake
- Leading with accuracy or ROC-AUC alone. Both flatter rare-class models; add PR-AUC and the precision and recall you will operate at.
- Resampling the validation or test set. Evaluation data must keep the real-world base rate, or precision becomes meaningless.
- “SMOTE always helps.” On strong learners such as gradient boosting it often matches or underperforms threshold tuning, and on high-dimensional or categorical data its interpolated points can be unrealistic.
- Using weighted-model probabilities as risk estimates. They are inflated by design; correct them or calibrate with
CalibratedClassifierCVon untouched data. - Too few positives in validation. With 30 positives, one extra catch moves recall by over three points. Use repeated stratified cross-validation or report confidence intervals.
Verify the behavior
Two quick checks make the claims above testable; run them on the Step 4 data and model. Random scores should produce an average precision close to the positive rate, and the undersampling correction should bring the mean prediction back to the base rate without changing the ranking:
r = np.random.default_rng(0).random(len(y_te))
assert abs(average_precision_score(y_te, r) - y_te.mean()) < 0.005
p = beta * p_s / (beta * p_s - p_s + 1)
assert abs(p.mean() - y_te.mean()) < 0.002
assert round(roc_auc_score(y_te, p_s), 6) == round(roc_auc_score(y_te, p), 6)
print("baseline and correction behave as expected")Follow-up questions
- When is ROC-AUC still useful? For comparing ranking quality across datasets with different base rates, since it ignores prevalence; just never report it alone when positives are rare.
- How do you pick the threshold from costs? With calibrated probabilities, flag when
p > C_FP / (C_FP + C_FN); otherwise sweep thresholds on validation data and minimise total cost. - What if positives are extremely rare, say 0.01%? Consider framing it as anomaly detection, collecting more labelled positives through active learning, or splitting the problem into a high-recall filter followed by a precise second stage.
- Do class weights affect logistic regression’s coefficients? Mostly the intercept, which shifts by about the log of the weight ratio; slopes change less. See logistic regression interview questions.
Interview exercise
A fraud model was trained after keeping only 10% of non-fraud transactions. For one transaction it outputs 0.3. The true fraud rate is 0.5%. What is the corrected fraud probability, and if a missed fraud costs $500 while reviewing a transaction costs $5, should this transaction be reviewed?
Answer and reasoning
Undersampling multiplied the odds of fraud by 1 / β = 10, so the correction divides the odds by 10: p = 0.1 × 0.3 / (0.1 × 0.3 - 0.3 + 1) = 0.03 / 0.73 ≈ 0.041. The raw 0.3 overstated the risk more than sevenfold. For the decision, review when the expected loss from ignoring the transaction exceeds the review cost: 0.041 × $500 ≈ $20.50 > $5, so yes, review it. Equivalently, the cost-based threshold is 5 / (5 + 500) ≈ 0.0099 on the corrected probability, and 0.041 is above it. Using the uncorrected 0.3 would give the right decision here by luck but would badly overstate expected losses in any aggregate forecast.
Continue learning
- Practise in the Machine Learning chapter and the machine learning MCQs.
- Choose the operating point with precision vs recall and F1, and keep resampling inside the folds with cross-validation and data leakage.
- Reference: Google’s imbalanced datasets module, scikit-learn’s probability calibration guide and the imbalanced-learn documentation.