Bias-Variance Tradeoff Explained: ML Interview Guide
What bias and variance really measure, how to read validation and learning curves, and which fix to reach for when a model underfits or overfits.
Machine learning interviews: supervised and unsupervised learning, bias-variance, overfitting and regularization, evaluation metrics, trees and ensembles, feature engineering and validation.
Official reference: scikit-learn user guide
What bias and variance really measure, how to read validation and learning curves, and which fix to reach for when a model underfits or overfits.
How k-fold cross-validation works, which splitter suits grouped or time data, and how leakage turns pure noise into an 85% accurate model.
How gradient descent updates weights, how batch, stochastic and mini-batch differ, and why learning rate and feature scaling decide convergence.
How to train and evaluate a classifier with rare positives: PR-AUC over accuracy, class weights vs thresholds, and resampling without leakage.
How logistic regression turns a linear score into a probability, why it uses log loss, how to read odds ratios and what separation breaks.
How to spot overfitting, why L1 zeroes weights while L2 only shrinks them, and how scaling, early stopping and dropout fit into regularization.
Precision, recall and F1 from the confusion matrix, how the threshold trades one for the other, and how to pick one that meets a target.
Why a random forest averages away variance while gradient boosting chips away bias, with runnable code, early stopping and when to choose each.
The difference is what feedback the model learns from.
There are hybrids too: semi-supervised learning uses a few labels plus many unlabelled rows, and self-supervised learning invents labels from the data itself, which is how language models are pretrained on next-token prediction.
Likely follow-up: Is anomaly detection supervised or unsupervised? · Why is reinforcement learning harder to evaluate offline?
Each split answers a different question. The training set fits the parameters. The validation set (or cross-validation on the training data) is used to choose between models, features and hyperparameters. The test set is touched once at the end to estimate how the chosen model performs on unseen data.
With only train and test, people tune on the test set. Every time you look at a test score and change something, information leaks from the test set into your choices, so the final number is optimistically biased. With enough tuning you can overfit the test set itself.
The split must mirror production: split by time for forecasting, by user or patient when rows from the same entity are correlated, and stratify by class when labels are rare.
Likely follow-up: How big should the test set be? · When would you use cross-validation instead of a fixed validation set?
Expected error on new data splits into bias squared, variance and irreducible noise. Bias is error from wrong assumptions: a straight line fitted to a curve misses in the same way whatever sample you train on. Variance is sensitivity to the particular sample: a very deep tree changes a lot when you swap a few rows. Making a model more flexible usually lowers bias and raises variance.
Learning curves show which one you have. High bias: training and validation error are both high and close together; more data does not help, so add features, reduce regularization or use a more flexible model. High variance: training error is low but validation error is much higher; the gap shrinks with more data, stronger regularization, simpler models or ensembling such as bagging.
Likely follow-up: Why does bagging reduce variance but not bias? · Where does the double descent phenomenon fit in?
Overfitting is when a model learns noise and quirks of the training sample instead of the pattern, so it scores well on training data and poorly on new data. You detect it by comparing training and validation scores: a large, persistent gap is the signal, and a validation loss that starts rising while training loss keeps falling is the classic picture for iterative models.
My first fixes, roughly in order:
Before any of that I would check for leakage, because a too-good validation score can hide a broken split.
Likely follow-up: Can a model overfit the validation set? · What does underfitting look like?
Both add a penalty on weight size to the loss. L2 (ridge) adds the sum of squared weights; it shrinks all weights smoothly toward zero but rarely makes any exactly zero, and it handles correlated features by spreading weight across them. L1 (lasso) adds the sum of absolute weights; it drives many weights to exactly zero, so it doubles as feature selection.
The sparsity comes from geometry and the gradient. The L1 penalty has a constant slope of size lambda however small the weight is, so a weight whose contribution to the loss is weaker than that is pushed all the way to zero. The L2 gradient shrinks with the weight, so the push fades near zero. Geometrically the L1 constraint region is a diamond whose corners lie on the axes.
Elastic net mixes both. Features must be scaled first, or the penalty punishes features just for their units.
Likely follow-up: How does lasso behave with two highly correlated features? · What is the Bayesian interpretation of L1 and L2?
K-fold splits the training data into k folds. The model is trained k times, each time on k minus 1 folds and scored on the held-out fold, and you report the mean and spread of the k scores. Every row is used for validation exactly once, so the estimate is less noisy than a single split.
The variant depends on what makes rows dependent:
cv=5.All preprocessing must happen inside each fold, which is why I put it in a Pipeline.
Likely follow-up: What is nested cross-validation for? · Why is leave-one-out rarely worth it?
Leakage is when information that will not be available at prediction time reaches the model during training or evaluation, so offline scores look better than production ever will.
Three common forms:
refund_issued when predicting fraud, or a field filled in after the outcome.To catch it I look for scores that are too good, check which features dominate importance, ask for each feature when it becomes known, compare offline metrics with a shadow deployment, and keep all fitted preprocessing inside a Pipeline evaluated with the right splitter.
Likely follow-up: How can target encoding leak, and how do you prevent it? · Is deduplication part of leakage prevention?
First I change the evaluation, because 99.5% accuracy is what predicting "never fraud" scores. I use precision, recall, PR-AUC and a cost-based view of the confusion matrix, with a stratified split and enough positives in the test set to make the numbers stable.
Then training options, cheapest first:
class_weight='balanced') so mistakes on positives cost more.If I resample or reweight, predicted probabilities no longer match real-world rates, so I recalibrate if downstream code uses them as probabilities. Often threshold tuning alone is enough.
Likely follow-up: Why does oversampling before the split inflate scores? · When would you use anomaly detection instead?
From the confusion matrix: precision is TP / (TP + FP), the share of predicted positives that are real. Recall is TP / (TP + FN), the share of real positives you caught. F1 is their harmonic mean, 2PR / (P + R), which is low if either one is low.
The choice depends on which error costs more. Optimise recall when missing a positive is expensive: cancer screening, fraud you must review, security alerts that feed a human queue. Optimise precision when false alarms are expensive: auto-blocking a customer payment, sending a sales team after leads, or a spam filter that hides real email.
In practice you fix one and maximise the other, for example the best precision at 90% recall, and you pick the threshold to get there.
Likely follow-up: Why the harmonic mean and not the arithmetic mean? · What is F-beta?
The ROC curve plots true positive rate against false positive rate across all thresholds. ROC-AUC equals the probability that a random positive gets a higher score than a random negative, so it measures ranking quality and ignores calibration and the threshold.
On heavily imbalanced data the false positive rate has a huge denominator. A model that produces 1,000 false alarms out of 1,000,000 negatives still has an FPR of 0.1%, so ROC-AUC can look excellent while most alerts are wrong. The precision-recall curve uses precision instead, which counts false positives against the alerts you raise, so PR-AUC (average precision) shows that pain. Its baseline is the positive rate, not 0.5.
I report ROC-AUC for overall ranking, PR-AUC when positives are rare, and then a precision and recall at the operating threshold.
Likely follow-up: What does an AUC of 0.3 tell you? · Is ROC-AUC affected by multiplying all scores by 2?
MAE is the mean absolute error: every unit of error counts the same, so it is robust to outliers and easy to explain ("we are off by 4.8 units on average"). Minimising it targets the median.
RMSE squares errors before averaging, so large errors dominate. Use it when big misses are disproportionately costly, such as under-forecasting demand for critical stock. Minimising squared error targets the mean. RMSE is always at least MAE, and a big gap between them signals a few large errors.
R-squared is the share of variance explained relative to predicting the mean. It is unitless, which helps compare across targets, but it can be negative on test data and says nothing about whether errors are acceptable in business terms. For targets spanning orders of magnitude I also look at MAPE or error on a log scale.
Likely follow-up: Why is MAPE dangerous when targets are near zero?
Log loss (binary cross-entropy) is the average of minus log of the probability the model gave to the true class. Predicting 0.9 for a true positive costs about 0.105; predicting 0.01 costs about 4.6. A constant 0.5 on a binary problem scores ln 2, about 0.693, which is a useful baseline.
It is the negative log-likelihood of a Bernoulli model, so minimising it gives maximum-likelihood estimates, and it is smooth and differentiable, which gradient-based training needs. Accuracy is a step function of the threshold and gives no gradient.
Accuracy only checks which side of 0.5 you land on. Log loss rewards calibrated, confident probabilities and punishes confident mistakes hard, so one prediction of 0.0001 for a real positive can dominate the average. I use it when probabilities feed decisions such as pricing or ranking, and I clip predictions away from 0 and 1.
Likely follow-up: What is the Brier score? · How do you calibrate a model?
Ordinary least squares assumes: a linear relationship between features and target (after any transformations you apply), independent errors, constant error variance (homoscedasticity), roughly normal errors for exact confidence intervals, and no perfect multicollinearity.
For pure prediction, linearity matters most: if the relationship is curved, the model is biased whatever else you do, and you fix it with transformed or interaction features. Correlated errors and changing variance mostly break the standard errors and p-values, not the point predictions. Multicollinearity makes individual coefficients unstable and hard to interpret but often barely changes predictions; ridge regression helps.
If the task is inference, "does price affect demand", all of them matter, and I check residual plots, use robust standard errors and look at variance inflation factors.
Likely follow-up: How do you read a residual plot? · Why is the closed-form solution not used for very large feature counts?
Logistic regression computes a linear score z = w.x + b and passes it through the sigmoid, 1 / (1 + e^-z), to get a probability between 0 and 1. It is a regression on the log-odds: log(p / (1 - p)) is modelled as a linear function of the features. That is where the name comes from; it only becomes a classifier when you apply a threshold.
It is trained by minimising log loss, which is convex here, so gradient-based solvers find the global optimum. There is no closed-form solution, unlike linear regression.
Interpretation is a strength: a weight of 0.7 means one unit of that feature multiplies the odds by e^0.7, about 2. The decision boundary is linear in the features, so curved boundaries need engineered features. In scikit-learn it is L2-regularized by default with C=1.0.
Likely follow-up: Why not train logistic regression with squared error? · How does it extend to more than two classes?
All three update weights by stepping against the gradient of the loss: w = w - lr * grad. They differ in how much data computes each gradient.
Too high a learning rate overshoots, oscillates or diverges with the loss climbing to infinity or NaN. Too low converges painfully slowly or stalls on plateaus. The fixes are learning-rate schedules, adaptive optimizers such as Adam, and feature scaling so one learning rate suits every direction.
Likely follow-up: Why does feature scaling speed up gradient descent? · What does momentum add?
A tree is built greedily from the root. At each node it tries every feature and candidate threshold and picks the split that most reduces impurity in the children, weighted by their size. For classification impurity is usually Gini (1 minus the sum of squared class shares) or entropy, where the reduction is called information gain. For regression it is variance, the mean squared error.
Splitting stops when a node is pure or a limit is hit. Without limits, the tree keeps splitting until each leaf holds one class, often a single row, so it memorises noise: training accuracy near 100% and high variance on new data.
You control it with max_depth, min_samples_leaf, min_samples_split, cost-complexity pruning (ccp_alpha), or by averaging many trees in a random forest. Trees need no feature scaling and handle interactions naturally.
Likely follow-up: Gini or entropy: does the choice matter much? · Why are trees unstable to small data changes?
A random forest trains many deep trees, each on a bootstrap sample of the rows, and at each split considers only a random subset of features (max_features, often the square root of the feature count for classification). Predictions are averaged or voted.
Each deep tree has low bias and high variance. Averaging reduces variance, but only if the trees make different errors. Bootstrapping alone leaves trees correlated because they all split on the same strong feature first; random feature subsets decorrelate them, so the average is much more stable than any single tree.
Useful side effects: out-of-bag rows give a free validation estimate, adding trees does not cause overfitting (it just stops helping), and it works well with little tuning. Costs are model size, slower prediction and weaker extrapolation outside the training range.
Likely follow-up: Why can a random forest not extrapolate a trend? · What does max_features control?
Bagging trains models independently and in parallel on bootstrap samples, then averages them. It mainly reduces variance, so it suits deep, low-bias trees. Random forest is the standard example.
Boosting trains models sequentially, each one fitting the mistakes of the ensemble so far. Gradient boosting fits each small tree to the negative gradient of the loss (the residuals for squared error) and adds it with a learning rate. It mainly reduces bias, using shallow trees, and usually wins on tabular accuracy.
I would pick a random forest when I want a strong baseline with little tuning, noisy labels, or parallel training. I pick gradient boosting (XGBoost, LightGBM, CatBoost or scikit-learn's HistGradientBoostingClassifier) when accuracy matters and I can tune learning rate, depth and number of rounds with early stopping, because boosting can overfit if you keep adding trees.
Likely follow-up: Why does a lower learning rate need more trees?
Both are gradient-boosted trees with engineering and regularization on top.
gamma), handles missing values by learning a default direction, and offers a fast histogram method.num_leaves.I tune in this order: a moderate learning rate with early stopping to pick the number of rounds, then tree complexity (max_depth or num_leaves, min_child_weight or min_data_in_leaf), then row and column subsampling, then the L1 and L2 penalties. Finally I lower the learning rate and retrain.
num_leaves limitsLikely follow-up: What does CatBoost do differently with categorical features?
A linear SVM finds the hyperplane that separates the classes with the maximum margin, the widest gap to the nearest points. Only those nearest points, the support vectors, define the boundary. The soft-margin version allows some points inside the margin or misclassified, with hinge loss as the penalty.
C sets that tradeoff. Large C punishes violations heavily: a narrow margin that fits the training data closely, lower bias and higher variance. Small C allows more violations for a wider, smoother margin.
The kernel trick replaces dot products with a kernel function, such as RBF or polynomial, which equals a dot product in a higher-dimensional feature space without computing that space. That lets the SVM draw curved boundaries. For RBF, gamma controls how local each point's influence is. SVMs need scaled features and scale poorly beyond tens of thousands of rows.
Likely follow-up: Why do SVMs not output probabilities directly? · What does a very large gamma do?
k-NN stores the training set. To predict, it finds the k closest training points by a distance such as Euclidean and takes a majority vote (classification) or the mean (regression), optionally weighted by distance. There is no training step, so prediction does all the work.
k is the bias-variance dial: k = 1 follows every noisy point (high variance), a large k smooths toward the overall majority (high bias). I choose it by cross-validation and prefer an odd k for binary problems to avoid ties.
It struggles when:
Likely follow-up: How does approximate nearest-neighbour search relate to vector databases?
k-means alternates two steps: assign every point to its nearest centroid, then move each centroid to the mean of its points. It repeats until assignments stop changing, which minimises within-cluster sum of squares (inertia). It only finds a local optimum, so scikit-learn uses k-means++ initialisation and several restarts.
To choose k, I look at the elbow in inertia versus k (inertia always falls as k grows, so you look for diminishing returns), the silhouette score, and above all whether the clusters are useful for the business question.
It is the wrong tool when clusters are not roughly spherical and similar in size, when there are strong outliers (means get dragged), or for categorical data. Alternatives: DBSCAN or HDBSCAN for arbitrary shapes and noise, Gaussian mixtures for elliptical clusters with soft assignments, k-modes for categories. Always scale features first.
Likely follow-up: Why does inertia always decrease as k increases?
PCA finds orthogonal directions, the principal components, along which the centred data has the most variance. They are the eigenvectors of the covariance matrix (computed in practice with an SVD), and the eigenvalues tell you how much variance each one explains. Projecting onto the top components reduces dimensions while keeping as much variance as possible.
Pitfalls:
Likely follow-up: How do you choose the number of components?
Naive Bayes applies Bayes' theorem, P(class given features) proportional to P(features given class) times P(class), and makes the naive assumption that features are conditionally independent given the class. Then the likelihood is just the product of per-feature probabilities, which you estimate by counting.
The assumption is almost always false; in text, "New" and "York" are clearly dependent. It still works because classification only needs the right class to score highest, not accurate probabilities. The model is very fast, needs little data, and handles thousands of word features easily, which is why it is a strong baseline for spam and topic classification.
Practical details: use Laplace smoothing (alpha) so an unseen word does not zero the product, compute in log space to avoid underflow, and do not trust its probabilities without calibration because they tend to be pushed to extremes.
Likely follow-up: Multinomial versus Bernoulli versus Gaussian naive Bayes?
Scaling matters wherever a model uses distances, dot products with a penalty, or gradient descent:
Tree-based models (decision trees, random forests, gradient boosting) do not need it: a split threshold is unaffected by any monotonic rescaling.
Standardization (zero mean, unit variance) is my default. Min-max to [0, 1] suits bounded inputs like pixels but is sensitive to outliers, as is the mean; RobustScaler uses the median and IQR. Whichever I pick, it is fitted on training data only.
Likely follow-up: Does scaling change the predictions of an unregularized linear regression?
One-hot encoding works for a few dozen categories but creates 50,000 sparse columns here, and rare merchants get unreliable weights. My options:
TargetEncoder does cross-fitting inside fit_transform.I also plan for categories that appear only in production, with an explicit unknown bucket.
Likely follow-up: Why is label encoding risky for linear models?
Hyperparameters are set before training (tree depth, learning rate, regularization strength), so I choose them by validation performance, usually with cross-validation.
Rules I follow: tune on validation folds and report on an untouched test set, use nested CV when I need an unbiased estimate of the whole tuning procedure, use early stopping for boosting rounds, and stop when gains are within the noise across folds.
Likely follow-up: What is successive halving? · Why sample the learning rate on a log scale?
I separate global explanations (what drives the model overall) from local ones (why this customer got this score).
Globally I avoid the default impurity-based importances, which are computed on training data and favour high-cardinality and continuous features. Permutation importance on a held-out set is more honest: shuffle one feature and measure the drop in score, though correlated features share and hide each other's importance. Partial dependence plots show how the prediction changes with a feature.
Locally I use SHAP values, which split one prediction into additive feature contributions that sum to the difference from the average prediction; TreeSHAP computes them exactly and quickly for tree ensembles.
I also say what these are not: they explain the model, not causation in the world. In regulated settings such as credit, I might prefer a monotonic-constrained model or a scorecard so explanations are faithful by construction.
Likely follow-up: How do correlated features distort permutation importance?
I would work through the usual causes in order of likelihood.
The lasting fixes are a shared feature pipeline, out-of-time validation, shadow deployment before launch and monitoring of inputs and outcomes.
Likely follow-up: How would you detect drift without labels?
The threshold turns scores into decisions, and the right one depends on the cost of each error and the base rate, not on the model. If a missed fraud costs 50 times a false alarm, the cost-minimising threshold on a calibrated probability is about 1 / 51, far below 0.5. I choose thresholds on a validation set, either by minimising expected cost or by meeting a constraint such as "at least 90% recall", and check them on the test set.
A model is calibrated when its probabilities match observed frequencies: of the cases scored 0.2, about 20% are positive. You check it with a reliability diagram or the Brier score. Boosted trees, SVMs and naive Bayes are often miscalibrated, and class weights or resampling deliberately distort probabilities. CalibratedClassifierCV with sigmoid (Platt) or isotonic calibration fixes that, on data not used for training. Calibration matters whenever probabilities are used as numbers, such as expected-loss calculations.
Likely follow-up: Does calibration change ROC-AUC?
No questions match that filter.
Prefer multiple choice? All 20 Machine Learning MCQs with answers →