Ch. 28

Machine Learning interview questions & answers

Machine learning interviews: supervised and unsupervised learning, bias-variance, overfitting and regularization, evaluation metrics, trees and ensembles, feature engineering and validation.

30 interview questions20 quiz questions8 notes
your progress0%

Notes in this chapter

Filter all notes →

30 Machine Learning interview questions study by subtopic

30 questions
  1. 1.What is the difference between supervised, unsupervised and reinforcement learning? Give an example of each.easy

    The difference is what feedback the model learns from.

    • Supervised learning trains on labelled examples, pairs of input and known answer. Predicting house prices (regression) or flagging spam (classification) are supervised.
    • Unsupervised learning has no labels and looks for structure: clustering customers into segments with k-means, or compressing features with PCA.
    • Reinforcement learning has an agent that takes actions in an environment and receives a reward, sometimes much later. It learns a policy that maximises long-term reward, as in game playing or ad bidding.

    There are hybrids too: semi-supervised learning uses a few labels plus many unlabelled rows, and self-supervised learning invents labels from the data itself, which is how language models are pretrained on next-token prediction.

    What interviewers listen for
    • Supervised: labelled input and target pairs
    • Unsupervised: no labels, find structure
    • Reinforcement: actions, rewards, delayed feedback
    • Mentions semi- or self-supervised hybrids

    Likely follow-up: Is anomaly detection supervised or unsupervised? · Why is reinforcement learning harder to evaluate offline?

  2. 2.Why do we split data into training, validation and test sets? What goes wrong if you only have train and test?easy

    Each split answers a different question. The training set fits the parameters. The validation set (or cross-validation on the training data) is used to choose between models, features and hyperparameters. The test set is touched once at the end to estimate how the chosen model performs on unseen data.

    With only train and test, people tune on the test set. Every time you look at a test score and change something, information leaks from the test set into your choices, so the final number is optimistically biased. With enough tuning you can overfit the test set itself.

    The split must mirror production: split by time for forecasting, by user or patient when rows from the same entity are correlated, and stratify by class when labels are rare.

    What interviewers listen for
    • Train fits, validation selects, test estimates
    • Tuning on the test set biases the estimate
    • Split by time or group when rows are correlated
    • Stratify for rare classes

    Likely follow-up: How big should the test set be? · When would you use cross-validation instead of a fixed validation set?

  3. 3.Explain the bias-variance tradeoff. How do you tell from learning curves which one is hurting your model?mid

    Expected error on new data splits into bias squared, variance and irreducible noise. Bias is error from wrong assumptions: a straight line fitted to a curve misses in the same way whatever sample you train on. Variance is sensitivity to the particular sample: a very deep tree changes a lot when you swap a few rows. Making a model more flexible usually lowers bias and raises variance.

    Learning curves show which one you have. High bias: training and validation error are both high and close together; more data does not help, so add features, reduce regularization or use a more flexible model. High variance: training error is low but validation error is much higher; the gap shrinks with more data, stronger regularization, simpler models or ensembling such as bagging.

    What interviewers listen for
    • Error = bias squared + variance + noise
    • High bias: both errors high and close
    • High variance: large train-validation gap
    • Different fixes for each case

    Likely follow-up: Why does bagging reduce variance but not bias? · Where does the double descent phenomenon fit in?

  4. 4.What is overfitting, how do you detect it, and what are your first five fixes?easy

    Overfitting is when a model learns noise and quirks of the training sample instead of the pattern, so it scores well on training data and poorly on new data. You detect it by comparing training and validation scores: a large, persistent gap is the signal, and a validation loss that starts rising while training loss keeps falling is the classic picture for iterative models.

    My first fixes, roughly in order:

    • Get more or cleaner data, or use data augmentation.
    • Simplify: fewer features, shallower trees, fewer parameters.
    • Add regularization: L1 or L2 penalties, dropout in neural networks.
    • Use early stopping on a validation set.
    • Use ensembles such as bagging or random forests, which average away variance.

    Before any of that I would check for leakage, because a too-good validation score can hide a broken split.

    What interviewers listen for
    • Train score high, validation score much lower
    • Rising validation loss during training
    • More data, simpler model, regularization
    • Early stopping and ensembling

    Likely follow-up: Can a model overfit the validation set? · What does underfitting look like?

  5. 5.What is the difference between L1 and L2 regularization, and why does L1 produce sparse models?mid

    Both add a penalty on weight size to the loss. L2 (ridge) adds the sum of squared weights; it shrinks all weights smoothly toward zero but rarely makes any exactly zero, and it handles correlated features by spreading weight across them. L1 (lasso) adds the sum of absolute weights; it drives many weights to exactly zero, so it doubles as feature selection.

    The sparsity comes from geometry and the gradient. The L1 penalty has a constant slope of size lambda however small the weight is, so a weight whose contribution to the loss is weaker than that is pushed all the way to zero. The L2 gradient shrinks with the weight, so the push fades near zero. Geometrically the L1 constraint region is a diamond whose corners lie on the axes.

    Elastic net mixes both. Features must be scaled first, or the penalty punishes features just for their units.

    What interviewers listen for
    • L2 sums squares and shrinks smoothly
    • L1 sums absolute values and zeroes weights
    • Constant L1 gradient explains sparsity
    • Scale features before penalizing

    Likely follow-up: How does lasso behave with two highly correlated features? · What is the Bayesian interpretation of L1 and L2?

  6. 6.How does k-fold cross-validation work, and which variant do you pick for imbalanced, grouped or time-series data?mid

    K-fold splits the training data into k folds. The model is trained k times, each time on k minus 1 folds and scored on the held-out fold, and you report the mean and spread of the k scores. Every row is used for validation exactly once, so the estimate is less noisy than a single split.

    The variant depends on what makes rows dependent:

    • StratifiedKFold keeps class proportions in each fold; scikit-learn uses it by default for classifiers when you pass cv=5.
    • GroupKFold keeps all rows from one user, patient or session in the same fold, so the model is never scored on someone it trained on.
    • TimeSeriesSplit always trains on the past and validates on the future.

    All preprocessing must happen inside each fold, which is why I put it in a Pipeline.

    What interviewers listen for
    • k fits, each row validated once
    • Report mean and standard deviation
    • Stratified, group and time-series variants
    • Preprocessing inside the folds

    Likely follow-up: What is nested cross-validation for? · Why is leave-one-out rarely worth it?

  7. 7.What is data leakage? Give three concrete examples and explain how you would catch them.hard

    Leakage is when information that will not be available at prediction time reaches the model during training or evaluation, so offline scores look better than production ever will.

    Three common forms:

    • Target leakage: a feature is a consequence of the label, such as refund_issued when predicting fraud, or a field filled in after the outcome.
    • Train-test contamination: fitting a scaler, imputer, feature selector or target encoder on the full dataset before splitting.
    • Split leakage: random splits when rows are correlated, such as several images of one patient or future rows used to predict the past.

    To catch it I look for scores that are too good, check which features dominate importance, ask for each feature when it becomes known, compare offline metrics with a shadow deployment, and keep all fitted preprocessing inside a Pipeline evaluated with the right splitter.

    What interviewers listen for
    • Information unavailable at prediction time
    • Target leakage, contamination, split leakage
    • Suspiciously high scores and dominant features
    • Pipelines and time or group splits

    Likely follow-up: How can target encoding leak, and how do you prevent it? · Is deduplication part of leakage prevention?

  8. 8.Your fraud dataset has 0.5% positives. How do you train and evaluate a model on it?mid

    First I change the evaluation, because 99.5% accuracy is what predicting "never fraud" scores. I use precision, recall, PR-AUC and a cost-based view of the confusion matrix, with a stratified split and enough positives in the test set to make the numbers stable.

    Then training options, cheapest first:

    • Keep the data as is, train a model that outputs good scores, and tune the decision threshold for the business tradeoff.
    • Use class weights (class_weight='balanced') so mistakes on positives cost more.
    • Resample: undersample negatives or oversample positives (SMOTE), only inside the training folds, never before splitting.

    If I resample or reweight, predicted probabilities no longer match real-world rates, so I recalibrate if downstream code uses them as probabilities. Often threshold tuning alone is enough.

    What interviewers listen for
    • Accuracy is misleading
    • PR-AUC, precision, recall and costs
    • Threshold tuning and class weights
    • Resample only inside training folds

    Likely follow-up: Why does oversampling before the split inflate scores? · When would you use anomaly detection instead?

  9. 9.Define precision, recall and F1. When do you optimise for precision and when for recall?easy

    From the confusion matrix: precision is TP / (TP + FP), the share of predicted positives that are real. Recall is TP / (TP + FN), the share of real positives you caught. F1 is their harmonic mean, 2PR / (P + R), which is low if either one is low.

    The choice depends on which error costs more. Optimise recall when missing a positive is expensive: cancer screening, fraud you must review, security alerts that feed a human queue. Optimise precision when false alarms are expensive: auto-blocking a customer payment, sending a sales team after leads, or a spam filter that hides real email.

    In practice you fix one and maximise the other, for example the best precision at 90% recall, and you pick the threshold to get there.

    What interviewers listen for
    • Precision = TP / (TP + FP)
    • Recall = TP / (TP + FN)
    • F1 is the harmonic mean
    • Choose by the cost of each error

    Likely follow-up: Why the harmonic mean and not the arithmetic mean? · What is F-beta?

  10. 10.What does ROC-AUC measure, and why can PR-AUC be the better metric on imbalanced data?hard

    The ROC curve plots true positive rate against false positive rate across all thresholds. ROC-AUC equals the probability that a random positive gets a higher score than a random negative, so it measures ranking quality and ignores calibration and the threshold.

    On heavily imbalanced data the false positive rate has a huge denominator. A model that produces 1,000 false alarms out of 1,000,000 negatives still has an FPR of 0.1%, so ROC-AUC can look excellent while most alerts are wrong. The precision-recall curve uses precision instead, which counts false positives against the alerts you raise, so PR-AUC (average precision) shows that pain. Its baseline is the positive rate, not 0.5.

    I report ROC-AUC for overall ranking, PR-AUC when positives are rare, and then a precision and recall at the operating threshold.

    What interviewers listen for
    • ROC: TPR versus FPR over thresholds
    • AUC is a ranking probability
    • FPR hides false positives when negatives dominate
    • PR-AUC baseline equals the positive rate

    Likely follow-up: What does an AUC of 0.3 tell you? · Is ROC-AUC affected by multiplying all scores by 2?

  11. 11.When would you report RMSE instead of MAE for a regression model, and what does R-squared add?easy

    MAE is the mean absolute error: every unit of error counts the same, so it is robust to outliers and easy to explain ("we are off by 4.8 units on average"). Minimising it targets the median.

    RMSE squares errors before averaging, so large errors dominate. Use it when big misses are disproportionately costly, such as under-forecasting demand for critical stock. Minimising squared error targets the mean. RMSE is always at least MAE, and a big gap between them signals a few large errors.

    R-squared is the share of variance explained relative to predicting the mean. It is unitless, which helps compare across targets, but it can be negative on test data and says nothing about whether errors are acceptable in business terms. For targets spanning orders of magnitude I also look at MAPE or error on a log scale.

    What interviewers listen for
    • MAE: linear, robust, median
    • RMSE: punishes large errors, mean
    • RMSE at least MAE; gap means outliers
    • R-squared is relative and can be negative

    Likely follow-up: Why is MAPE dangerous when targets are near zero?

  12. 12.What is log loss, why is it used to train classifiers, and how does it differ from accuracy?mid

    Log loss (binary cross-entropy) is the average of minus log of the probability the model gave to the true class. Predicting 0.9 for a true positive costs about 0.105; predicting 0.01 costs about 4.6. A constant 0.5 on a binary problem scores ln 2, about 0.693, which is a useful baseline.

    It is the negative log-likelihood of a Bernoulli model, so minimising it gives maximum-likelihood estimates, and it is smooth and differentiable, which gradient-based training needs. Accuracy is a step function of the threshold and gives no gradient.

    Accuracy only checks which side of 0.5 you land on. Log loss rewards calibrated, confident probabilities and punishes confident mistakes hard, so one prediction of 0.0001 for a real positive can dominate the average. I use it when probabilities feed decisions such as pricing or ranking, and I clip predictions away from 0 and 1.

    What interviewers listen for
    • Mean of minus log probability of the true class
    • Equals negative log-likelihood
    • Differentiable, unlike accuracy
    • Rewards calibration, punishes confident errors

    Likely follow-up: What is the Brier score? · How do you calibrate a model?

  13. 13.What are the assumptions of linear regression, and which ones actually matter for prediction?mid

    Ordinary least squares assumes: a linear relationship between features and target (after any transformations you apply), independent errors, constant error variance (homoscedasticity), roughly normal errors for exact confidence intervals, and no perfect multicollinearity.

    For pure prediction, linearity matters most: if the relationship is curved, the model is biased whatever else you do, and you fix it with transformed or interaction features. Correlated errors and changing variance mostly break the standard errors and p-values, not the point predictions. Multicollinearity makes individual coefficients unstable and hard to interpret but often barely changes predictions; ridge regression helps.

    If the task is inference, "does price affect demand", all of them matter, and I check residual plots, use robust standard errors and look at variance inflation factors.

    What interviewers listen for
    • Linearity, independence, constant variance, normality
    • Linearity matters most for prediction
    • Others mainly affect standard errors
    • Multicollinearity destabilises coefficients

    Likely follow-up: How do you read a residual plot? · Why is the closed-form solution not used for very large feature counts?

  14. 14.How does logistic regression work, and why is it called regression if it is used for classification?easy

    Logistic regression computes a linear score z = w.x + b and passes it through the sigmoid, 1 / (1 + e^-z), to get a probability between 0 and 1. It is a regression on the log-odds: log(p / (1 - p)) is modelled as a linear function of the features. That is where the name comes from; it only becomes a classifier when you apply a threshold.

    It is trained by minimising log loss, which is convex here, so gradient-based solvers find the global optimum. There is no closed-form solution, unlike linear regression.

    Interpretation is a strength: a weight of 0.7 means one unit of that feature multiplies the odds by e^0.7, about 2. The decision boundary is linear in the features, so curved boundaries need engineered features. In scikit-learn it is L2-regularized by default with C=1.0.

    What interviewers listen for
    • Sigmoid of a linear score
    • Linear in the log-odds
    • Trained with convex log loss
    • Weights read as odds ratios

    Likely follow-up: Why not train logistic regression with squared error? · How does it extend to more than two classes?

  15. 15.Compare batch, stochastic and mini-batch gradient descent. What happens when the learning rate is too high or too low?mid

    All three update weights by stepping against the gradient of the loss: w = w - lr * grad. They differ in how much data computes each gradient.

    • Batch uses the whole dataset: an exact gradient and smooth progress, but every step is expensive and it does not fit large data in memory.
    • Stochastic (SGD) uses one example: very cheap, noisy steps that can bounce around the minimum, though the noise can help escape poor regions.
    • Mini-batch (say 32 to 1,024 rows) is the practical default: vectorised hardware use and moderate noise.

    Too high a learning rate overshoots, oscillates or diverges with the loss climbing to infinity or NaN. Too low converges painfully slowly or stalls on plateaus. The fixes are learning-rate schedules, adaptive optimizers such as Adam, and feature scaling so one learning rate suits every direction.

    What interviewers listen for
    • Update rule w = w - lr * grad
    • Batch exact, SGD noisy, mini-batch practical
    • High rate diverges, low rate crawls
    • Scaling and schedules help convergence

    Likely follow-up: Why does feature scaling speed up gradient descent? · What does momentum add?

  16. 16.How does a decision tree choose its splits, and why do unconstrained trees overfit?easy

    A tree is built greedily from the root. At each node it tries every feature and candidate threshold and picks the split that most reduces impurity in the children, weighted by their size. For classification impurity is usually Gini (1 minus the sum of squared class shares) or entropy, where the reduction is called information gain. For regression it is variance, the mean squared error.

    Splitting stops when a node is pure or a limit is hit. Without limits, the tree keeps splitting until each leaf holds one class, often a single row, so it memorises noise: training accuracy near 100% and high variance on new data.

    You control it with max_depth, min_samples_leaf, min_samples_split, cost-complexity pruning (ccp_alpha), or by averaging many trees in a random forest. Trees need no feature scaling and handle interactions naturally.

    What interviewers listen for
    • Greedy search over features and thresholds
    • Gini, entropy or variance reduction
    • Unlimited depth memorises noise
    • Depth, leaf size and pruning limits

    Likely follow-up: Gini or entropy: does the choice matter much? · Why are trees unstable to small data changes?

  17. 17.How does a random forest reduce overfitting compared with a single decision tree?mid

    A random forest trains many deep trees, each on a bootstrap sample of the rows, and at each split considers only a random subset of features (max_features, often the square root of the feature count for classification). Predictions are averaged or voted.

    Each deep tree has low bias and high variance. Averaging reduces variance, but only if the trees make different errors. Bootstrapping alone leaves trees correlated because they all split on the same strong feature first; random feature subsets decorrelate them, so the average is much more stable than any single tree.

    Useful side effects: out-of-bag rows give a free validation estimate, adding trees does not cause overfitting (it just stops helping), and it works well with little tuning. Costs are model size, slower prediction and weaker extrapolation outside the training range.

    What interviewers listen for
    • Bootstrap rows plus random feature subsets
    • Averaging cuts variance of deep trees
    • Decorrelation is the key
    • Out-of-bag estimate; more trees do not overfit

    Likely follow-up: Why can a random forest not extrapolate a trend? · What does max_features control?

  18. 18.What is the difference between bagging and boosting? When would you choose a random forest over gradient boosting?mid

    Bagging trains models independently and in parallel on bootstrap samples, then averages them. It mainly reduces variance, so it suits deep, low-bias trees. Random forest is the standard example.

    Boosting trains models sequentially, each one fitting the mistakes of the ensemble so far. Gradient boosting fits each small tree to the negative gradient of the loss (the residuals for squared error) and adds it with a learning rate. It mainly reduces bias, using shallow trees, and usually wins on tabular accuracy.

    I would pick a random forest when I want a strong baseline with little tuning, noisy labels, or parallel training. I pick gradient boosting (XGBoost, LightGBM, CatBoost or scikit-learn's HistGradientBoostingClassifier) when accuracy matters and I can tune learning rate, depth and number of rounds with early stopping, because boosting can overfit if you keep adding trees.

    What interviewers listen for
    • Bagging: parallel, averages, lowers variance
    • Boosting: sequential, fits residuals, lowers bias
    • Boosting needs tuning and early stopping
    • Forest as a robust baseline

    Likely follow-up: Why does a lower learning rate need more trees?

  19. 19.What makes XGBoost and LightGBM faster or better than classic gradient boosting, and which hyperparameters do you tune first?hard

    Both are gradient-boosted trees with engineering and regularization on top.

    • XGBoost uses second-order (gradient and Hessian) information, adds L1 and L2 penalties on leaf weights plus a minimum gain to split (gamma), handles missing values by learning a default direction, and offers a fast histogram method.
    • LightGBM bins features into histograms, grows trees leaf-wise (always splitting the leaf with the largest gain) rather than level-wise, and samples rows by gradient size (GOSS) and bundles sparse features. That makes it very fast on large data, but leaf-wise trees overfit small data unless you cap num_leaves.

    I tune in this order: a moderate learning rate with early stopping to pick the number of rounds, then tree complexity (max_depth or num_leaves, min_child_weight or min_data_in_leaf), then row and column subsampling, then the L1 and L2 penalties. Finally I lower the learning rate and retrain.

    What interviewers listen for
    • Histogram binning and second-order gradients
    • Leaf-wise growth needs num_leaves limits
    • Built-in missing-value handling
    • Early stopping first, then complexity and sampling

    Likely follow-up: What does CatBoost do differently with categorical features?

  20. 20.How does a support vector machine work, and what are the kernel trick and the C parameter?mid

    A linear SVM finds the hyperplane that separates the classes with the maximum margin, the widest gap to the nearest points. Only those nearest points, the support vectors, define the boundary. The soft-margin version allows some points inside the margin or misclassified, with hinge loss as the penalty.

    C sets that tradeoff. Large C punishes violations heavily: a narrow margin that fits the training data closely, lower bias and higher variance. Small C allows more violations for a wider, smoother margin.

    The kernel trick replaces dot products with a kernel function, such as RBF or polynomial, which equals a dot product in a higher-dimensional feature space without computing that space. That lets the SVM draw curved boundaries. For RBF, gamma controls how local each point's influence is. SVMs need scaled features and scale poorly beyond tens of thousands of rows.

    What interviewers listen for
    • Maximum-margin hyperplane
    • Support vectors define the boundary
    • C trades margin width against violations
    • Kernels give non-linear boundaries cheaply

    Likely follow-up: Why do SVMs not output probabilities directly? · What does a very large gamma do?

  21. 21.How does k-nearest neighbours work, how do you choose k, and where does it struggle?easy

    k-NN stores the training set. To predict, it finds the k closest training points by a distance such as Euclidean and takes a majority vote (classification) or the mean (regression), optionally weighted by distance. There is no training step, so prediction does all the work.

    k is the bias-variance dial: k = 1 follows every noisy point (high variance), a large k smooths toward the overall majority (high bias). I choose it by cross-validation and prefer an odd k for binary problems to avoid ties.

    It struggles when:

    • Features are on different scales; distance is dominated by the largest unit, so scaling is mandatory.
    • There are many dimensions; distances concentrate and neighbours stop being meaningfully close (the curse of dimensionality).
    • Data is large; naive prediction is O(n) per query, so you need KD-trees, ball trees or approximate nearest-neighbour indexes.
    What interviewers listen for
    • Lazy learner: vote of nearest points
    • Small k high variance, large k high bias
    • Scaling is mandatory
    • Curse of dimensionality and query cost

    Likely follow-up: How does approximate nearest-neighbour search relate to vector databases?

  22. 22.How does k-means work, how do you choose k, and when is it the wrong algorithm?mid

    k-means alternates two steps: assign every point to its nearest centroid, then move each centroid to the mean of its points. It repeats until assignments stop changing, which minimises within-cluster sum of squares (inertia). It only finds a local optimum, so scikit-learn uses k-means++ initialisation and several restarts.

    To choose k, I look at the elbow in inertia versus k (inertia always falls as k grows, so you look for diminishing returns), the silhouette score, and above all whether the clusters are useful for the business question.

    It is the wrong tool when clusters are not roughly spherical and similar in size, when there are strong outliers (means get dragged), or for categorical data. Alternatives: DBSCAN or HDBSCAN for arbitrary shapes and noise, Gaussian mixtures for elliptical clusters with soft assignments, k-modes for categories. Always scale features first.

    What interviewers listen for
    • Assign to nearest centroid, recompute means
    • Local optimum; k-means++ and restarts
    • Elbow, silhouette and usefulness
    • Fails on odd shapes, outliers, unscaled data

    Likely follow-up: Why does inertia always decrease as k increases?

  23. 23.What does PCA do, and what are its pitfalls when used before a supervised model?mid

    PCA finds orthogonal directions, the principal components, along which the centred data has the most variance. They are the eigenvectors of the covariance matrix (computed in practice with an SVD), and the eigenvalues tell you how much variance each one explains. Projecting onto the top components reduces dimensions while keeping as much variance as possible.

    Pitfalls:

    • PCA is scale sensitive: unscaled, a feature measured in dollars dominates one measured in years. Standardize first.
    • It is unsupervised: the directions with the most variance are not necessarily the ones that predict the target, so you can throw away a small but crucial signal.
    • Components are mixtures of features and hard to explain.
    • It is linear, so curved structure needs kernel PCA, or UMAP and t-SNE for visualisation only.
    • Fit it on training data only, inside the pipeline, or it leaks.
    What interviewers listen for
    • Orthogonal directions of maximum variance
    • Eigenvectors of covariance, via SVD
    • Standardize first
    • Unsupervised: may drop predictive signal

    Likely follow-up: How do you choose the number of components?

  24. 24.What is naive about naive Bayes, and why does it still work well for text classification?easy

    Naive Bayes applies Bayes' theorem, P(class given features) proportional to P(features given class) times P(class), and makes the naive assumption that features are conditionally independent given the class. Then the likelihood is just the product of per-feature probabilities, which you estimate by counting.

    The assumption is almost always false; in text, "New" and "York" are clearly dependent. It still works because classification only needs the right class to score highest, not accurate probabilities. The model is very fast, needs little data, and handles thousands of word features easily, which is why it is a strong baseline for spam and topic classification.

    Practical details: use Laplace smoothing (alpha) so an unseen word does not zero the product, compute in log space to avoid underflow, and do not trust its probabilities without calibration because they tend to be pushed to extremes.

    What interviewers listen for
    • Conditional independence given the class
    • Product of per-feature likelihoods
    • Ranking survives the wrong assumption
    • Laplace smoothing and log space

    Likely follow-up: Multinomial versus Bernoulli versus Gaussian naive Bayes?

  25. 25.Which models need feature scaling and which do not? Standardization or min-max normalization?easy

    Scaling matters wherever a model uses distances, dot products with a penalty, or gradient descent:

    • k-NN, k-means, SVMs and PCA, because distance or variance is dominated by large-unit features.
    • Regularized linear and logistic regression, because the penalty treats all weights alike, so units change which features get shrunk.
    • Neural networks and anything trained with gradient descent, because badly scaled features make the loss surface elongated and convergence slow.

    Tree-based models (decision trees, random forests, gradient boosting) do not need it: a split threshold is unaffected by any monotonic rescaling.

    Standardization (zero mean, unit variance) is my default. Min-max to [0, 1] suits bounded inputs like pixels but is sensitive to outliers, as is the mean; RobustScaler uses the median and IQR. Whichever I pick, it is fitted on training data only.

    What interviewers listen for
    • Distance, penalty and gradient methods need it
    • Trees do not
    • Standardize by default; min-max for bounded data
    • Fit the scaler on training data only

    Likely follow-up: Does scaling change the predictions of an unregularized linear regression?

  26. 26.How do you encode a categorical feature with 50,000 distinct values, such as a merchant id?mid

    One-hot encoding works for a few dozen categories but creates 50,000 sparse columns here, and rare merchants get unreliable weights. My options:

    • Target encoding: replace each category with the mean target for it, shrunk toward the global mean for rare categories. It must be computed out of fold, or the encoding leaks each row's own label. scikit-learn's TargetEncoder does cross-fitting inside fit_transform.
    • Frequency or count encoding: how common the merchant is, often surprisingly predictive.
    • Grouping: keep the top N categories and bucket the rest as "other", or map to a higher level such as merchant category.
    • Hashing: map categories into a fixed number of buckets, accepting collisions.
    • Learned embeddings in a neural network, or native categorical support in LightGBM and CatBoost.

    I also plan for categories that appear only in production, with an explicit unknown bucket.

    What interviewers listen for
    • One-hot does not scale to high cardinality
    • Target encoding computed out of fold
    • Frequency, grouping, hashing, embeddings
    • Handle unseen categories

    Likely follow-up: Why is label encoding risky for linear models?

  27. 27.How do you tune hyperparameters? Compare grid search, random search and Bayesian optimization.mid

    Hyperparameters are set before training (tree depth, learning rate, regularization strength), so I choose them by validation performance, usually with cross-validation.

    • Grid search tries every combination. It is exhaustive but its cost multiplies with each parameter, and it wastes trials on parameters that do not matter.
    • Random search samples combinations from distributions. With the same budget it explores more distinct values of the parameters that matter, so it usually beats grid search; sample rates like learning rate on a log scale.
    • Bayesian optimization (Optuna, scikit-optimize) builds a model of score versus parameters and picks promising points next, with pruning of bad trials. Best when each fit is expensive.

    Rules I follow: tune on validation folds and report on an untouched test set, use nested CV when I need an unbiased estimate of the whole tuning procedure, use early stopping for boosting rounds, and stop when gains are within the noise across folds.

    What interviewers listen for
    • Grid is exhaustive and expensive
    • Random search is better per budget
    • Bayesian methods for costly fits
    • Tuning score is optimistic: keep a test set

    Likely follow-up: What is successive halving? · Why sample the learning rate on a log scale?

  28. 28.How would you explain a gradient boosting model to a stakeholder or a regulator?hard

    I separate global explanations (what drives the model overall) from local ones (why this customer got this score).

    Globally I avoid the default impurity-based importances, which are computed on training data and favour high-cardinality and continuous features. Permutation importance on a held-out set is more honest: shuffle one feature and measure the drop in score, though correlated features share and hide each other's importance. Partial dependence plots show how the prediction changes with a feature.

    Locally I use SHAP values, which split one prediction into additive feature contributions that sum to the difference from the average prediction; TreeSHAP computes them exactly and quickly for tree ensembles.

    I also say what these are not: they explain the model, not causation in the world. In regulated settings such as credit, I might prefer a monotonic-constrained model or a scorecard so explanations are faithful by construction.

    What interviewers listen for
    • Global versus local explanations
    • Impurity importance is biased
    • Permutation importance on held-out data
    • SHAP for per-prediction contributions; not causal

    Likely follow-up: How do correlated features distort permutation importance?

  29. 29.A model scored 0.92 AUC offline but performs poorly after launch. How do you debug it?hard

    I would work through the usual causes in order of likelihood.

    • Leakage or a bad split: a feature that is only known after the outcome, or a random split where a time or group split was needed. I recheck when each top feature is available and rerun validation with an out-of-time split.
    • Training-serving skew: the feature is computed differently online (different code path, units, missing-value defaults, time zone, stale lookups). I log the online features and compare their distributions with training.
    • Drift: the input distribution or the relationship to the label has changed since the training window.
    • Wrong metric or threshold: AUC measures ranking, but production uses a threshold, and the base rate or costs may differ.
    • Feedback effects: the model changes the behaviour it is measured on.

    The lasting fixes are a shared feature pipeline, out-of-time validation, shadow deployment before launch and monitoring of inputs and outcomes.

    What interviewers listen for
    • Leakage and split mismatch
    • Training-serving skew
    • Data and concept drift
    • Threshold and base-rate mismatch
    • Shadow deployment and monitoring

    Likely follow-up: How would you detect drift without labels?

  30. 30.Why is 0.5 not always the right classification threshold, and what does it mean for a model to be calibrated?hard

    The threshold turns scores into decisions, and the right one depends on the cost of each error and the base rate, not on the model. If a missed fraud costs 50 times a false alarm, the cost-minimising threshold on a calibrated probability is about 1 / 51, far below 0.5. I choose thresholds on a validation set, either by minimising expected cost or by meeting a constraint such as "at least 90% recall", and check them on the test set.

    A model is calibrated when its probabilities match observed frequencies: of the cases scored 0.2, about 20% are positive. You check it with a reliability diagram or the Brier score. Boosted trees, SVMs and naive Bayes are often miscalibrated, and class weights or resampling deliberately distort probabilities. CalibratedClassifierCV with sigmoid (Platt) or isotonic calibration fixes that, on data not used for training. Calibration matters whenever probabilities are used as numbers, such as expected-loss calculations.

    What interviewers listen for
    • Threshold follows costs and base rate
    • Pick it on validation data
    • Calibrated means probabilities match frequencies
    • Platt or isotonic calibration on held-out data

    Likely follow-up: Does calibration change ROC-AUC?

Prefer multiple choice? All 20 Machine Learning MCQs with answers →

esc