Machine Learning · cheat sheet

Machine Learning

Learning types, data splits, bias-variance, regularization, cross-validation, leakage, metrics and the classic algorithms, in the form ML interviews test them.

The facts a machine learning round keeps coming back to, from “why is accuracy useless here” to “why is your forest flat beyond the data”. API names are scikit-learn 1.9.

Learning types

Type Learns from Typical tasks Examples
Supervised labelled pairs (x, y) classification, regression spam filter, price prediction
Unsupervised x only clustering, dimensionality reduction, anomaly detection k-means segments, PCA
Reinforcement actions and rewards from an environment sequential decisions game agents, bidding, robotics
Semi-supervised few labels plus many unlabelled rows label propagation, pseudo-labels medical images with scarce labels
Self-supervised labels made from the data itself pretraining next-token prediction, masked images

Splits & validation

  • Train fits parameters, validation (or CV on train) picks models and hyperparameters, test is used once for the final estimate. Every look at the test set that changes a decision leaks it.
  • Typical: 60/20/20 or 80/20 plus CV. Big data can use a much smaller share for test; small data needs CV.
  • Split the way production works: by time for forecasting, by group (user, patient, device) when rows are correlated, stratified for rare classes.
Splitter Use when
KFold(5, shuffle=True) i.i.d. rows, regression
StratifiedKFold classification; default for classifiers when you pass cv=5 (no shuffle)
GroupKFold / StratifiedGroupKFold several rows per entity
TimeSeriesSplit ordered data: train on the past, test on the future
Nested CV an unbiased estimate of a tuning procedure (inner loop tunes, outer loop scores)

Bias, variance, overfitting

  • Expected error = bias squared + variance + irreducible noise.
  • High bias (underfit): train and validation error both high and close. Fix: more features, more flexible model, less regularization, train longer.
  • High variance (overfit): train error low, validation error much higher. Fix: more data, simpler model, regularization, dropout, early stopping, bagging.
  • More flexibility usually lowers bias and raises variance: deeper trees, smaller k in k-NN, larger C in SVMs and logistic regression, higher polynomial degree.
  • Learning curves (score versus training size) tell you whether more data will help: only when there is a gap.

Regularization

Technique Effect Notes
L2 / ridge penalty lambda * sum(w^2): shrinks all weights handles correlated features; rarely exactly zero
L1 / lasso penalty lambda * sum(abs(w)): many weights exactly 0 feature selection; picks one of correlated features arbitrarily
Elastic net mix of L1 and L2 l1_ratio between 0 and 1
Dropout randomly zeroes activations in training neural networks only; off at inference
Early stopping stop when validation loss stops improving boosting (n_iter_no_change), neural nets
Tree limits max_depth, min_samples_leaf, ccp_alpha pruning for trees
  • In scikit-learn C is the inverse of strength (LogisticRegression, SVC): smaller C, stronger penalty. alpha (Ridge, Lasso) is the strength itself.
  • LogisticRegression() is L2-regularized with C=1.0 by default. Since 1.8, penalty is deprecated: use l1_ratio (0 is L2, 1 is L1) and C=np.inf for none.
  • Scale features before any penalty, or units decide which weights shrink.

Data leakage

  • Target leakage: a feature only known after the outcome (refund_issued to predict fraud, discharge_date to predict admission).
  • Contamination: fitting a scaler, imputer, PCA, feature selector, target encoder or resampler on all rows before splitting.
  • Split leakage: random split on time-ordered or grouped data; duplicate rows across train and test.
  • Smells: a validation score that is too good, one feature dominating importance, a big offline-to-online drop.
  • Cure: put every fitted step in a Pipeline, evaluate with the right splitter, and ask “when is this value known?” for each feature.
pre = ColumnTransformer([
    ("num", make_pipeline(SimpleImputer(strategy="median"), StandardScaler()), ["age", "income"]),
    ("cat", OneHotEncoder(handle_unknown="ignore"), ["plan"]),
])
model = make_pipeline(pre, LogisticRegression(C=1.0, max_iter=1000))
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
scores = cross_val_score(model, X_train, y_train, cv=cv, scoring="roc_auc")  # preprocessing refit per fold
python

Classification metrics

Predicted positive Predicted negative
Actual positive TP FN (type II)
Actual negative FP (type I) TN
Metric Formula Read it as
Accuracy (TP + TN) / all useless when classes are imbalanced
Precision TP / (TP + FP) of the alerts, how many were real
Recall (TPR, sensitivity) TP / (TP + FN) of the real positives, how many we caught
Specificity (TNR) TN / (TN + FP) of the negatives, how many we cleared
FPR FP / (FP + TN) 1 - specificity; x-axis of ROC
F1 2PR / (P + R) harmonic mean, low if either is low
F-beta weights recall beta times as much as precision F2 for recall-heavy problems
ROC-AUC area under TPR vs FPR P(random positive scores above random negative); 0.5 is random
PR-AUC / average precision area under precision vs recall baseline = positive rate; better for rare positives
Log loss mean of -log p(true class) rewards calibrated probabilities; 0.693 = constant 0.5
Brier score mean squared error of probabilities calibration plus sharpness
  • AUC is threshold-free and unchanged by any monotonic transform of scores; precision, recall and F1 depend on the threshold.
  • Choose the threshold on validation data by cost: with calibrated probabilities, predict positive when p > C_FP / (C_FP + C_FN).
  • Multiclass averaging: macro (every class equal), weighted (by support), micro (pool all decisions; equals accuracy for single-label problems).

Regression metrics

Metric Notes
MAE average absolute error; robust; optimal constant is the median
MSE / RMSE squares errors, punishes big misses; optimal constant is the mean; RMSE is at least MAE
R-squared 1 - SSE / SST; share of variance explained; can be negative on test data
MAPE percentage error; explodes near zero targets
RMSLE error on log scale; for targets spanning orders of magnitude

Class imbalance

  1. Fix the metric first: PR-AUC, recall at a fixed precision, expected cost.
  2. Stratified splits with enough positives in test.
  3. Tune the decision threshold (often enough on its own).
  4. class_weight="balanced" or scale_pos_weight (XGBoost).
  5. Resample inside the training folds only (undersample, SMOTE via imbalanced-learn); recalibrate probabilities afterwards.

Algorithms at a glance

Model Idea Scale features? Watch out for
Linear regression least squares, closed form or GD for regularization or GD outliers, non-linearity, collinearity
Logistic regression sigmoid of a linear score, log loss yes (penalty) linear boundary; weights are log-odds
Decision tree greedy splits by Gini, entropy or variance no overfits without depth or leaf limits; unstable
Random forest bagged deep trees + random feature subsets no cannot extrapolate; large models
Gradient boosting shallow trees fit to the loss gradient, in sequence no needs early stopping; learning rate vs rounds
SVM maximum margin, hinge loss, kernels yes C, gamma; slow above about 100k rows
k-NN vote or mean of k nearest points yes curse of dimensionality; slow queries
Naive Bayes Bayes’ rule with conditionally independent features no probabilities poorly calibrated; needs smoothing
k-means assign to nearest centroid, move centroids to means yes spherical clusters; choose k; outliers
PCA orthogonal directions of maximum variance (SVD) yes unsupervised: may drop the predictive signal

Gradient descent

  • Update: w = w - lr * gradient. Batch (all rows), stochastic (one row), mini-batch (32 to 1,024 rows, the default).
  • Learning rate too high: oscillation or divergence (loss goes to inf or NaN). Too low: slow progress, stuck on plateaus.
  • Convex losses (linear and logistic regression) have one global minimum; neural nets do not.
  • Unscaled features stretch the loss surface and force a tiny learning rate: standardize.
  • Momentum, RMSProp and Adam adapt the step; schedules decay it.

Trees & ensembles

  • Gini = 1 - sum(p_k^2); entropy = -sum(p_k * log(p_k)). Choice rarely matters; Gini is a bit faster.
  • Bagging (parallel, averages) cuts variance; boosting (sequential, fits residuals) cuts bias.
  • Random forest: more trees never overfit, they just stop helping; out-of-bag score is a free validation estimate.
  • Boosting knobs, in tuning order: learning rate + early stopping, depth or num_leaves, min samples per leaf, row and column subsampling, L1 and L2 on leaf weights.
  • XGBoost: second-order gradients, regularized leaves, sparsity-aware missing values. LightGBM: histograms, leaf-wise growth (cap num_leaves), GOSS. CatBoost: ordered target statistics for categories. scikit-learn: HistGradientBoostingClassifier (native NaN and categorical support).

Features & tuning

  • Scale for distance, penalty and gradient methods; never needed for trees. StandardScaler default, MinMaxScaler for bounded inputs, RobustScaler with outliers.
  • Categoricals: one-hot for low cardinality; target encoding (out of fold: TargetEncoder), frequency, hashing or embeddings for high cardinality; always an “unknown” bucket.
  • Missing values: impute inside the pipeline, add a “was missing” indicator when missingness carries signal.
  • Tuning: random search beats grid search for the same budget; Bayesian optimization (Optuna) when fits are expensive; sample learning rate and regularization on a log scale.
  • Interpretability: permutation importance on held-out data, partial dependence, SHAP for single predictions. Impurity importance is biased toward high-cardinality features. None of these are causal.

Quick answers

  • Why not accuracy for fraud? Predicting “never fraud” scores 99.5%.
  • L1 or L2? L1 when you want sparsity or feature selection, L2 for stable shrinkage with correlated features.
  • Why does my model ace validation and fail in production? Leakage, training-serving skew, drift or a threshold tuned for a different base rate.
  • Why is it called logistic regression? It is a linear regression on the log-odds.
  • Random forest or boosting? Forest for a quick robust baseline; boosting for top tabular accuracy with tuning.
  • How to choose k in k-means? Elbow, silhouette, and whether the segments are useful; inertia always falls with k.
  • What does AUC 0.3 mean? Ranking is worse than random; flipping the scores gives 0.7, so check for a label or sign bug.
  • Does a tree need scaling? No: thresholds are unaffected by monotonic transforms.
  • Is dropout used at prediction time? No; it is switched off and activations are already scaled during training.

Practise these in the Machine Learning chapter.

esc