The facts a machine learning round keeps coming back to, from “why is accuracy useless here” to “why is your forest flat beyond the data”. API names are scikit-learn 1.9.
Learning types
| Type | Learns from | Typical tasks | Examples |
|---|---|---|---|
| Supervised | labelled pairs (x, y) | classification, regression | spam filter, price prediction |
| Unsupervised | x only | clustering, dimensionality reduction, anomaly detection | k-means segments, PCA |
| Reinforcement | actions and rewards from an environment | sequential decisions | game agents, bidding, robotics |
| Semi-supervised | few labels plus many unlabelled rows | label propagation, pseudo-labels | medical images with scarce labels |
| Self-supervised | labels made from the data itself | pretraining | next-token prediction, masked images |
Splits & validation
- Train fits parameters, validation (or CV on train) picks models and hyperparameters, test is used once for the final estimate. Every look at the test set that changes a decision leaks it.
- Typical: 60/20/20 or 80/20 plus CV. Big data can use a much smaller share for test; small data needs CV.
- Split the way production works: by time for forecasting, by group (user, patient, device) when rows are correlated, stratified for rare classes.
| Splitter | Use when |
|---|---|
KFold(5, shuffle=True) |
i.i.d. rows, regression |
StratifiedKFold |
classification; default for classifiers when you pass cv=5 (no shuffle) |
GroupKFold / StratifiedGroupKFold |
several rows per entity |
TimeSeriesSplit |
ordered data: train on the past, test on the future |
| Nested CV | an unbiased estimate of a tuning procedure (inner loop tunes, outer loop scores) |
Bias, variance, overfitting
- Expected error = bias squared + variance + irreducible noise.
- High bias (underfit): train and validation error both high and close. Fix: more features, more flexible model, less regularization, train longer.
- High variance (overfit): train error low, validation error much higher. Fix: more data, simpler model, regularization, dropout, early stopping, bagging.
- More flexibility usually lowers bias and raises variance: deeper trees, smaller k in k-NN, larger
Cin SVMs and logistic regression, higher polynomial degree. - Learning curves (score versus training size) tell you whether more data will help: only when there is a gap.
Regularization
| Technique | Effect | Notes |
|---|---|---|
| L2 / ridge | penalty lambda * sum(w^2): shrinks all weights |
handles correlated features; rarely exactly zero |
| L1 / lasso | penalty lambda * sum(abs(w)): many weights exactly 0 |
feature selection; picks one of correlated features arbitrarily |
| Elastic net | mix of L1 and L2 | l1_ratio between 0 and 1 |
| Dropout | randomly zeroes activations in training | neural networks only; off at inference |
| Early stopping | stop when validation loss stops improving | boosting (n_iter_no_change), neural nets |
| Tree limits | max_depth, min_samples_leaf, ccp_alpha |
pruning for trees |
- In scikit-learn
Cis the inverse of strength (LogisticRegression,SVC): smallerC, stronger penalty.alpha(Ridge,Lasso) is the strength itself. LogisticRegression()is L2-regularized withC=1.0by default. Since 1.8,penaltyis deprecated: usel1_ratio(0 is L2, 1 is L1) andC=np.inffor none.- Scale features before any penalty, or units decide which weights shrink.
Data leakage
- Target leakage: a feature only known after the outcome (
refund_issuedto predict fraud,discharge_dateto predict admission). - Contamination: fitting a scaler, imputer, PCA, feature selector, target encoder or resampler on all rows before splitting.
- Split leakage: random split on time-ordered or grouped data; duplicate rows across train and test.
- Smells: a validation score that is too good, one feature dominating importance, a big offline-to-online drop.
- Cure: put every fitted step in a
Pipeline, evaluate with the right splitter, and ask “when is this value known?” for each feature.
pre = ColumnTransformer([
("num", make_pipeline(SimpleImputer(strategy="median"), StandardScaler()), ["age", "income"]),
("cat", OneHotEncoder(handle_unknown="ignore"), ["plan"]),
])
model = make_pipeline(pre, LogisticRegression(C=1.0, max_iter=1000))
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
scores = cross_val_score(model, X_train, y_train, cv=cv, scoring="roc_auc") # preprocessing refit per foldClassification metrics
| Predicted positive | Predicted negative | |
|---|---|---|
| Actual positive | TP | FN (type II) |
| Actual negative | FP (type I) | TN |
| Metric | Formula | Read it as |
|---|---|---|
| Accuracy | (TP + TN) / all | useless when classes are imbalanced |
| Precision | TP / (TP + FP) | of the alerts, how many were real |
| Recall (TPR, sensitivity) | TP / (TP + FN) | of the real positives, how many we caught |
| Specificity (TNR) | TN / (TN + FP) | of the negatives, how many we cleared |
| FPR | FP / (FP + TN) | 1 - specificity; x-axis of ROC |
| F1 | 2PR / (P + R) | harmonic mean, low if either is low |
| F-beta | weights recall beta times as much as precision | F2 for recall-heavy problems |
| ROC-AUC | area under TPR vs FPR | P(random positive scores above random negative); 0.5 is random |
| PR-AUC / average precision | area under precision vs recall | baseline = positive rate; better for rare positives |
| Log loss | mean of -log p(true class) |
rewards calibrated probabilities; 0.693 = constant 0.5 |
| Brier score | mean squared error of probabilities | calibration plus sharpness |
- AUC is threshold-free and unchanged by any monotonic transform of scores; precision, recall and F1 depend on the threshold.
- Choose the threshold on validation data by cost: with calibrated probabilities, predict positive when
p > C_FP / (C_FP + C_FN). - Multiclass averaging: macro (every class equal), weighted (by support), micro (pool all decisions; equals accuracy for single-label problems).
Regression metrics
| Metric | Notes |
|---|---|
| MAE | average absolute error; robust; optimal constant is the median |
| MSE / RMSE | squares errors, punishes big misses; optimal constant is the mean; RMSE is at least MAE |
| R-squared | 1 - SSE / SST; share of variance explained; can be negative on test data |
| MAPE | percentage error; explodes near zero targets |
| RMSLE | error on log scale; for targets spanning orders of magnitude |
Class imbalance
- Fix the metric first: PR-AUC, recall at a fixed precision, expected cost.
- Stratified splits with enough positives in test.
- Tune the decision threshold (often enough on its own).
class_weight="balanced"orscale_pos_weight(XGBoost).- Resample inside the training folds only (undersample, SMOTE via imbalanced-learn); recalibrate probabilities afterwards.
Algorithms at a glance
| Model | Idea | Scale features? | Watch out for |
|---|---|---|---|
| Linear regression | least squares, closed form or GD | for regularization or GD | outliers, non-linearity, collinearity |
| Logistic regression | sigmoid of a linear score, log loss | yes (penalty) | linear boundary; weights are log-odds |
| Decision tree | greedy splits by Gini, entropy or variance | no | overfits without depth or leaf limits; unstable |
| Random forest | bagged deep trees + random feature subsets | no | cannot extrapolate; large models |
| Gradient boosting | shallow trees fit to the loss gradient, in sequence | no | needs early stopping; learning rate vs rounds |
| SVM | maximum margin, hinge loss, kernels | yes | C, gamma; slow above about 100k rows |
| k-NN | vote or mean of k nearest points | yes | curse of dimensionality; slow queries |
| Naive Bayes | Bayes’ rule with conditionally independent features | no | probabilities poorly calibrated; needs smoothing |
| k-means | assign to nearest centroid, move centroids to means | yes | spherical clusters; choose k; outliers |
| PCA | orthogonal directions of maximum variance (SVD) | yes | unsupervised: may drop the predictive signal |
Gradient descent
- Update:
w = w - lr * gradient. Batch (all rows), stochastic (one row), mini-batch (32 to 1,024 rows, the default). - Learning rate too high: oscillation or divergence (loss goes to inf or NaN). Too low: slow progress, stuck on plateaus.
- Convex losses (linear and logistic regression) have one global minimum; neural nets do not.
- Unscaled features stretch the loss surface and force a tiny learning rate: standardize.
- Momentum, RMSProp and Adam adapt the step; schedules decay it.
Trees & ensembles
- Gini =
1 - sum(p_k^2); entropy =-sum(p_k * log(p_k)). Choice rarely matters; Gini is a bit faster. - Bagging (parallel, averages) cuts variance; boosting (sequential, fits residuals) cuts bias.
- Random forest: more trees never overfit, they just stop helping; out-of-bag score is a free validation estimate.
- Boosting knobs, in tuning order: learning rate + early stopping, depth or
num_leaves, min samples per leaf, row and column subsampling, L1 and L2 on leaf weights. - XGBoost: second-order gradients, regularized leaves, sparsity-aware missing values. LightGBM: histograms, leaf-wise growth (cap
num_leaves), GOSS. CatBoost: ordered target statistics for categories. scikit-learn:HistGradientBoostingClassifier(native NaN and categorical support).
Features & tuning
- Scale for distance, penalty and gradient methods; never needed for trees.
StandardScalerdefault,MinMaxScalerfor bounded inputs,RobustScalerwith outliers. - Categoricals: one-hot for low cardinality; target encoding (out of fold:
TargetEncoder), frequency, hashing or embeddings for high cardinality; always an “unknown” bucket. - Missing values: impute inside the pipeline, add a “was missing” indicator when missingness carries signal.
- Tuning: random search beats grid search for the same budget; Bayesian optimization (Optuna) when fits are expensive; sample learning rate and regularization on a log scale.
- Interpretability: permutation importance on held-out data, partial dependence, SHAP for single predictions. Impurity importance is biased toward high-cardinality features. None of these are causal.
Quick answers
- Why not accuracy for fraud? Predicting “never fraud” scores 99.5%.
- L1 or L2? L1 when you want sparsity or feature selection, L2 for stable shrinkage with correlated features.
- Why does my model ace validation and fail in production? Leakage, training-serving skew, drift or a threshold tuned for a different base rate.
- Why is it called logistic regression? It is a linear regression on the log-odds.
- Random forest or boosting? Forest for a quick robust baseline; boosting for top tabular accuracy with tuning.
- How to choose k in k-means? Elbow, silhouette, and whether the segments are useful; inertia always falls with k.
- What does AUC 0.3 mean? Ranking is worse than random; flipping the scores gives 0.7, so check for a label or sign bug.
- Does a tree need scaling? No: thresholds are unaffected by monotonic transforms.
- Is dropout used at prediction time? No; it is switched off and activations are already scaled during training.
Practise these in the Machine Learning chapter.