MLOps · cheat sheet

MLOps

ML lifecycle, versioning, MLflow tracking and registry, pipelines, feature stores, serving, rollouts, drift, retraining, CI/CD for ML, GPU cost and LLMOps.

What an MLOps round checks: can you get a model to production repeatably, roll it out safely and notice when it stops working. Tool names are MLflow 3.16, scikit-learn 1.9 and SciPy 1.18.

Lifecycle & maturity

  • Loop, not a line: frame (metric, baseline) → data (collect, validate, split like production) → experiment (tracked) → pipeline (repeatable training) → validate (vs champion, slices) → deploy (shadow, canary, A/B) → monitor → retrain or retire.
  • MLOps = DevOps + two more versioned inputs (data, model) + models that decay without any deploy.
  • CI (test code and data contracts) · CD (ship pipeline and service) · CT (continuous training produces new models).
Google Cloud level What is automated Typical pain
0: manual nothing; notebook hands a file to engineering rare releases, no lineage, no monitoring
1: pipeline automation the training pipeline: retrains on new data with data and model validation pipeline code changes still manual
2: CI/CD automation building, testing and deploying the pipeline itself needs real platform investment

Reproducibility: what to record per run

Input Record Tooling
Code git commit, clean tree MLflow tags mlflow.source.git.commit when run inside a repo
Data immutable snapshot path or table version + content hash dated paths, Delta/Iceberg time travel, DVC, lakeFS
Config all hyperparameters, feature list, seeds mlflow.log_params
Environment image digest or lockfile; CUDA and driver for GPU MLflow writes requirements.txt / conda.yaml with the model
Artifact model file + checksum + signature registry version
  • “Latest” is not a version. Bit-exact GPU retraining may be impossible (non-deterministic kernels): keep the original artifact.

MLflow tracking & registry

mlflow.set_experiment("churn")
with mlflow.start_run(run_name="logreg-c1"):
    mlflow.log_params({"C": 1.0})
    mlflow.log_metric("val_auc", auc)                      # step=... for curves
    mlflow.sklearn.log_model(model, name="model",
        signature=infer_signature(X_tr, model.predict(X_tr)),
        registered_model_name="churn-model")               # creates version N
MlflowClient().set_registered_model_alias("churn-model", "champion", 3)
model = mlflow.pyfunc.load_model("models:/churn-model@champion")
python
  • Params are immutable per run (changing one raises); metrics keep a history, run.data.metrics holds the latest.
  • search_runs(filter_string="metrics.val_auc > 0.9"): keys need metrics. / params. / tags. prefixes.
  • Registry: named model, immutable numbered versions, lineage to the run. Aliases (@champion, @challenger) are pointers you move; stages are deprecated since 2.9. New versions never move an alias.
  • Promotion and rollback = move the alias. models:/name/latest changes on every registration: never serve it.

Pipelines & orchestration

  • DAG: extract window → validate data (schema, nulls, ranges, counts) → features + splits → train (tracked) → evaluate vs champion → register if passing, else alert.
  • Steps: containerised, idempotent, explicit inputs/outputs. Orchestrators: Airflow, Kubeflow Pipelines, Vertex AI Pipelines, Dagster, Prefect.

Features & training-serving skew

Cause of skew Fix
Feature logic written twice (SQL vs service code) one definition: feature store or shared library
Training features computed “as of today” point-in-time join (pd.merge_asof, feature store)
Scaler/encoder refitted at serving ship preprocessing inside the model (Pipeline, pyfunc)
Online store stale or missing values freshness monitoring, alert on default-value rate
  • Feature store: offline store (full timestamped history, for training) + online store (latest per entity, Redis/DynamoDB, for serving) + materialisation job between them.
  • Log the features actually served, then compare with training distributions.

Serving patterns

Batch Online Streaming
Trigger schedule request event
Latency hours ms (p99 SLO) seconds
Typical churn scores, nightly recs fraud at checkout, search ranking real-time features, alerts
Risk stale for new entities tail latency, availability state, ordering
  • Latency: SLO on p95/p99, not mean. Culprits: feature lookups, cold starts, CPU throttling, oversized inputs. Fixes: batched lookups, warm-up before readiness, quantise/distil, ONNX Runtime or TensorRT, timeouts with fallback.
  • Dynamic batching on GPUs: more throughput, more queueing delay. Cap the wait.

Rollouts

Strategy Users see new model? Measures Rollback
Shadow no (mirrored, logged only) latency, errors, skew, disagreement turn off mirror
Canary small share, stepped up guardrails on real users traffic weight to 0 / alias back
A/B test random 50/50 (or other) causal effect on business metric stop the test
Blue-green all at switch quick swap switch back
  • Route by stable hash of user id (zlib.crc32, SHA-256), never Python hash() (randomised per process).
  • A/B hygiene: pre-registered metric, sample size from minimum effect, no peeking, check sample ratio mismatch.
  • Rollback needs the old model loadable and its features still produced.

Monitoring & drift

Type What changes Detect with Needs labels?
Data drift P(X) PSI, KS, chi-square, null rates no
Prediction drift score distribution PSI on scores, approval rate no
Concept drift P(y given X) accuracy, calibration, proxy labels yes
Label shift P(y) class rates yes
def psi(ref, cur, bins=10, eps=1e-6):
    edges = np.quantile(ref, np.linspace(0, 1, bins + 1)); edges[[0, -1]] = -np.inf, np.inf
    r = np.clip(np.histogram(ref, edges)[0] / len(ref), eps, None)
    c = np.clip(np.histogram(cur, edges)[0] / len(cur), eps, None)
    return float(np.sum((c - r) * np.log(c / r)))
python
  • PSI bands: < 0.1 stable, 0.1 to 0.25 moderate, > 0.25 significant. Symmetric; empty bins need epsilon.
  • KS (scipy.stats.ks_2samp): max CDF gap, numeric only. With large samples p is tiny for trivial shifts: alert on the statistic.
  • Layered monitoring: service health → data quality → drift → prediction drift → proxy labels → ground truth. Log prediction id, model version, features and timestamp to join late labels.

Retraining triggers

Trigger Good when Watch out
Schedule decay rate is known retrains for nothing, or too late
Performance drop labels arrive fast delayed labels
Drift labels slow drift is not always harm
Event known shifts (launch, pricing) manual
  • Every retrained model passes the same gate. Retraining cannot fix broken upstream data.

CI/CD for ML: tests

  • PR: feature unit tests, data contract tests, smoke training on a fixture, serving test (loads, signature, latency).
  • Promotion gate: same fresh holdout for champion and challenger, margin above noise, slice and calibration checks, behavioural tests, size/latency budget, fairness where relevant.

GPU cost

  • Measure utilisation first. Training: spot instances + checkpointing, mixed precision (bf16), fix data-loader starvation, early-stopped searches.
  • Serving: right-size (CPU often fine), quantise/distil, dynamic batching, multi-model GPUs (Triton, MIG), autoscale on queue depth or concurrency, scale to zero for rare models.
  • Kubernetes: request nvidia.com/gpu, readiness after weights load and warm-up, keep warm replicas (cold start = node + image + weights).

LLMOps basics

  • Version prompts like code: registry with immutable versions and aliases (prompts:/support-reply@production via mlflow.genai.load_prompt).
  • The release unit = prompt version + model id + parameters + tools + retrieval config. A provider model change is a release.
  • Eval gate in CI: versioned dataset of real and adversarial cases; deterministic checks (format, citations), reference metrics, LLM-as-judge (mlflow.genai.evaluate with scorers) calibrated on human labels; compare to production with tolerances; track cost and latency.
  • Trace every request with prompt and model versions; feed failures back into the eval set.

Practise these in the MLOps chapter.

esc