What an MLOps round checks: can you get a model to production repeatably, roll it out safely and notice when it stops working. Tool names are MLflow 3.16, scikit-learn 1.9 and SciPy 1.18.
Lifecycle & maturity
Loop, not a line: frame (metric, baseline) → data (collect, validate, split like production) → experiment (tracked) → pipeline (repeatable training) → validate (vs champion, slices) → deploy (shadow, canary, A/B) → monitor → retrain or retire.
MLOps = DevOps + two more versioned inputs (data, model) + models that decay without any deploy.
CI (test code and data contracts) · CD (ship pipeline and service) · CT (continuous training produces new models).
Google Cloud level
What is automated
Typical pain
0: manual
nothing; notebook hands a file to engineering
rare releases, no lineage, no monitoring
1: pipeline automation
the training pipeline: retrains on new data with data and model validation
pipeline code changes still manual
2: CI/CD automation
building, testing and deploying the pipeline itself
needs real platform investment
Reproducibility: what to record per run
Input
Record
Tooling
Code
git commit, clean tree
MLflow tags mlflow.source.git.commit when run inside a repo
Data
immutable snapshot path or table version + content hash
dated paths, Delta/Iceberg time travel, DVC, lakeFS
Config
all hyperparameters, feature list, seeds
mlflow.log_params
Environment
image digest or lockfile; CUDA and driver for GPU
MLflow writes requirements.txt / conda.yaml with the model
Artifact
model file + checksum + signature
registry version
“Latest” is not a version. Bit-exact GPU retraining may be impossible (non-deterministic kernels): keep the original artifact.
Registry: named model, immutable numbered versions, lineage to the run. Aliases (@champion, @challenger) are pointers you move; stages are deprecated since 2.9. New versions never move an alias.
Promotion and rollback = move the alias. models:/name/latest changes on every registration: never serve it.
Pipelines & orchestration
DAG: extract window → validate data (schema, nulls, ranges, counts) → features + splits → train (tracked) → evaluate vs champion → register if passing, else alert.
ship preprocessing inside the model (Pipeline, pyfunc)
Online store stale or missing values
freshness monitoring, alert on default-value rate
Feature store: offline store (full timestamped history, for training) + online store (latest per entity, Redis/DynamoDB, for serving) + materialisation job between them.
Log the features actually served, then compare with training distributions.
Serving patterns
Batch
Online
Streaming
Trigger
schedule
request
event
Latency
hours
ms (p99 SLO)
seconds
Typical
churn scores, nightly recs
fraud at checkout, search ranking
real-time features, alerts
Risk
stale for new entities
tail latency, availability
state, ordering
Latency: SLO on p95/p99, not mean. Culprits: feature lookups, cold starts, CPU throttling, oversized inputs. Fixes: batched lookups, warm-up before readiness, quantise/distil, ONNX Runtime or TensorRT, timeouts with fallback.
Dynamic batching on GPUs: more throughput, more queueing delay. Cap the wait.
Rollouts
Strategy
Users see new model?
Measures
Rollback
Shadow
no (mirrored, logged only)
latency, errors, skew, disagreement
turn off mirror
Canary
small share, stepped up
guardrails on real users
traffic weight to 0 / alias back
A/B test
random 50/50 (or other)
causal effect on business metric
stop the test
Blue-green
all at switch
quick swap
switch back
Route by stable hash of user id (zlib.crc32, SHA-256), never Python hash() (randomised per process).
A/B hygiene: pre-registered metric, sample size from minimum effect, no peeking, check sample ratio mismatch.
Rollback needs the old model loadable and its features still produced.
PSI bands: < 0.1 stable, 0.1 to 0.25 moderate, > 0.25 significant. Symmetric; empty bins need epsilon.
KS (scipy.stats.ks_2samp): max CDF gap, numeric only. With large samples p is tiny for trivial shifts: alert on the statistic.
Layered monitoring: service health → data quality → drift → prediction drift → proxy labels → ground truth. Log prediction id, model version, features and timestamp to join late labels.
Retraining triggers
Trigger
Good when
Watch out
Schedule
decay rate is known
retrains for nothing, or too late
Performance drop
labels arrive fast
delayed labels
Drift
labels slow
drift is not always harm
Event
known shifts (launch, pricing)
manual
Every retrained model passes the same gate. Retraining cannot fix broken upstream data.
CI/CD for ML: tests
PR: feature unit tests, data contract tests, smoke training on a fixture, serving test (loads, signature, latency).
Promotion gate: same fresh holdout for champion and challenger, margin above noise, slice and calibration checks, behavioural tests, size/latency budget, fairness where relevant.
Serving: right-size (CPU often fine), quantise/distil, dynamic batching, multi-model GPUs (Triton, MIG), autoscale on queue depth or concurrency, scale to zero for rare models.
Kubernetes: request nvidia.com/gpu, readiness after weights load and warm-up, keep warm replicas (cold start = node + image + weights).
LLMOps basics
Version prompts like code: registry with immutable versions and aliases (prompts:/support-reply@production via mlflow.genai.load_prompt).
The release unit = prompt version + model id + parameters + tools + retrieval config. A provider model change is a release.
Eval gate in CI: versioned dataset of real and adversarial cases; deterministic checks (format, citations), reference metrics, LLM-as-judge (mlflow.genai.evaluate with scorers) calibrated on human labels; compare to production with tolerances; track cost and latency.
Trace every request with prompt and model versions; feed failures back into the eval set.