“What does CI/CD look like for a machine learning system?” is where MLOps interviews separate people who have deployed one model from people who have kept several running. A good answer explains why ordinary CI/CD is not enough, names the extra loop (continuous training), lists concrete tests, and describes the automated gate that decides whether a newly trained model may replace the current one. Senior candidates are then asked when retraining should happen and how they would stop automation from shipping a regression.
Before you start
You should know what a CI pipeline does for normal software (run tests on each pull request, build an artifact, deploy it), how a model is trained and evaluated, and what a model registry is. Code was run with Python 3.14, pandas 3.0, scikit-learn 1.9 and pytest; outputs in comments are from real runs. A champion is the model in production; a challenger is the newly trained candidate.
The short answer
ML systems need three loops. CI runs on every code change: unit tests for feature logic, data contract tests, a small smoke training run and serving tests. CD deploys the training pipeline and the serving service. Continuous training (CT) runs the deployed pipeline on new data, on a schedule or a trigger, and produces challenger models. A challenger is registered only after an automated gate compares it with the champion on the same fresh holdout, overall and on important slices, with latency and size limits, and it then goes through shadow or canary before promotion. Google Cloud’s guide describes this progression as level 0 (manual), level 1 (automated training pipeline) and level 2 (CI/CD for the pipeline itself).
How it works
In ordinary software, the commit determines the artifact. In ML, the artifact (the model) also depends on data that changes without any commit, so there are two kinds of release. A pipeline release happens when code changes: CI tests it, CD deploys the new pipeline version. A model release happens when the deployed pipeline trains on new data: the gate tests the model, and the rollout process promotes it.
The pipeline is a sequence of steps, each testable on its own:
SCHEMA = {"amount": "float64", "country": "str", "orders_30d": "int64", "label": "int64"}
def validate(df: pd.DataFrame) -> list[str]:
problems = []
for col, dtype in SCHEMA.items():
if col not in df.columns:
problems.append(f"missing column {col}")
elif str(df[col].dtype) != dtype:
problems.append(f"{col}: {df[col].dtype} != {dtype}")
if "amount" in df and df["amount"].isna().mean() > 0.01:
problems.append(f"amount null rate {df['amount'].isna().mean():.0%}")
if len(df) < 1_000:
problems.append(f"only {len(df)} rows")
return problems
print(validate(broken)) # ['missing column orders_30d', 'amount null rate 50%']The full DAG is: extract a data window, validate it, build features and splits, train with tracking, evaluate against the champion, register if the gate passes, otherwise stop and alert. An orchestrator (Airflow, Kubeflow Pipelines, Vertex AI Pipelines, Dagster, Prefect) runs it with retries and passes artifacts between steps. Each step is containerised and idempotent, so a failed run can resume.
Step-by-step walkthrough
Step 1: Test the code on every pull request
CI cannot run full training on every commit, but it can prove the pipeline works end to end on a tiny fixture:
from pipeline import make_data, validate, train, evaluate, gate
def test_schema_violation_is_reported():
df = make_data(2_000, 0).drop(columns="orders_30d")
assert "missing column orders_30d" in validate(df)
def test_smoke_training_learns_something():
train_df, holdout = make_data(3_000, 0), make_data(1_000, 1)
assert not validate(train_df)
assert evaluate(train(train_df), holdout)["overall"] > 0.65
def test_gate_blocks_slice_regression():
ok, reasons = gate({"overall": 0.81, "IN": 0.62}, {"overall": 0.80, "IN": 0.70})
assert not ok and reasons == ["slice IN: 0.620 vs 0.700"]
# 3 passed in 2.92sAdd unit tests for feature functions (nulls, time zones, empty inputs) and a serving test that loads the packaged model and calls it with the expected signature. A minimal GitHub Actions workflow runs them:
name: ml-ci
on:
pull_request:
paths: ["pipeline/**", "tests/**", "requirements.txt"]
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.13"
cache: pip
- run: pip install -r requirements.txt
- run: pytest -qStep 2: Gate every trained model
The gate compares challenger and champion on the same, recent holdout neither model trained on. It requires a margin larger than run-to-run noise, and it checks slices, because an overall gain can hide a segment getting worse:
def gate(challenger: dict, champion: dict, min_gain=0.005, max_slice_drop=0.01):
reasons = []
if challenger["overall"] < champion["overall"] + min_gain:
reasons.append(f"overall {challenger['overall']:.3f} vs {champion['overall']:.3f}")
for k in champion:
if k != "overall" and challenger[k] < champion[k] - max_slice_drop:
reasons.append(f"slice {k}: {challenger[k]:.3f} vs {champion[k]:.3f}")
return (not reasons), reasonsAdd calibration if scores are used as probabilities, latency and artifact size limits, and behavioural tests on known hard cases. A model that passes is registered with a challenger alias and moves to shadow or canary; promotion to champion happens only after live checks.
Step 3: Choose retraining triggers
Retrain on a schedule matched to how fast the model decays, plus triggers for measured degradation and persistent drift:
def should_retrain(today, last_trained, psi_top_features, live_auc, baseline_auc):
if live_auc is not None and live_auc < baseline_auc - 0.03:
return True, f"performance: live AUC {live_auc:.3f} vs {baseline_auc:.3f}"
drifted = [f for f, v in psi_top_features.items() if v > 0.25]
if drifted:
return True, f"drift: {', '.join(drifted)}"
if (today - last_trained).days >= 28:
return True, "schedule: 28 days since last training"
return False, "no trigger"
# (True, 'drift: orders_30d') with PSI 0.31 on orders_30dWhatever fires, the result goes through the same validation and gate. Retraining is a request to produce a candidate, not a decision to deploy one.
Step 4: Move up the maturity levels only as far as needed
At level 0 a data scientist trains in a notebook and hands over a file. At level 1 the pipeline is automated and deployed, so the system retrains on new data with automated data and model validation; you deploy the pipeline, not just a model. At level 2 changes to the pipeline code are themselves built, tested and deployed automatically, letting teams try new features and architectures quickly. A model retrained twice a year may be fine at level 1; a platform with dozens of models and frequent modelling changes needs level 2.
Worked scenario
A payments company retrained its fraud model nightly and deployed the result automatically if overall AUC did not drop. One week, an ingestion change dropped most rows from India, a market where customers with many recent orders were riskier rather than safer, the opposite of other countries. The nightly model trained on mostly non-India data. On the holdout its overall AUC was 0.806 against the champion’s 0.803, so the “no drop” rule passed it, and it shipped.
Re-running the gate from Step 2 on the same models tells the real story:
# champion {'overall': 0.803, 'DE': 0.71, 'IN': 0.703, 'US': 0.715}
# buggy {'overall': 0.806, 'DE': 0.734, 'IN': 0.623, 'US': 0.727}
print(gate(m_bug, m_champ))
# (False, ['overall 0.806 vs 0.803', 'slice IN: 0.623 vs 0.703'])The improvement is inside the noise margin, and the India slice fell by 0.08. A correctly trained challenger on the full data scored 0.816 overall and improved every slice, and passed. The team added three things: a row-count-per-country check in data validation, which would have stopped the run before training, the slice-aware gate with a minimum gain, and a canary step before any nightly model reached full traffic.
Common mistake
- Automating retraining without automating validation. That just ships regressions faster.
- Comparing models on different holdouts. The challenger’s validation split and the champion’s old test score are not comparable.
- “Any improvement wins.” Differences of 0.001 AUC are noise; set a margin from repeated runs with different seeds.
- Only overall metrics. Slices are where the damage hides.
- Full training in CI on every commit. Too slow and too expensive; use a smoke run on a fixture and keep full training in the pipeline.
Verify the behavior
Run pytest -q locally and in CI; the three tests above passed in about three seconds. Then break things on purpose: remove a column from the fixture and confirm the validation test fails, train a deliberately weaker model and confirm the gate rejects it, and feed the trigger function a PSI of 0.31 and confirm it requests retraining. A gate that has never rejected anything has not been tested.
Follow-up questions
- How do you test a pipeline step that reads from the warehouse? Separate the I/O from the logic: test the logic on fixtures, and run a small integration test against a staging dataset.
- How do you estimate the noise margin? Train the same configuration with several seeds or bootstrap the holdout, and use the spread of the metric.
- Where does a feature store fit? It gives training and serving the same feature definitions, and level 1 pipelines usually read from it.
- What do you do when the gate keeps rejecting every new model? Investigate rather than loosen it: data quality, a changed label definition, or a champion evaluated on an easier period.
Interview exercise
Your churn model is retrained monthly by a data scientist in a notebook, exported as a pickle and deployed by an engineer through a ticket. It takes three weeks and nobody can say which data trained the current model. Outline the first three improvements you would make, in order, and why.
Answer and reasoning
First, tracking and a registry: every training run logs params, metrics, a data snapshot id and the git commit, and the deployed model is loaded by registry alias. That gives lineage and turns deployment into an alias move instead of a ticket. Second, convert the notebook into a pipeline with data validation, training, evaluation against the champion and registration, scheduled monthly by an orchestrator, which is level 1 and removes the manual handover. Third, CI on the pipeline code with unit tests, data contract tests and a smoke training run, so changes stop breaking it. Monitoring and canary rollouts come next. This order fixes the most painful problems first: nobody knowing what is deployed, then the slow manual release, then the fragility of changes.
Continue learning
- Practise in the MLOps chapter and the MLOps MCQs.
- Learn which drift signals should trigger retraining in data drift vs concept drift, and how gated models reach users in shadow, canary and A/B deployments.
- Reference: Google Cloud’s MLOps continuous delivery and automation pipelines and the Kubeflow Pipelines documentation.