Ch. 31 · MLOps

Shadow, Canary and A/B Deployments for ML Models

How shadow mode, canary releases and A/B tests de-risk a new model, what each can measure, sticky user bucketing, and rolling back safely.

~8 min readintermediateupdated Oct 6, 2026

“Your new model beats the old one offline. How do you get it into production safely?” Interviewers ask this to see whether you know that an offline win is a hypothesis, not a result. The expected answer names a sequence of rollout strategies, explains what each one can and cannot tell you, and covers the mechanics people get wrong in practice: how users are assigned to variants, which metrics stop a rollout, and what “rollback” actually requires for a model.

Before you start

You should be familiar with offline evaluation (holdout sets, AUC or similar), the idea of a model registry, and basic hypothesis testing. Code is Python 3.14 with SciPy 1.18; outputs in comments are from real runs. The champion is the model currently serving, the challenger the candidate. A guardrail metric is one that must not get worse, such as error rate or latency, as opposed to the primary metric you hope to improve.

The short answer

Roll out in stages. Shadow first: mirror live requests to the challenger, log its predictions, serve only the champion’s. That validates latency, errors, input features and disagreement with zero user risk, but it cannot show how users respond to the new predictions. Then a canary: send a small share of users, say 5%, to the challenger and step up only while guardrails hold, with automatic rollback on a breach. If the goal is a business outcome, run an A/B test: random, sticky assignment, a pre-registered primary metric and a planned sample size. Rollback means moving traffic or a registry alias back to a version that is still loadable and still compatible with today’s features.

How it works

The three strategies answer different questions:

Strategy Users see challenger output? Answers Cannot answer
Shadow no Does it run correctly on real traffic? Do users behave better?
Canary small share, increasing Is it safe for real users? Small effects on business metrics
A/B test planned share, fixed duration What is the causal effect on the outcome? Nothing quickly

All of them depend on assigning requests to a variant. For a canary or A/B test, assignment must be sticky per user: one person should not alternate between models within a session, and both groups should stay stable for the whole experiment. The usual technique hashes the user id with an experiment-specific salt:

import hashlib
from collections import Counter

def bucket(user_id: str, experiment: str) -> int:
    digest = hashlib.sha256(f"{experiment}:{user_id}".encode()).digest()
    return int.from_bytes(digest[:8], "big") % 100

def variant(user_id: str, experiment: str, canary_pct: int) -> str:
    return "challenger" if bucket(user_id, experiment) < canary_pct else "champion"

print(bucket("user-42", "churn-v8"), bucket("user-42", "churn-v8"), bucket("user-42", "ranker-v3"))
# 48 48 5
counts = Counter(variant(f"user-{i}", "churn-v8", 5) for i in range(100_000))
print(counts["challenger"], counts["champion"])     # 4992 95008
python

The same user always gets the same bucket for one experiment, different experiments get independent buckets because of the salt, and raising the percentage from 5 to 25 keeps everyone already in the canary inside it, since buckets 0 to 4 are a subset of 0 to 24. Python’s built-in hash() would break all of this: string hashing is randomised per process, so four API replicas would disagree about every user.

Step-by-step walkthrough

Step 1: Shadow the challenger off the critical path

The live request must never wait for, or fail because of, the shadow model. Mirror asynchronously and swallow every shadow error:

import asyncio, time

async def champion(x):   await asyncio.sleep(0.010); return 0.12
async def challenger(x): await asyncio.sleep(0.300); return 0.31   # slow new model

shadow_log, background = [], set()

async def shadow_call(x, live_score):
    try:
        s = await asyncio.wait_for(challenger(x), timeout=1.0)
        shadow_log.append({"live": live_score, "shadow": s})
    except Exception as e:                       # never reaches the user
        shadow_log.append({"live": live_score, "error": type(e).__name__})

async def handle(x):
    live = await champion(x)
    task = asyncio.create_task(shadow_call(x, live))   # fire and forget
    background.add(task); task.add_done_callback(background.discard)
    return live
# handle() returned 0.12 in 11 ms although the challenger takes 300 ms
python

In production the mirror is often done by the infrastructure instead (Istio traffic mirroring, or a queue that a separate shadow service consumes). Compare the shadow log with the live log: latency distribution, error rate, feature null rates, score distribution and the share of decisions that would flip.

Step 2: Define guardrails and the canary plan before starting

Write down the steps (5%, 25%, 50%, 100%), how long each lasts, and the rule that aborts. A rule needs both a practical threshold and enough evidence, or noise will trigger rollbacks:

from scipy.stats import chi2_contingency

def canary_check(base_err, base_n, can_err, can_n, max_ratio=1.2, alpha=0.01):
    table = [[base_err, base_n - base_err], [can_err, can_n - can_err]]
    p = chi2_contingency(table).pvalue
    ratio = (can_err / can_n) / (base_err / base_n)
    return ("rollback" if ratio > max_ratio and p < alpha else "continue"), round(ratio, 2), f"{p:.2g}"

print(canary_check(2_280, 380_000, 180, 20_000))   # ('rollback', 1.5, '1.6e-07')
print(canary_check(2_280, 380_000, 126, 20_000))   # ('continue', 1.05, '0.63')
python

On Kubernetes, Argo Rollouts and Flagger automate exactly this loop: shift weight, query metrics, promote or abort.

Step 3: Run an A/B test when the question is business impact

Pick one primary metric (conversion, revenue per user, loss rate), compute the sample size from the smallest effect worth detecting, and run for the planned duration, usually whole weeks. Do not stop the first time p dips below 0.05; repeated peeking inflates false positives. Check the sample ratio: a 50/50 test with a noticeably uneven split means assignment or logging is broken.

Step 4: Keep rollback real

Rollback works only if the previous model is still deployed or instantly loadable, and if its inputs still exist. With a registry, rollback is moving the champion alias back or setting the challenger’s traffic weight to zero. If the new model required a new feature and the old feature pipeline was deleted, there is nothing to roll back to.

Worked scenario

A team released a new ranking model to 5% of traffic using random.random() < 0.05 per request. After a week the canary looked great: 4% higher click-through. They went to 100% and click-through fell. Two things had gone wrong. Per-request assignment meant most users saw both models within a session, so clicks on the challenger’s results were partly driven by items the champion had shown a minute earlier. And the canary group was dominated by heavy users, simply because they made more requests, so the comparison was not like for like.

They restarted with salted SHA-256 bucketing on user id, as above, which made assignment sticky and per user. They added a sample ratio check to the dashboard; one day it caught the canary at 19,000 of 400,000 users instead of the configured 5%:

from scipy.stats import chisquare
print(f"{chisquare([19_000, 381_000], f_exp=[20_000, 380_000]).pvalue:.2g}")   # 4e-13
python

The cause was a mobile client version that bypassed the router. With that fixed, a two-week A/B test showed a real but smaller gain of 1.2%, and the rollout went ahead with the old model kept loadable for a month.

Common mistake

  • “Shadow mode proved users like it.” Shadow outputs are never shown, so it says nothing about user behaviour.
  • Random assignment per request. Users flip between variants, which contaminates both groups.
  • Choosing metrics after seeing the data. With enough metrics, one will look significant by chance.
  • Treating a canary as an experiment. A 5% canary for a day is a safety check; it is usually too small and short to measure modest business effects.
  • Assuming rollback is free. Models depend on features, schemas and downstream thresholds; all must still match.

Verify the behavior

Three checks belong in tests. Bucketing is deterministic: calling bucket() twice, or in two processes, returns the same value, while python -c "print(hash('user-42') % 100)" run three times printed 97, 91 and 29 for us. Growth is monotonic: users in the 5% canary are a subset of the 25% canary (our check printed True with 25,239 users at 25%). Shadow isolation: with a challenger that sleeps for 300 ms, the live handler still returns in about 10 ms, and with a challenger that raises, the live response is unchanged and the error is logged.

Follow-up questions

  • When would you use a multi-armed bandit instead of an A/B test? When exploiting the better variant during the test matters more than a clean effect estimate, for example short-lived promotions.
  • What is interleaving? For rankers, mixing results from both models in one list and counting which model’s items get clicked; it needs far fewer users than an A/B test.
  • How do you handle delayed outcomes, such as loan defaults? Use guardrails and proxy metrics for the canary, and keep a long-running holdout to measure the true outcome later.
  • What is blue-green deployment? Two full environments with an instant switch between them; quick rollback, but everyone moves at once, so it does not limit exposure the way a canary does.

Interview exercise

You are replacing a fraud model. The challenger catches more fraud offline at the same false-positive rate, but it uses a new device-fingerprint feature from a vendor API with occasional timeouts. Chargeback labels arrive after about 45 days. Design the rollout.

Answer and reasoning

Start with two weeks of shadow mode to measure the vendor API’s timeout rate under real load, the challenger’s latency and how often its decisions differ from the champion’s; the timeout behaviour decides whether the new feature needs a fallback value the model was trained to handle. Then a canary by hashed card or customer id at 5% and 25%, guarded by decline rate, manual review volume, latency and feature fallback rate, because real fraud labels will not arrive in time to guard anything. Keep the champion loadable and the old features flowing so rollback is an alias move. Finally keep a small holdout on the champion for at least 45 days so chargeback rates can be compared once labels arrive. The real verdict on fraud caught comes from that comparison, not from the canary.

Continue learning

More in MLOps

esc