Batch vs Online Inference: Choosing How to Serve a Model
When to score models in a nightly batch job and when behind a low-latency API, with tail-latency budgets, fallbacks and the hybrid pattern.
MLOps interviews: training pipelines, experiment tracking, model registries, batch and real-time serving, monitoring, data and concept drift, feature stores and retraining.
Official reference: Google Cloud: MLOps pipelines
When to score models in a nightly batch job and when behind a low-latency API, with tail-latency budgets, fallbacks and the hybrid pattern.
How CI, CD and continuous training fit together for ML, which tests gate a pipeline, retraining triggers, and Google's MLOps maturity levels.
How data drift, concept drift and prediction drift differ, how PSI and the KS test measure them, and why p-value alerts drown your team in noise.
Treat prompts as versioned releases: register and alias them in MLflow, trace every call, and block regressions with an evaluation gate in CI.
Log params, metrics and models with MLflow 3, compare runs with search_runs, and promote or roll back models by moving registry aliases.
What it takes to rebuild a model months later: code commit, data snapshot and hash, config, seeds, environment and the registered artifact.
How shadow mode, canary releases and A/B tests de-risk a new model, what each can measure, sticky user bucketing, and rolling back safely.
Why models that ace offline tests fail live: duplicated feature code, refitted scalers, future-leaking joins, and how feature stores stop them.
MLOps is the set of practices that gets machine learning models into production and keeps them working there: versioning, automated training pipelines, testing, deployment, monitoring and retraining.
It borrows CI/CD, infrastructure as code and observability from DevOps, but an ML system has two extra moving parts besides code: data and the trained model. That changes three things:
The goal is that a new model reaches production through an automated, auditable path rather than a notebook and a copied pickle file.
Likely follow-up: What is the smallest MLOps setup you would recommend for a team shipping its first model? · Who owns a model in production: data science or platform?
The important point is that it is a loop, not a line. In a mature team most of the effort goes into everything after the first experiment: pipelines, validation, rollout and monitoring.
Likely follow-up: Where do most ML projects fail in this lifecycle? · When would you retire a model instead of retraining it?
Batch inference scores many rows on a schedule, for example a nightly job that writes churn scores for every customer into a table. Online inference scores one request at a time behind an API, within a latency budget, for example fraud checks during checkout.
Choose by asking when the inputs become known and how fresh the prediction must be. If the decision can use yesterday's data and the set of entities is known in advance, batch is cheaper, simpler and easier to debug: you can inspect the output before anyone uses it. If the prediction depends on information only available at request time (the basket, the search query, the transaction amount) or the entity space is too large to precompute, you need online serving, with its latency, scaling and availability concerns.
A common middle ground is precomputing in batch and looking the result up online, or streaming features with online scoring.
Likely follow-up: What would make you move a batch model to online serving? · How do you handle a user who is new since the last batch run?
Experiment tracking records every training run so you can compare, reproduce and audit results instead of relying on notebook names and memory.
For each run I log:
mlflow.log_params.step argument of mlflow.log_metric.Runs are grouped into experiments and can be queried, for example mlflow.search_runs(filter_string="metrics.val_auc > 0.9"). Autologging (mlflow.autolog()) covers many frameworks, but I still log the data version myself, because nothing logs that automatically.
Likely follow-up: What happens if you call log_param twice with different values in one run? · How would you track a hyperparameter search with hundreds of trials?
A model registry is a catalogue of named models with numbered, immutable versions, each linked to the run that produced it, so you know the code, parameters, metrics and data behind it.
A bucket of files gives you storage but none of the workflow. The registry adds:
abc123 and everything logged there.@champion or @challenger. Serving loads models:/churn-model@champion, and promotion or rollback is just moving the alias to another version.MLflow's older fixed stages (Staging, Production, Archived) are deprecated in favour of aliases and tags since MLflow 2.9, because aliases are flexible and several can point at different versions at once.
Likely follow-up: How does a serving process find out the champion alias has moved? · Would you use one registry per environment or one shared registry?
Data drift (covariate shift) is a change in the distribution of inputs, P(X). A loan model trained on applicants aged 30 to 50 starts receiving mostly students. The relationship between features and outcome may still hold, but the model is now working in regions it saw little of.
Concept drift is a change in the relationship between inputs and target, P(y given X). The same transaction pattern that was legitimate last year is now typical of a new fraud scheme. Inputs can look identical while the right answer has changed.
The practical difference is detection. Data drift can be measured immediately by comparing feature distributions with a reference window (PSI, KS test). Concept drift usually needs labels, which often arrive late, so you watch proxies such as prediction drift and business metrics until ground truth arrives.
Not every data drift hurts accuracy, so drift alerts should trigger investigation, not automatic retraining.
Likely follow-up: Give an example of data drift that does not hurt the model. · What is label shift and how does it differ from both?
Because a model is a function of code and data. If the training table is overwritten every night, you cannot rebuild last month's model, explain a regression or show an auditor what the model learned from.
Options, from simplest:
s3://ml/churn/2026-10-01/) and log that path and a content hash with the run.Whichever you use, the run must record the exact version it read. A dataset name such as customers_latest is not a version.
Likely follow-up: How do you version a 5 TB dataset cheaply? · How do data retention or deletion requests interact with data versioning?
I package three things together: the model artifact, the exact inference code including preprocessing, and the environment.
Pipeline or an MLflow pyfunc model, so the serving side cannot apply different feature logic.infer_signature and validates requests against it.requirements.txt and conda.yaml), because a pickle loaded under a different scikit-learn version can fail or behave differently.Then the same image is promoted through environments, never rebuilt per environment.
Likely follow-up: Would you bake the model into the image or download it at start-up? · Why is pickle risky as a model format?
Five things, all linked to the registered model version:
Even then, bit-for-bit reproduction may be impossible: GPU kernels can be non-deterministic and parallel reductions sum in different orders. So I would say what I can promise: the same artifact exactly, and a retrained model within a stated tolerance on the original evaluation set.
Likely follow-up: How would you make PyTorch training as deterministic as possible? · How long would you keep training snapshots?
Training-serving skew is any difference between the features or behaviour a model saw in training and what it gets in production. The model looks good offline and underperforms live, with no error anywhere.
Common causes:
Prevention: one feature definition used by both paths (a feature store or shared library), preprocessing shipped inside the model artifact, point-in-time correct training sets, and logging the features actually served so you can compare them with training values.
Likely follow-up: How would you detect skew that already exists in production? · Is training-serving skew the same as drift?
A feature store manages feature definitions once and serves them to both training and inference. Feast, Tecton, Databricks and Vertex AI all follow the same shape.
The value is consistency and reuse: the same definition produces training and serving values, and teams share features instead of rebuilding them. The cost is another system to operate, so a team with one batch model may not need one.
Likely follow-up: How would you build a point-in-time join without a feature store? · What happens if the materialisation job falls behind?
A notebook depends on hidden state, execution order and one person's laptop. A pipeline makes training repeatable, schedulable, reviewable and observable.
A typical training pipeline is a DAG of steps:
Orchestrators such as Airflow, Kubeflow Pipelines, Vertex AI Pipelines, Dagster or Prefect handle scheduling, retries, dependencies and passing artifacts between steps. Each step should be a containerised, idempotent unit with explicit inputs and outputs, so a failed run can resume and any step can be rerun with the same inputs.
Likely follow-up: How do you test a pipeline step? · Airflow or Kubeflow Pipelines: what would decide it for you?
PSI bins the reference distribution (usually into deciles), measures the share of current data in each bin and sums (cur - ref) * ln(cur / ref). Common rules of thumb: below 0.1 stable, 0.1 to 0.25 moderate shift, above 0.25 significant. It works for numeric and categorical features and is symmetric.
The two-sample KS test compares empirical cumulative distributions; the statistic is the largest vertical gap between them, and scipy.stats.ks_2samp returns it with a p-value. It only applies to continuous numeric features.
Pitfalls:
import numpy as np
def psi(reference, current, bins=10, eps=1e-6):
edges = np.quantile(reference, np.linspace(0, 1, bins + 1))
edges[0], edges[-1] = -np.inf, np.inf
ref = np.histogram(reference, edges)[0] / len(reference)
cur = np.histogram(current, edges)[0] / len(current)
ref, cur = np.clip(ref, eps, None), np.clip(cur, eps, None)
return float(np.sum((cur - ref) * np.log(cur / ref)))Likely follow-up: How would you monitor drift for a high-cardinality categorical feature? · Which reference window would you compare against?
I would monitor in layers, from fastest to slowest signal.
This only works if every prediction is logged with its id, model version, input features and timestamp, so later labels can be joined back.
Likely follow-up: How would you detect calibration drift? · What would you alert on versus only show on a dashboard?
In a shadow deployment the new model receives a copy of live traffic and makes predictions, but its outputs are only logged; users still get the current model's answer.
It is good for checking things that offline tests miss, with zero user risk:
What it cannot measure is the effect of acting on its predictions. If the new recommender would show different items, you never see whether users click them, because they were never shown. For that you need a canary or an A/B test.
Implementation detail: mirror the request asynchronously so the shadow model can never add latency or failures to the live path, and watch the extra cost.
Likely follow-up: How would you implement traffic mirroring on Kubernetes? · How do you compare outputs when labels are delayed?
Route a small share of real traffic, say 5%, to the new version, compare it with the current version over a fixed window, then step up (5, 25, 50, 100%) only if guardrails hold.
Guardrails are defined before the rollout: error rate, p99 latency, prediction distribution, and fast business proxies such as approval rate or click-through. Ideally an automated analysis compares canary against baseline and aborts on breach (Argo Rollouts and Flagger do this on Kubernetes).
For rollback, the old version must still be running or instantly loadable. With a registry, rollback is moving the champion alias back to the previous version, or shifting the traffic weight back to 0. Keep the previous model's feature pipeline and schema compatible, otherwise rollback breaks.
Route by a stable hash of user id rather than per request, so one user gets consistent behaviour and metrics are not mixed within a session.
Likely follow-up: How long should each canary step last? · What makes a model rollback harder than a code rollback?
Offline AUC measures ranking on historical data. The business cares about an outcome such as revenue, conversion or losses, and the two can disagree: the new model may be slower, may be better only on segments that do not matter, or may change user behaviour in ways the historical data cannot show.
An A/B test randomly assigns users (by a stable hash of user id) to the champion or challenger and compares a primary metric chosen in advance, with guardrail metrics such as latency, complaints or fairness.
The statistics matter: compute the sample size from the minimum effect you care about, run for the planned duration (often whole weeks to cover weekly cycles), and avoid stopping as soon as the p-value dips below 0.05, which inflates false positives. Check sample ratio mismatch: if the split was 50/50 but you see 52/48, the assignment is broken and the result is not trustworthy.
Likely follow-up: When would you use a multi-armed bandit instead of an A/B test? · What if offline and online results disagree?
Each trigger has a place.
In practice I start with a schedule plus alerts, and whichever trigger fires, the retrained model still goes through the same validation gate against the champion. Automatic retraining without automatic validation just automates shipping regressions. Also note that retraining cannot fix a broken upstream feature; check data quality first.
Likely follow-up: How much history should each retrain use? · How would you detect that retraining made things worse?
There are three loops. CI runs on every change to code or pipeline definitions. CD deploys the pipeline and the serving service. CT (continuous training) runs the pipeline to produce new models, which then go through their own promotion.
Tests on each pull request:
Model promotion tests, run after training:
Only a model that passes is registered and moves to shadow or canary.
Likely follow-up: How do you keep a smoke training run fast? · What is a behavioural test for a model?
First, measure where the time goes for slow requests: feature lookup, preprocessing, model execution or serialization. Averages hide this; the mean can look fine while the tail ruins checkout.
Common causes and fixes:
On GPU servers, dynamic batching raises throughput but adds queueing delay, so cap the batch wait time to protect the tail.
Likely follow-up: How would you set a timeout and fallback for the model call itself? · How does dynamic batching trade latency for throughput?
The point I would add in an interview: most teams do not need level 2 for every model. Choose based on how often the data changes and how often you change the modelling approach.
Likely follow-up: Which component would you add first to move from level 0 to level 1? · What breaks when a level 0 team scales to ten models?
Compare challenger and champion on the same, fresh holdout data that neither was trained on, ideally the most recent period.
If everything passes, register and alias it as challenger and send it to shadow or canary; promotion to champion happens only after live checks.
Likely follow-up: How do you estimate run-to-run noise? · What if the challenger wins overall but loses badly on one slice?
Treat data as an interface with a contract, and validate it at every boundary.
Tools such as Great Expectations, TensorFlow Data Validation, Pandera or dbt tests implement these checks. The organisational fix matters as much: a named owner for every upstream table the model reads.
Likely follow-up: Where would you put validation: training, serving or both? · How strict should checks be before they cause alert fatigue?
First measure utilisation per workload; idle or underused GPUs are usually the biggest waste.
Training:
Serving:
Then make cost visible: tag spend per team and model, and add cost per 1,000 predictions to the model review.
Likely follow-up: What breaks when training on spot instances? · When is CPU inference the right choice?
It happens when a model's predictions influence the data it is later trained on. A recommender only shows items it already ranks highly, users can only click what they see, and the next model learns that those items are even better. Popular items get more popular; new items never get a chance. Fraud models are similar: blocked transactions never get a real outcome label, so the model never learns whether it was right.
Mitigations:
The key interview point is that offline evaluation on logged data inherits the loop, so it can look great while the system narrows.
Likely follow-up: How do you evaluate a new recommender offline when the data came from the old one? · How big should a holdout group be?
I would rule out the cheap causes first, in order.
Mitigation depends on the cause: roll back, fix the feature, retrain on recent data, or adjust the threshold as a stopgap while the real fix lands.
Likely follow-up: What would you have wanted in place before this happened? · When is adjusting the threshold an acceptable fix?
In an LLM app the prompt is effectively the model's code: a one-word change can alter tone, refusals, output format or cost. If prompts live as strings scattered in application code, you cannot tell which prompt produced a bad answer or roll back quickly.
What I would do:
prompts:/support-reply@production, so promotion and rollback do not need a code deploy.import mlflow
p = mlflow.genai.register_prompt(
name="support-reply",
template="You are a support agent. Cite the policy. Question: {{question}}",
commit_message="cite policy",
)
mlflow.genai.set_prompt_alias("support-reply", alias="production", version=p.version)
prompt = mlflow.genai.load_prompt("prompts:/support-reply@production")
print(prompt.format(question="Where is my refund?"))Likely follow-up: How would you A/B test two prompt versions? · What do you do when the model provider deprecates the model you use?
Build a versioned evaluation dataset of real, anonymised inputs: common cases, known failures, adversarial and policy-sensitive examples, each with an expected answer or grading criteria.
On every change to a prompt, model, retrieval setting or tool, CI runs the candidate over the dataset and scores it with:
The gate compares with the current production version: fail if a key score drops by more than a set tolerance, or any safety check fails. Because outputs are non-deterministic, use temperature 0 where possible, repeat flaky cases and set tolerances from measured variance. Production traces feed new failure cases back into the dataset.
Likely follow-up: How do you know the LLM judge is reliable? · How large does the evaluation set need to be?
Features: static customer features from batch jobs, plus streaming aggregates (transactions in the last 10 minutes, distinct merchants today) computed with Kafka and Flink and written to an online store such as Redis. The same definitions backfill the offline store with timestamps, so training uses point-in-time joins.
Serving: a stateless scoring service, autoscaled, loading the model by registry alias. One batched feature lookup, a gradient-boosted model (fast on CPU), strict timeouts, and a rules-based fallback if features or the model are unavailable, because checkout must not fail.
Logging: every request's features, score, model version and decision, keyed by transaction id.
Labels: chargebacks arrive weeks later; join them back. Blocked transactions have no outcome, so keep a tiny randomised review sample to avoid a feedback loop.
Lifecycle: weekly retraining with a champion gate, shadow then canary rollouts, drift and approval-rate monitoring, and alerting on fallback rate.
Likely follow-up: How do you pick the decision threshold? · How would you handle a sudden new fraud pattern before retraining?
GPU inference servers barely use the CPU, so the CPU percentage says almost nothing about load. GPU utilisation is also misleading: a GPU can report 100% while serving comfortably, and with batching more traffic raises throughput before it hurts latency.
Better scaling signals are queue depth, in-flight requests per replica (concurrency) or latency against the target. KServe with Knative scales on concurrency, and KEDA or a custom-metrics HPA can scale on Prometheus metrics from the server.
Other things that matter:
nvidia.com/gpu: 1), since GPUs are not shared by default; use MIG or time-slicing for small models.Likely follow-up: How does Knative scale-to-zero interact with large models? · How would you share one GPU between several small models?
No questions match that filter.
Prefer multiple choice? All 20 MLOps MCQs with answers →