Ch. 31

MLOps interview questions & answers

MLOps interviews: training pipelines, experiment tracking, model registries, batch and real-time serving, monitoring, data and concept drift, feature stores and retraining.

30 interview questions20 quiz questions8 notes
your progress0%

Notes in this chapter

Filter all notes →

30 MLOps interview questions study by subtopic

30 questions
  1. 1.What is MLOps, and how is it different from ordinary DevOps?easy

    MLOps is the set of practices that gets machine learning models into production and keeps them working there: versioning, automated training pipelines, testing, deployment, monitoring and retraining.

    It borrows CI/CD, infrastructure as code and observability from DevOps, but an ML system has two extra moving parts besides code: data and the trained model. That changes three things:

    • A build is not reproducible from the commit alone; you also need the data snapshot, the parameters and the environment.
    • Tests must cover data (schema, ranges, freshness) and model quality, not just code paths.
    • A model can degrade with no deploy at all, because the world it predicts changes. So production needs drift and quality monitoring, and pipelines need continuous training as well as continuous delivery.

    The goal is that a new model reaches production through an automated, auditable path rather than a notebook and a copied pickle file.

    What interviewers listen for
    • Data and model are versioned artifacts alongside code
    • Tests cover data and model quality, not just code
    • Models decay without deploys, so monitoring is essential
    • Adds continuous training to CI/CD

    Likely follow-up: What is the smallest MLOps setup you would recommend for a team shipping its first model? · Who owns a model in production: data science or platform?

  2. 2.Walk me through the lifecycle of a machine learning model, from problem to retirement.easy
    • Frame the problem: the business decision, the prediction target, the metric and a baseline (often a rule or the current process).
    • Data: collect, label, validate and split the data the way production will see it.
    • Experiment: feature engineering, model selection and tuning, with every run tracked.
    • Train in a pipeline: the winning recipe becomes a repeatable pipeline that produces a versioned model.
    • Validate: compare with the current champion on holdout data and important slices, plus latency and size checks.
    • Deploy: batch job or online service, rolled out through shadow, canary or A/B testing.
    • Monitor: service health, input drift, prediction drift and, once labels arrive, real quality.
    • Retrain or retire: on a schedule or a trigger, feeding back into the training pipeline.

    The important point is that it is a loop, not a line. In a mature team most of the effort goes into everything after the first experiment: pipelines, validation, rollout and monitoring.

    What interviewers listen for
    • Starts with a metric and a baseline
    • Experimentation becomes a repeatable pipeline
    • Validation gate before deployment
    • Monitoring feeds retraining: it is a loop

    Likely follow-up: Where do most ML projects fail in this lifecycle? · When would you retire a model instead of retraining it?

  3. 3.What is the difference between batch inference and online inference? How do you choose?easy

    Batch inference scores many rows on a schedule, for example a nightly job that writes churn scores for every customer into a table. Online inference scores one request at a time behind an API, within a latency budget, for example fraud checks during checkout.

    Choose by asking when the inputs become known and how fresh the prediction must be. If the decision can use yesterday's data and the set of entities is known in advance, batch is cheaper, simpler and easier to debug: you can inspect the output before anyone uses it. If the prediction depends on information only available at request time (the basket, the search query, the transaction amount) or the entity space is too large to precompute, you need online serving, with its latency, scaling and availability concerns.

    A common middle ground is precomputing in batch and looking the result up online, or streaming features with online scoring.

    What interviewers listen for
    • Batch: scheduled, many rows, results stored
    • Online: per request, latency budget
    • Decide by input availability and freshness
    • Precompute-then-lookup hybrid

    Likely follow-up: What would make you move a batch model to online serving? · How do you handle a user who is new since the last batch run?

  4. 4.What is experiment tracking, and what would you log for each training run in a tool such as MLflow?easy

    Experiment tracking records every training run so you can compare, reproduce and audit results instead of relying on notebook names and memory.

    For each run I log:

    • Parameters: hyperparameters and important config, with mlflow.log_params.
    • Metrics: validation scores, plus per-epoch curves using the step argument of mlflow.log_metric.
    • Code version: the git commit (MLflow tags it automatically when the script runs inside a repository).
    • Data version: a dataset hash, snapshot id or DVC revision.
    • Environment: library versions, captured when the model is logged.
    • Artifacts: the model with its signature, plots, confusion matrices, feature importances.

    Runs are grouped into experiments and can be queried, for example mlflow.search_runs(filter_string="metrics.val_auc > 0.9"). Autologging (mlflow.autolog()) covers many frameworks, but I still log the data version myself, because nothing logs that automatically.

    What interviewers listen for
    • Params, metrics, code, data, environment, artifacts
    • Metric history via step
    • Data version must be logged explicitly
    • Runs are queryable for comparison

    Likely follow-up: What happens if you call log_param twice with different values in one run? · How would you track a hyperparameter search with hundreds of trials?

  5. 5.What is a model registry and why not just store model files in S3?easy

    A model registry is a catalogue of named models with numbered, immutable versions, each linked to the run that produced it, so you know the code, parameters, metrics and data behind it.

    A bucket of files gives you storage but none of the workflow. The registry adds:

    • Lineage: version 7 points to run abc123 and everything logged there.
    • Deployment pointers: in MLflow, aliases such as @champion or @challenger. Serving loads models:/churn-model@champion, and promotion or rollback is just moving the alias to another version.
    • Governance: descriptions, tags, approvals and an audit trail of who promoted what.

    MLflow's older fixed stages (Staging, Production, Archived) are deprecated in favour of aliases and tags since MLflow 2.9, because aliases are flexible and several can point at different versions at once.

    What interviewers listen for
    • Named models, immutable numbered versions
    • Lineage back to the training run
    • Aliases decouple serving from version numbers
    • Stages are deprecated in favour of aliases

    Likely follow-up: How does a serving process find out the champion alias has moved? · Would you use one registry per environment or one shared registry?

  6. 6.What is the difference between data drift and concept drift?easy

    Data drift (covariate shift) is a change in the distribution of inputs, P(X). A loan model trained on applicants aged 30 to 50 starts receiving mostly students. The relationship between features and outcome may still hold, but the model is now working in regions it saw little of.

    Concept drift is a change in the relationship between inputs and target, P(y given X). The same transaction pattern that was legitimate last year is now typical of a new fraud scheme. Inputs can look identical while the right answer has changed.

    The practical difference is detection. Data drift can be measured immediately by comparing feature distributions with a reference window (PSI, KS test). Concept drift usually needs labels, which often arrive late, so you watch proxies such as prediction drift and business metrics until ground truth arrives.

    Not every data drift hurts accuracy, so drift alerts should trigger investigation, not automatic retraining.

    What interviewers listen for
    • Data drift: P(X) changes
    • Concept drift: P(y given X) changes
    • Data drift detectable without labels
    • Drift is a signal to investigate, not proof of harm

    Likely follow-up: Give an example of data drift that does not hurt the model. · What is label shift and how does it differ from both?

  7. 7.Why do you need to version data, and how would you do it?easy

    Because a model is a function of code and data. If the training table is overwritten every night, you cannot rebuild last month's model, explain a regression or show an auditor what the model learned from.

    Options, from simplest:

    • Immutable snapshots: write each training set to a dated, never-modified path (s3://ml/churn/2026-10-01/) and log that path and a content hash with the run.
    • Table formats with time travel: Delta Lake, Apache Iceberg or Hudi let you query a table as of a version or timestamp.
    • DVC or lakeFS: git-like versioning for files and data lakes; DVC stores small pointer files in git and the data in remote storage.

    Whichever you use, the run must record the exact version it read. A dataset name such as customers_latest is not a version.

    What interviewers listen for
    • Model depends on code and data
    • Snapshots, time travel tables or DVC/lakeFS
    • Log the exact version with each run
    • Latest is not a version

    Likely follow-up: How do you version a 5 TB dataset cheaply? · How do data retention or deletion requests interact with data versioning?

  8. 8.How do you package a trained model so it can be deployed reliably?easy

    I package three things together: the model artifact, the exact inference code including preprocessing, and the environment.

    • One artifact for preprocessing and model: a scikit-learn Pipeline or an MLflow pyfunc model, so the serving side cannot apply different feature logic.
    • A signature: the expected input columns, types and output shape. MLflow infers it with infer_signature and validates requests against it.
    • Pinned dependencies: library versions recorded with the model (MLflow writes requirements.txt and conda.yaml), because a pickle loaded under a different scikit-learn version can fail or behave differently.
    • A container image: a small multi-stage Docker image with a serving framework such as FastAPI, MLflow serving, KServe or Triton, health and readiness endpoints, and the model either baked in or pulled by registry alias at start-up.

    Then the same image is promoted through environments, never rebuilt per environment.

    What interviewers listen for
    • Preprocessing travels with the model
    • Input signature and validation
    • Pinned library versions
    • Immutable container promoted across environments

    Likely follow-up: Would you bake the model into the image or download it at start-up? · Why is pickle risky as a model format?

  9. 9.A regulator asks you to reproduce the exact model you deployed eight months ago. What must you have recorded?mid

    Five things, all linked to the registered model version:

    • Code: the git commit of the training pipeline, with no uncommitted changes (fail the run if the tree is dirty).
    • Data: an immutable snapshot or table version plus a content hash, including the label definition and the query that built it.
    • Configuration: every hyperparameter and feature flag, and the random seeds.
    • Environment: the container image digest or a lockfile, down to library versions, CUDA and driver versions for GPU training.
    • The artifact itself: the model file and its checksum, so you can also prove what was served without retraining.

    Even then, bit-for-bit reproduction may be impossible: GPU kernels can be non-deterministic and parallel reductions sum in different orders. So I would say what I can promise: the same artifact exactly, and a retrained model within a stated tolerance on the original evaluation set.

    What interviewers listen for
    • Code commit, data snapshot, config and seeds
    • Environment pinned by image digest or lockfile
    • Keep the original artifact and checksum
    • Admit GPU non-determinism; define a tolerance

    Likely follow-up: How would you make PyTorch training as deterministic as possible? · How long would you keep training snapshots?

  10. 10.What is training-serving skew? Give causes and how you prevent it.mid

    Training-serving skew is any difference between the features or behaviour a model saw in training and what it gets in production. The model looks good offline and underperforms live, with no error anywhere.

    Common causes:

    • Two implementations of the same feature: SQL in the training pipeline, Java in the service, with different null handling or time zones.
    • Leaky offline features: training used values computed after the prediction time, such as a 30-day count recalculated today.
    • Stateful preprocessing refitted at serving, for example fitting a scaler on a single request.
    • Different data sources or freshness: a feature that is complete in the warehouse but delayed in the online store.

    Prevention: one feature definition used by both paths (a feature store or shared library), preprocessing shipped inside the model artifact, point-in-time correct training sets, and logging the features actually served so you can compare them with training values.

    What interviewers listen for
    • Offline and online features differ silently
    • Duplicate feature code is the classic cause
    • Point-in-time correctness for training data
    • Log served features and compare

    Likely follow-up: How would you detect skew that already exists in production? · Is training-serving skew the same as drift?

  11. 11.What is a feature store? Explain the offline store, the online store and point-in-time joins.mid

    A feature store manages feature definitions once and serves them to both training and inference. Feast, Tecton, Databricks and Vertex AI all follow the same shape.

    • Offline store: the full history of feature values with timestamps, in a warehouse or lake. Used to build training sets.
    • Online store: only the latest value per entity, in a low-latency key-value store such as Redis or DynamoDB. Used at prediction time. A materialisation job copies fresh values from offline to online.
    • Point-in-time join: for each training label with timestamp t, fetch the feature value as it was at t, not the latest value. Otherwise training sees the future and offline metrics are inflated.

    The value is consistency and reuse: the same definition produces training and serving values, and teams share features instead of rebuilding them. The cost is another system to operate, so a team with one batch model may not need one.

    What interviewers listen for
    • One definition for training and serving
    • Offline: history; online: latest per entity
    • Point-in-time joins prevent leakage
    • Not free: justify the operational cost

    Likely follow-up: How would you build a point-in-time join without a feature store? · What happens if the materialisation job falls behind?

  12. 12.Why should model training run in an orchestrated pipeline instead of a notebook? What does the pipeline look like?mid

    A notebook depends on hidden state, execution order and one person's laptop. A pipeline makes training repeatable, schedulable, reviewable and observable.

    A typical training pipeline is a DAG of steps:

    • Extract data for a defined window.
    • Validate it: schema, null rates, ranges, row counts.
    • Build features and the train, validation and test splits.
    • Train, logging everything to the tracker.
    • Evaluate against the current champion, overall and on key slices.
    • Register the model if it passes the gate, otherwise stop and alert.

    Orchestrators such as Airflow, Kubeflow Pipelines, Vertex AI Pipelines, Dagster or Prefect handle scheduling, retries, dependencies and passing artifacts between steps. Each step should be a containerised, idempotent unit with explicit inputs and outputs, so a failed run can resume and any step can be rerun with the same inputs.

    What interviewers listen for
    • Notebooks hide state and are not schedulable
    • DAG: extract, validate, features, train, evaluate, register
    • Steps are containerised and idempotent
    • Orchestrator handles retries and lineage

    Likely follow-up: How do you test a pipeline step? · Airflow or Kubeflow Pipelines: what would decide it for you?

  13. 13.How do the Population Stability Index and the Kolmogorov-Smirnov test detect drift, and what are their pitfalls?mid

    PSI bins the reference distribution (usually into deciles), measures the share of current data in each bin and sums (cur - ref) * ln(cur / ref). Common rules of thumb: below 0.1 stable, 0.1 to 0.25 moderate shift, above 0.25 significant. It works for numeric and categorical features and is symmetric.

    The two-sample KS test compares empirical cumulative distributions; the statistic is the largest vertical gap between them, and scipy.stats.ks_2samp returns it with a p-value. It only applies to continuous numeric features.

    Pitfalls:

    • KS p-values depend on sample size. With a million rows, a shift of 0.01 standard deviations is highly significant and irrelevant. Alert on the statistic or an effect size, not p below 0.05.
    • PSI depends on binning, and an empty bin makes it infinite without a small epsilon.
    • Testing hundreds of features daily produces false alarms; prioritise important features and require persistence over several windows.
    import numpy as np
    
    def psi(reference, current, bins=10, eps=1e-6):
        edges = np.quantile(reference, np.linspace(0, 1, bins + 1))
        edges[0], edges[-1] = -np.inf, np.inf
        ref = np.histogram(reference, edges)[0] / len(reference)
        cur = np.histogram(current, edges)[0] / len(current)
        ref, cur = np.clip(ref, eps, None), np.clip(cur, eps, None)
        return float(np.sum((cur - ref) * np.log(cur / ref)))
    What interviewers listen for
    • PSI: binned share differences, 0.1 and 0.25 thresholds
    • KS: max gap between CDFs, numeric only
    • Large samples make tiny shifts significant
    • Binning, empty bins and multiple testing

    Likely follow-up: How would you monitor drift for a high-cardinality categorical feature? · Which reference window would you compare against?

  14. 14.Labels for your loan default model arrive 12 months after the prediction. How do you monitor it in the meantime?mid

    I would monitor in layers, from fastest to slowest signal.

    • Service health: error rate, latency, missing or default-valued features, request volume. These catch most incidents.
    • Data quality and drift: null rates, out-of-range values and PSI per important feature against the training or recent-stable window.
    • Prediction drift: the score distribution and approval rate. A sudden jump in average predicted risk with no policy change is a strong warning.
    • Early proxies: labels that correlate with the final one but arrive sooner, such as a missed first payment at 30 or 60 days.
    • Ground truth: when the 12-month labels arrive, compute AUC and calibration by cohort, joined on the logged prediction id.

    This only works if every prediction is logged with its id, model version, input features and timestamp, so later labels can be joined back.

    What interviewers listen for
    • Health, data quality, drift first
    • Prediction distribution as a proxy
    • Early proxy labels
    • Log predictions to join labels later

    Likely follow-up: How would you detect calibration drift? · What would you alert on versus only show on a dashboard?

  15. 15.What is a shadow deployment for a model, and what can it tell you and what can it not?mid

    In a shadow deployment the new model receives a copy of live traffic and makes predictions, but its outputs are only logged; users still get the current model's answer.

    It is good for checking things that offline tests miss, with zero user risk:

    • Latency, memory and error rate under real traffic.
    • Training-serving skew: does it get the features you expected, with the same distributions?
    • Agreement with the champion: how often and where the decisions differ.
    • Once labels arrive, accuracy on real production traffic.

    What it cannot measure is the effect of acting on its predictions. If the new recommender would show different items, you never see whether users click them, because they were never shown. For that you need a canary or an A/B test.

    Implementation detail: mirror the request asynchronously so the shadow model can never add latency or failures to the live path, and watch the extra cost.

    What interviewers listen for
    • Receives mirrored traffic, outputs not served
    • Validates latency, errors, skew, agreement
    • Cannot measure effect on user behaviour
    • Mirror asynchronously, off the critical path

    Likely follow-up: How would you implement traffic mirroring on Kubernetes? · How do you compare outputs when labels are delayed?

  16. 16.How would you run a canary release for a new model version, and how do you roll back?mid

    Route a small share of real traffic, say 5%, to the new version, compare it with the current version over a fixed window, then step up (5, 25, 50, 100%) only if guardrails hold.

    Guardrails are defined before the rollout: error rate, p99 latency, prediction distribution, and fast business proxies such as approval rate or click-through. Ideally an automated analysis compares canary against baseline and aborts on breach (Argo Rollouts and Flagger do this on Kubernetes).

    For rollback, the old version must still be running or instantly loadable. With a registry, rollback is moving the champion alias back to the previous version, or shifting the traffic weight back to 0. Keep the previous model's feature pipeline and schema compatible, otherwise rollback breaks.

    Route by a stable hash of user id rather than per request, so one user gets consistent behaviour and metrics are not mixed within a session.

    What interviewers listen for
    • Small traffic share, staged increases
    • Guardrails agreed before rollout
    • Rollback = alias or traffic weight change
    • Sticky routing by hashed user id

    Likely follow-up: How long should each canary step last? · What makes a model rollback harder than a code rollback?

  17. 17.The new model has better offline AUC. Why would you still A/B test it, and how?mid

    Offline AUC measures ranking on historical data. The business cares about an outcome such as revenue, conversion or losses, and the two can disagree: the new model may be slower, may be better only on segments that do not matter, or may change user behaviour in ways the historical data cannot show.

    An A/B test randomly assigns users (by a stable hash of user id) to the champion or challenger and compares a primary metric chosen in advance, with guardrail metrics such as latency, complaints or fairness.

    The statistics matter: compute the sample size from the minimum effect you care about, run for the planned duration (often whole weeks to cover weekly cycles), and avoid stopping as soon as the p-value dips below 0.05, which inflates false positives. Check sample ratio mismatch: if the split was 50/50 but you see 52/48, the assignment is broken and the result is not trustworthy.

    What interviewers listen for
    • Offline metric is a proxy for the business metric
    • Randomise by stable user hash
    • Pre-registered primary and guardrail metrics
    • Sample size, no peeking, sample ratio mismatch

    Likely follow-up: When would you use a multi-armed bandit instead of an A/B test? · What if offline and online results disagree?

  18. 18.When should a model be retrained: on a schedule, on drift, or on performance drop?mid

    Each trigger has a place.

    • Scheduled (daily, weekly, monthly): simplest and predictable. Pick the cadence from how fast performance decays, which you can measure by evaluating old models on newer data.
    • Performance-based: retrain when measured quality falls below a threshold. Most direct, but only possible when labels arrive quickly.
    • Drift-based: retrain when input or prediction drift persists. Useful when labels are slow, but drift does not always mean worse predictions, so it can retrain for nothing.
    • Event-based: new data volume, a product launch or a known shift such as a pricing change.

    In practice I start with a schedule plus alerts, and whichever trigger fires, the retrained model still goes through the same validation gate against the champion. Automatic retraining without automatic validation just automates shipping regressions. Also note that retraining cannot fix a broken upstream feature; check data quality first.

    What interviewers listen for
    • Schedule, performance, drift, event triggers
    • Measure decay to choose cadence
    • Retrained model still passes the gate
    • Retraining does not fix broken data

    Likely follow-up: How much history should each retrain use? · How would you detect that retraining made things worse?

  19. 19.What does CI/CD look like for a machine learning system? What do you test?mid

    There are three loops. CI runs on every change to code or pipeline definitions. CD deploys the pipeline and the serving service. CT (continuous training) runs the pipeline to produce new models, which then go through their own promotion.

    Tests on each pull request:

    • Unit tests for feature functions, including nulls, time zones and edge cases.
    • Data contract tests: schema, ranges and null rates on a sample.
    • A smoke training run on a small fixture dataset, proving the pipeline runs end to end and the loss goes down.
    • Serving tests: the packaged model loads, accepts the signature and returns valid outputs within a latency budget.

    Model promotion tests, run after training:

    • Quality against the champion on the same holdout, overall and per slice.
    • Behavioural checks: invariance, directional expectations, fairness metrics.
    • Size and latency limits.

    Only a model that passes is registered and moves to shadow or canary.

    What interviewers listen for
    • CI, CD and continuous training
    • Unit, data contract and smoke-training tests
    • Champion comparison and slice checks
    • Latency and size gates before rollout

    Likely follow-up: How do you keep a smoke training run fast? · What is a behavioural test for a model?

  20. 20.Your model API has a p50 of 20 ms but a p99 of 900 ms. How do you investigate and fix it?mid

    First, measure where the time goes for slow requests: feature lookup, preprocessing, model execution or serialization. Averages hide this; the mean can look fine while the tail ruins checkout.

    Common causes and fixes:

    • Slow feature fetches: a database call per feature or a cold cache. Batch the lookups, precompute, set timeouts with sensible fallback values.
    • Garbage collection or cold starts: warm the model at start-up with a dummy request and only mark the pod ready afterwards.
    • Resource contention: CPU throttling from tight limits, noisy neighbours, too many worker threads. Check throttling metrics and right-size requests.
    • Large inputs: long texts or many candidates. Cap input size or split the work.
    • The model itself: distil, quantise, prune, or compile with ONNX Runtime or TensorRT.

    On GPU servers, dynamic batching raises throughput but adds queueing delay, so cap the batch wait time to protect the tail.

    What interviewers listen for
    • Break down latency per stage for slow requests
    • Feature lookups and cold starts are frequent culprits
    • Check CPU throttling
    • Model optimisation and capped dynamic batching

    Likely follow-up: How would you set a timeout and fallback for the model call itself? · How does dynamic batching trade latency for throughput?

  21. 21.Describe the MLOps maturity levels 0, 1 and 2 from Google Cloud's MLOps guide.mid
    • Level 0, manual process: data scientists train in notebooks and hand a model file to engineers, who deploy it. Releases are rare, there is no monitoring of model quality and no link between the deployed model and how it was made.
    • Level 1, ML pipeline automation: the training pipeline itself is automated and deployed, so the system can retrain continuously on new data with automated data and model validation. What gets deployed is the pipeline, not just the model. Adds a feature store, metadata tracking and triggers.
    • Level 2, CI/CD pipeline automation: changes to the pipeline code are themselves built, tested and deployed automatically, so teams can try new features, architectures and hyperparameters and ship pipeline changes quickly and safely.

    The point I would add in an interview: most teams do not need level 2 for every model. Choose based on how often the data changes and how often you change the modelling approach.

    What interviewers listen for
    • Level 0: manual, notebook handoff
    • Level 1: automated continuous training
    • Level 2: CI/CD for the pipeline itself
    • Match the level to the need

    Likely follow-up: Which component would you add first to move from level 0 to level 1? · What breaks when a level 0 team scales to ten models?

  22. 22.What checks would you put in an automated gate before a retrained model can replace the champion?mid

    Compare challenger and champion on the same, fresh holdout data that neither was trained on, ideally the most recent period.

    • Primary metric: must beat or match the champion by a margin larger than run-to-run noise, not by 0.001.
    • Slices: no meaningful regression on important segments (country, device, new users, high-value customers). An overall gain can hide a slice getting much worse.
    • Calibration, if scores are used as probabilities or thresholds.
    • Behavioural tests: known hard examples, invariance (changing a name does not change a credit score) and directional checks.
    • Operational: latency, memory and artifact size within budget; the signature matches what the service sends.
    • Fairness and policy checks where they apply.

    If everything passes, register and alias it as challenger and send it to shadow or canary; promotion to champion happens only after live checks.

    What interviewers listen for
    • Same fresh holdout for both models
    • Margin above noise, not any improvement
    • Slice, calibration and behavioural checks
    • Operational limits; live checks still follow

    Likely follow-up: How do you estimate run-to-run noise? · What if the challenger wins overall but loses badly on one slice?

  23. 23.An upstream team renamed a column and your model silently scored with nulls for two days. How do you prevent this?mid

    Treat data as an interface with a contract, and validate it at every boundary.

    • Schema checks at pipeline ingestion and at the serving API: required columns, types, allowed categories. A missing column should fail loudly, not become NaN filled with a default.
    • Distribution checks: null rate, range, cardinality and row counts against expected bounds. A feature going from 1% null to 100% null is the clearest alert you can have.
    • Data contracts with the producing team: the schema is versioned, breaking changes need notice, and their CI runs your contract tests.
    • Serving-side monitoring of the share of requests where each feature fell back to a default value.

    Tools such as Great Expectations, TensorFlow Data Validation, Pandera or dbt tests implement these checks. The organisational fix matters as much: a named owner for every upstream table the model reads.

    What interviewers listen for
    • Fail loudly on schema changes
    • Null rate and range checks with alerts
    • Versioned data contracts with producers
    • Monitor default-value fallbacks at serving

    Likely follow-up: Where would you put validation: training, serving or both? · How strict should checks be before they cause alert fatigue?

  24. 24.GPU spend for training and serving has tripled. How do you bring it down without hurting quality?hard

    First measure utilisation per workload; idle or underused GPUs are usually the biggest waste.

    Training:

    • Use spot or preemptible instances with frequent checkpointing so interruptions only lose minutes.
    • Use mixed precision (bf16 or fp16) for faster steps and lower memory.
    • Fix input pipelines that starve the GPU: data loading is a common bottleneck.
    • Run smaller hyperparameter searches with early stopping, and try on samples first.

    Serving:

    • Right-size the hardware: many models run fine on CPU or smaller GPUs.
    • Quantise or distil models (int8, smaller student models).
    • Batch requests dynamically and share GPUs between small models (Triton multi-model serving, NVIDIA MIG partitions).
    • Autoscale on queue depth or concurrency, scale to zero for rare workloads, and move non-urgent work to batch.

    Then make cost visible: tag spend per team and model, and add cost per 1,000 predictions to the model review.

    What interviewers listen for
    • Measure utilisation first
    • Spot instances with checkpointing, mixed precision
    • Quantise, distil, batch, share GPUs
    • Autoscale and attribute cost per model

    Likely follow-up: What breaks when training on spot instances? · When is CPU inference the right choice?

  25. 25.What is a degenerate feedback loop in a production ML system, and how do you mitigate it?hard

    It happens when a model's predictions influence the data it is later trained on. A recommender only shows items it already ranks highly, users can only click what they see, and the next model learns that those items are even better. Popular items get more popular; new items never get a chance. Fraud models are similar: blocked transactions never get a real outcome label, so the model never learns whether it was right.

    Mitigations:

    • Exploration: show a small share of randomised or less-certain items (epsilon-greedy, bandits) and log the propensity of each choice.
    • Holdout traffic: a small group not affected by the model, giving unbiased labels and a baseline.
    • Correct for selection bias: inverse propensity weighting when training on logged data.
    • Monitor diversity and coverage of predictions, not just accuracy.

    The key interview point is that offline evaluation on logged data inherits the loop, so it can look great while the system narrows.

    What interviewers listen for
    • Predictions shape future training data
    • Recommenders and fraud blocking examples
    • Exploration and holdout traffic
    • Propensity logging and weighting

    Likely follow-up: How do you evaluate a new recommender offline when the data came from the old one? · How big should a holdout group be?

  26. 26.A conversion model that was stable for a year suddenly loses 15% of its precision. How do you debug it?hard

    I would rule out the cheap causes first, in order.

    • Did anything change on our side? Model version, feature pipeline, serving code, config, dependency upgrades. Check deploy logs around the start of the drop.
    • Is the measurement right? Label pipeline delays or definition changes can fake a drop. Compare with an independent business metric.
    • Data quality: per-feature null rates, default-value rates, ranges and cardinality, compared with the week before. A broken upstream feature is the most common cause.
    • Data drift: PSI per feature and by segment. Did a new traffic source, country or app version appear?
    • Concept drift: if inputs look the same but precision fell, the relationship changed: a competitor promotion, a policy change, seasonality.
    • Slices: find where the drop is concentrated; a global drop is often one segment.

    Mitigation depends on the cause: roll back, fix the feature, retrain on recent data, or adjust the threshold as a stopgap while the real fix lands.

    What interviewers listen for
    • Check our own changes first
    • Verify the labels and metric
    • Data quality before drift
    • Slice to localise; fix depends on cause

    Likely follow-up: What would you have wanted in place before this happened? · When is adjusting the threshold an acceptable fix?

  27. 27.In an LLM application, why should prompts be versioned like models, and how would you manage them?hard

    In an LLM app the prompt is effectively the model's code: a one-word change can alter tone, refusals, output format or cost. If prompts live as strings scattered in application code, you cannot tell which prompt produced a bad answer or roll back quickly.

    What I would do:

    • Store prompts in a prompt registry (MLflow Prompt Registry, Langfuse and similar) with immutable versions and commit messages.
    • Load them by alias, for example prompts:/support-reply@production, so promotion and rollback do not need a code deploy.
    • Record the full configuration as one versioned unit: prompt version, model name and version, temperature, tools and retrieval settings. Changing the underlying model is a release too.
    • Trace every request with the prompt version, so incidents can be traced to a version.
    • Promote a new version only after it passes the evaluation suite.
    import mlflow
    
    p = mlflow.genai.register_prompt(
        name="support-reply",
        template="You are a support agent. Cite the policy. Question: {{question}}",
        commit_message="cite policy",
    )
    mlflow.genai.set_prompt_alias("support-reply", alias="production", version=p.version)
    prompt = mlflow.genai.load_prompt("prompts:/support-reply@production")
    print(prompt.format(question="Where is my refund?"))
    What interviewers listen for
    • Prompt changes are behaviour changes
    • Registry with immutable versions and aliases
    • Version the whole config, including model
    • Trace prompt version per request

    Likely follow-up: How would you A/B test two prompt versions? · What do you do when the model provider deprecates the model you use?

  28. 28.How would you build an evaluation gate in CI for changes to an LLM-powered feature?hard

    Build a versioned evaluation dataset of real, anonymised inputs: common cases, known failures, adversarial and policy-sensitive examples, each with an expected answer or grading criteria.

    On every change to a prompt, model, retrieval setting or tool, CI runs the candidate over the dataset and scores it with:

    • Deterministic checks: valid JSON, required fields, length, banned content, citations present.
    • Reference-based metrics where there is a right answer: exact match, retrieval recall.
    • LLM-as-judge scorers for helpfulness, correctness against a reference and groundedness, calibrated against human labels so you know how far to trust them.
    • Cost and latency per request.

    The gate compares with the current production version: fail if a key score drops by more than a set tolerance, or any safety check fails. Because outputs are non-deterministic, use temperature 0 where possible, repeat flaky cases and set tolerances from measured variance. Production traces feed new failure cases back into the dataset.

    What interviewers listen for
    • Versioned eval dataset from real cases
    • Deterministic, reference and judge scorers
    • Compare against production with tolerances
    • Handle non-determinism; grow the dataset from traces

    Likely follow-up: How do you know the LLM judge is reliable? · How large does the evaluation set need to be?

  29. 29.Design the ML system for real-time fraud scoring at card checkout, with a 50 ms budget for the model call.hard

    Features: static customer features from batch jobs, plus streaming aggregates (transactions in the last 10 minutes, distinct merchants today) computed with Kafka and Flink and written to an online store such as Redis. The same definitions backfill the offline store with timestamps, so training uses point-in-time joins.

    Serving: a stateless scoring service, autoscaled, loading the model by registry alias. One batched feature lookup, a gradient-boosted model (fast on CPU), strict timeouts, and a rules-based fallback if features or the model are unavailable, because checkout must not fail.

    Logging: every request's features, score, model version and decision, keyed by transaction id.

    Labels: chargebacks arrive weeks later; join them back. Blocked transactions have no outcome, so keep a tiny randomised review sample to avoid a feedback loop.

    Lifecycle: weekly retraining with a champion gate, shadow then canary rollouts, drift and approval-rate monitoring, and alerting on fallback rate.

    What interviewers listen for
    • Streaming plus batch features, online and offline stores
    • Fast model, timeouts, rules fallback
    • Log features and decisions for label joins
    • Delayed labels and feedback loop handling

    Likely follow-up: How do you pick the decision threshold? · How would you handle a sudden new fraud pattern before retraining?

  30. 30.You serve a large model on GPUs in Kubernetes. Why is CPU-based autoscaling a poor fit, and what would you do instead?hard

    GPU inference servers barely use the CPU, so the CPU percentage says almost nothing about load. GPU utilisation is also misleading: a GPU can report 100% while serving comfortably, and with batching more traffic raises throughput before it hurts latency.

    Better scaling signals are queue depth, in-flight requests per replica (concurrency) or latency against the target. KServe with Knative scales on concurrency, and KEDA or a custom-metrics HPA can scale on Prometheus metrics from the server.

    Other things that matter:

    • Cold starts are slow: a new node, image pull and loading many GB of weights can take minutes. Keep a warm minimum, pre-pull images, cache weights on local disks and scale ahead of known peaks.
    • Request the GPU explicitly (nvidia.com/gpu: 1), since GPUs are not shared by default; use MIG or time-slicing for small models.
    • Readiness probes must wait until the model is loaded and warmed.
    • Batch work should go to a queue, not the latency-sensitive service.
    What interviewers listen for
    • CPU and GPU utilisation are poor signals
    • Scale on queue depth or concurrency
    • Cold starts: warm pools and weight caching
    • Explicit GPU requests and readiness after warm-up

    Likely follow-up: How does Knative scale-to-zero interact with large models? · How would you share one GPU between several small models?

Prefer multiple choice? All 20 MLOps MCQs with answers →

esc