MLOps MCQs multiple-choice questions with answers & explanations
All 20 MLOps quiz questions on one page. Pick an answer in your head, then open Show answer to check it and read why. Want a score and a timer? Take them as a quiz instead.
Official reference: Google Cloud: MLOps pipelines
- 1.easy
What does the last line print?
import mlflow with mlflow.start_run() as run: for step, loss in enumerate([0.9, 0.5, 0.3, 0.35]): mlflow.log_metric("loss", loss, step=step) print(mlflow.get_run(run.info.run_id).data.metrics["loss"])- A
0.3 - B
0.9 - C
[0.9, 0.5, 0.3, 0.35] - D
0.35
Show answer
Answer: D (
0.35)run.data.metricsholds the latest value logged for each metric key, here the value at the highest step,0.35. It is not the best value. The full series is available withMlflowClient().get_metric_history(run_id, "loss"). - A
- 2.mid
What happens on the second
log_paramcall?with mlflow.start_run(): mlflow.log_param("lr", 0.1) mlflow.log_param("lr", 0.2)- AThe value is overwritten with
0.2 - BBoth values are kept as a history
- CAn
MlflowException: changing param values is not allowed - DThe second call is silently ignored
Show answer
Answer: C (An
MlflowException: changing param values is not allowed)Parameters are immutable within a run, so MLflow raises
MlflowExceptionsaying that changing param values is not allowed. Logging the same value again is fine. Metrics, by contrast, are designed to be logged many times with a step. - AThe value is overwritten with
- 3.mid
Versions 1 and 2 of
fraudexist andchampionpoints at version 2. A new run then registers version 3. What does the last line print?client.set_registered_model_alias("fraud", "champion", 2) with mlflow.start_run(): mlflow.sklearn.log_model(model, name="model", registered_model_name="fraud") print(client.get_model_version_by_alias("fraud", "champion").version)- A
3 - B
2 - C
None - DIt raises because the alias is ambiguous
Show answer
Answer: B (
2)An alias is a pointer that only moves when you move it. Registering version 3 does not touch
champion, so serving code that loadsmodels:/fraud@championkeeps using version 2 until someone promotes version 3. That is what makes aliases safe deployment pointers. - A
- 4.mid
What happens when this runs?
mlflow.search_runs(experiment_names=["churn"], filter_string="loss < 0.5")- AReturns runs whose final loss is below 0.5
- BReturns every run, since the filter is ignored
- CRaises
MlflowExceptionabout an invalid attribute key - DReturns runs where any logged loss value was below 0.5
Show answer
Answer: C (Raises
MlflowExceptionabout an invalid attribute key)Filter strings need a prefix that says what the key is:
metrics.loss < 0.5,params.C = '1.0',tags.team = 'risk'. A bare name is treated as a run attribute, andlossis not one, so MLflow raises an invalid attribute key error. Metric filters compare the latest value. - 5.easy
A serving container loads
mlflow.pyfunc.load_model("models:/churn-model@champion"). What does this URI refer to?- AWhatever version the
championalias currently points to - BThe newest version of
churn-model - CThe run named
championin the churn experiment - DThe version in the deprecated Production stage
Show answer
Answer: A (Whatever version the
championalias currently points to)@championresolves an alias to one specific version at load time. Promotion and rollback become alias moves, with no code change. The newest version would be loaded withmodels:/churn-model/latest, which is risky because any registration changes it. - AWhatever version the
- 6.mid
A serving endpoint standardises each incoming request on its own before calling the model. What does it pass to the model for an amount of 120?
from sklearn.preprocessing import StandardScaler amount = [[120.0]] x = StandardScaler().fit_transform(amount) # inside the request handler print(x)- A
[[3.13]] - B
[[1.]] - C
[[0.]] - D
[[120.]]
Show answer
Answer: C (
[[0.]])Fitting a scaler on one row makes that row the mean, so every request becomes 0. This is a classic training-serving skew bug: the scaler must be fitted on training data and shipped with the model, for example inside a
Pipeline. Fitted on training amounts 20, 40, 60, 80 it would output about 3.13. - A
- 7.easy
Of 100 requests, 98 take 20 ms, one takes 900 ms and one takes 1,200 ms. Which summary is correct (NumPy default percentiles)?
import numpy as np lat = np.array([20] * 98 + [900, 1200], dtype=float) print(lat.mean(), np.percentile(lat, 50), np.percentile(lat, 99))- Amean 40.6, p50 20, p99 about 903
- Bmean 20, p50 20, p99 20
- Cmean 40.6, p50 40.6, p99 1200
- Dmean 1050, p50 20, p99 1200
Show answer
Answer: A (mean 40.6, p50 20, p99 about 903)
The mean (40.6 ms) looks healthy and the median is 20 ms, but the 99th percentile, interpolated between 900 and 1,200, is about 903 ms. Serving SLOs are written on tail percentiles because users notice the slow requests the average hides.
- 8.hard
Labels are at 1 and 5 March; the feature
orders_30dwas computed on 28 Feb (3), 3 Mar (4) and 6 Mar (9). What do the two joins return?pit = pd.merge_asof(labels.sort_values("ts"), feats.sort_values("ts"), on="ts", by="user") latest = labels.merge(feats.groupby("user", as_index=False).last(), on="user") print(pit["orders_30d"].tolist(), latest["orders_30d"].tolist())- A
[3, 4] [9, 9] - B
[3, 4] [3, 4] - C
[4, 9] [9, 9] - D
[9, 9] [3, 4]
Show answer
Answer: A (
[3, 4] [9, 9])merge_asoftakes, for each label, the last feature row at or before its timestamp, so the model sees what was known at prediction time: 3 and 4. Joining the latest value gives 9 to both labels, leaking a future value into training. Feature stores do the point-in-time version for you. - A
- 9.hard
Four API replicas assign users to the canary with
hash(user_id) % 100 < 5. What goes wrong?- AStrings hash differently in each process, so a user can switch between models
- BNothing: the same user always gets the same bucket
- CPython raises because strings are unhashable
- DAll users land in bucket 0
Show answer
Answer: A (Strings hash differently in each process, so a user can switch between models)
Python randomises
strhashing per process (unlessPYTHONHASHSEEDis fixed), so replicas and restarts disagree about a user's bucket and users flip between models, polluting the comparison. Use a stable hash such aszlib.crc32or SHA-256 of the id, salted per experiment. - 10.mid
Bin shares are reference
[0.5, 0.3, 0.2]and current[0.4, 0.4, 0.2]. What do the two calls print?import numpy as np psi = lambda ref, cur: np.sum((cur - ref) * np.log(cur / ref)) ref, cur = np.array([0.5, 0.3, 0.2]), np.array([0.4, 0.4, 0.2]) print(round(psi(ref, cur), 4), round(psi(cur, ref), 4))- A
0.0511 0.0511 - B
0.0511 -0.0511 - C
0.0511 0.0487 - D
0.1 0.1
Show answer
Answer: A (
0.0511 0.0511)Each term
(cur - ref) * ln(cur / ref)keeps its sign when both factors flip, so PSI is symmetric, unlike KL divergence. 0.05 is under the usual 0.1 threshold, so this would read as stable. - A
- 11.mid
A category had no rows in the reference window but 20% of rows today. Without any smoothing, what is the PSI?
ref = np.array([0.5, 0.5, 0.0]) cur = np.array([0.4, 0.4, 0.2]) print(np.sum((cur - ref) * np.log(cur / ref)))- A
0.2 - B
nan - C
inf - D
0.0
Show answer
Answer: C (
inf)The new bin contributes
0.2 * ln(0.2 / 0), which is infinite. Implementations clip shares to a small epsilon or add an "other" bucket so a brand new category produces a large but finite value. - A
- 12.hard
Two samples of 1,000,000 values differ in mean by 0.01 standard deviations. What does
ks_2sampreport?from scipy.stats import ks_2samp a = rng.normal(0, 1, 1_000_000) b = rng.normal(0.01, 1, 1_000_000) print(ks_2samp(a, b))- AA large statistic and a large p-value
- BA tiny statistic (around 0.004) and a p-value far below 0.05
- CA tiny statistic and a p-value near 1
- DAn error: the samples are too large
Show answer
Answer: B (A tiny statistic (around 0.004) and a p-value far below 0.05)
With huge samples, even a negligible shift is statistically significant: the statistic is around 0.004 while p is typically between 1e-5 and 1e-13, depending on the sample. Alerting on p below 0.05 would page someone every day; alert on the statistic, PSI or another effect size instead.
- 13.easy
A feature's PSI against the training data is 0.15. Using the common rule of thumb, how should it be read?
- ANo meaningful change
- BSevere shift: retrain immediately
- CModerate shift: investigate
- DPSI above 0.1 means the model is broken
Show answer
Answer: C (Moderate shift: investigate)
The usual bands are below 0.1 stable, 0.1 to 0.25 moderate, above 0.25 significant. Even a large PSI calls for investigation rather than automatic retraining, because drift on an unimportant feature may not hurt predictions.
- 14.easy
Transaction features look exactly like last quarter, but fraudsters now use a pattern the model learned was safe. What is this?
- AData drift
- BTraining-serving skew
- CConcept drift
- DLabel leakage
Show answer
Answer: C (Concept drift)
The inputs P(X) are unchanged but the relationship to the label P(y given X) has changed, which is concept drift. Input drift monitors will stay quiet, so you need label-based metrics or proxies such as chargeback rates.
- 15.mid
A new recommender runs in shadow mode for two weeks. Which question can shadow mode not answer?
- AIs its p99 latency within budget on real traffic?
- BDoes it receive the features it was trained on?
- CHow often does it disagree with the current model?
- DWould users click more on its recommendations?
Show answer
Answer: D (Would users click more on its recommendations?)
Shadow outputs are logged but never shown, so user reactions to them cannot be observed. Latency, feature health and disagreement are exactly what shadow mode is good for; behavioural impact needs a canary or A/B test.
- 16.hard
An A/B test was configured 50/50 but has 52,000 users in control and 48,000 in treatment. What should you conclude?
from scipy.stats import chisquare print(chisquare([52000, 48000]).pvalue) # 1.1e-36- AThe assignment or logging is broken; do not trust the results yet
- BFine: a 2% imbalance is normal noise
- CTreatment is worse, because users left it
- DRebalance by dropping 4,000 control users
Show answer
Answer: A (The assignment or logging is broken; do not trust the results yet)
This is a sample ratio mismatch: with 100,000 users such an imbalance has a p-value around 1e-36 under a true 50/50 split. Something is filtering users unevenly (crashes, redirects, bot filtering, logging), which biases every metric. Find the cause before reading the results.
- 17.mid
In Google Cloud's MLOps maturity model, what is the defining feature of level 1?
- AModels are trained manually in notebooks
- BPipeline code changes are built, tested and deployed by CI/CD
- CEvery model runs on GPUs
- DThe training pipeline is automated, enabling continuous training
Show answer
Answer: D (The training pipeline is automated, enabling continuous training)
Level 0 is a manual process, level 1 automates the ML pipeline so models retrain on new data with automated validation, and level 2 adds CI/CD for the pipeline code itself. Notebook training is level 0; automated pipeline deployment is level 2.
- 18.hard
You enable dynamic batching on a GPU model server and raise the maximum queue delay from 2 ms to 50 ms. What is the most likely effect?
- AHigher throughput and higher tail latency
- BLower throughput and lower latency
- CNo change: batching only affects training
- DLower GPU memory use and identical latency
Show answer
Answer: A (Higher throughput and higher tail latency)
Waiting longer lets the server build bigger batches, which uses the GPU more efficiently and raises throughput, but each request may wait up to 50 ms in the queue first, so p99 grows. The delay cap is the knob that trades one for the other.
- 19.mid
Two prompt versions are registered, then the alias is set. Which template does the application load?
mlflow.genai.register_prompt(name="support-reply", template="Answer politely: {{question}}") mlflow.genai.register_prompt(name="support-reply", template="Cite the policy: {{question}}") mlflow.genai.set_prompt_alias("support-reply", alias="production", version=1) p = mlflow.genai.load_prompt("prompts:/support-reply@production") print(p.version)- A
2, the newest version - BBoth, concatenated
- C
1, the version the alias points to - DIt raises: two versions exist
Show answer
Answer: C (
1, the version the alias points to)Prompt aliases work like model aliases:
@productionresolves to the version it was set to, version 1, regardless of newer registrations. Releasing version 2 or rolling back is an alias move, not a code deploy. - A
- 20.mid
A pipeline retrains nightly and automatically deploys the new model. An upstream feature started arriving as all nulls yesterday. What is the main risk?
- ANone: retraining adapts the model to the new data
- BThe pipeline fails to start because of the nulls
- COnly training time increases
- DThe model learns from broken data and the regression ships automatically
Show answer
Answer: D (The model learns from broken data and the regression ships automatically)
Retraining cannot fix broken inputs; it bakes them into the next model. Without data validation before training and a champion comparison before deployment, automation just ships the regression faster. Schema and null-rate checks should stop the run.
No questions match these filters. Try a different subtopic or clear the filters.