pencils ready ✎

MLOps MCQs multiple-choice questions with answers & explanations

All 20 MLOps quiz questions on one page. Pick an answer in your head, then open Show answer to check it and read why. Want a score and a timer? Take them as a quiz instead.

20 questions
  1. 1.

    What does the last line print?

    easy
    import mlflow
    
    with mlflow.start_run() as run:
        for step, loss in enumerate([0.9, 0.5, 0.3, 0.35]):
            mlflow.log_metric("loss", loss, step=step)
    
    print(mlflow.get_run(run.info.run_id).data.metrics["loss"])
    1. A0.3
    2. B0.9
    3. C[0.9, 0.5, 0.3, 0.35]
    4. D0.35
    Show answer

    Answer: D (0.35)

    run.data.metrics holds the latest value logged for each metric key, here the value at the highest step, 0.35. It is not the best value. The full series is available with MlflowClient().get_metric_history(run_id, "loss").

  2. 2.

    What happens on the second log_param call?

    mid
    with mlflow.start_run():
        mlflow.log_param("lr", 0.1)
        mlflow.log_param("lr", 0.2)
    1. AThe value is overwritten with 0.2
    2. BBoth values are kept as a history
    3. CAn MlflowException: changing param values is not allowed
    4. DThe second call is silently ignored
    Show answer

    Answer: C (An MlflowException: changing param values is not allowed)

    Parameters are immutable within a run, so MLflow raises MlflowException saying that changing param values is not allowed. Logging the same value again is fine. Metrics, by contrast, are designed to be logged many times with a step.

  3. 3.

    Versions 1 and 2 of fraud exist and champion points at version 2. A new run then registers version 3. What does the last line print?

    mid
    client.set_registered_model_alias("fraud", "champion", 2)
    
    with mlflow.start_run():
        mlflow.sklearn.log_model(model, name="model", registered_model_name="fraud")
    
    print(client.get_model_version_by_alias("fraud", "champion").version)
    1. A3
    2. B2
    3. CNone
    4. DIt raises because the alias is ambiguous
    Show answer

    Answer: B (2)

    An alias is a pointer that only moves when you move it. Registering version 3 does not touch champion, so serving code that loads models:/fraud@champion keeps using version 2 until someone promotes version 3. That is what makes aliases safe deployment pointers.

  4. 4.

    What happens when this runs?

    mid
    mlflow.search_runs(experiment_names=["churn"], filter_string="loss < 0.5")
    1. AReturns runs whose final loss is below 0.5
    2. BReturns every run, since the filter is ignored
    3. CRaises MlflowException about an invalid attribute key
    4. DReturns runs where any logged loss value was below 0.5
    Show answer

    Answer: C (Raises MlflowException about an invalid attribute key)

    Filter strings need a prefix that says what the key is: metrics.loss < 0.5, params.C = '1.0', tags.team = 'risk'. A bare name is treated as a run attribute, and loss is not one, so MLflow raises an invalid attribute key error. Metric filters compare the latest value.

  5. 5.

    A serving container loads mlflow.pyfunc.load_model("models:/churn-model@champion"). What does this URI refer to?

    easy
    1. AWhatever version the champion alias currently points to
    2. BThe newest version of churn-model
    3. CThe run named champion in the churn experiment
    4. DThe version in the deprecated Production stage
    Show answer

    Answer: A (Whatever version the champion alias currently points to)

    @champion resolves an alias to one specific version at load time. Promotion and rollback become alias moves, with no code change. The newest version would be loaded with models:/churn-model/latest, which is risky because any registration changes it.

  6. 6.

    A serving endpoint standardises each incoming request on its own before calling the model. What does it pass to the model for an amount of 120?

    mid
    from sklearn.preprocessing import StandardScaler
    
    amount = [[120.0]]
    x = StandardScaler().fit_transform(amount)   # inside the request handler
    print(x)
    1. A[[3.13]]
    2. B[[1.]]
    3. C[[0.]]
    4. D[[120.]]
    Show answer

    Answer: C ([[0.]])

    Fitting a scaler on one row makes that row the mean, so every request becomes 0. This is a classic training-serving skew bug: the scaler must be fitted on training data and shipped with the model, for example inside a Pipeline. Fitted on training amounts 20, 40, 60, 80 it would output about 3.13.

  7. 7.

    Of 100 requests, 98 take 20 ms, one takes 900 ms and one takes 1,200 ms. Which summary is correct (NumPy default percentiles)?

    easy
    import numpy as np
    lat = np.array([20] * 98 + [900, 1200], dtype=float)
    print(lat.mean(), np.percentile(lat, 50), np.percentile(lat, 99))
    1. Amean 40.6, p50 20, p99 about 903
    2. Bmean 20, p50 20, p99 20
    3. Cmean 40.6, p50 40.6, p99 1200
    4. Dmean 1050, p50 20, p99 1200
    Show answer

    Answer: A (mean 40.6, p50 20, p99 about 903)

    The mean (40.6 ms) looks healthy and the median is 20 ms, but the 99th percentile, interpolated between 900 and 1,200, is about 903 ms. Serving SLOs are written on tail percentiles because users notice the slow requests the average hides.

  8. 8.

    Labels are at 1 and 5 March; the feature orders_30d was computed on 28 Feb (3), 3 Mar (4) and 6 Mar (9). What do the two joins return?

    hard
    pit = pd.merge_asof(labels.sort_values("ts"), feats.sort_values("ts"),
                        on="ts", by="user")
    latest = labels.merge(feats.groupby("user", as_index=False).last(), on="user")
    print(pit["orders_30d"].tolist(), latest["orders_30d"].tolist())
    1. A[3, 4] [9, 9]
    2. B[3, 4] [3, 4]
    3. C[4, 9] [9, 9]
    4. D[9, 9] [3, 4]
    Show answer

    Answer: A ([3, 4] [9, 9])

    merge_asof takes, for each label, the last feature row at or before its timestamp, so the model sees what was known at prediction time: 3 and 4. Joining the latest value gives 9 to both labels, leaking a future value into training. Feature stores do the point-in-time version for you.

  9. 9.

    Four API replicas assign users to the canary with hash(user_id) % 100 < 5. What goes wrong?

    hard
    1. AStrings hash differently in each process, so a user can switch between models
    2. BNothing: the same user always gets the same bucket
    3. CPython raises because strings are unhashable
    4. DAll users land in bucket 0
    Show answer

    Answer: A (Strings hash differently in each process, so a user can switch between models)

    Python randomises str hashing per process (unless PYTHONHASHSEED is fixed), so replicas and restarts disagree about a user's bucket and users flip between models, polluting the comparison. Use a stable hash such as zlib.crc32 or SHA-256 of the id, salted per experiment.

  10. 10.

    Bin shares are reference [0.5, 0.3, 0.2] and current [0.4, 0.4, 0.2]. What do the two calls print?

    mid
    import numpy as np
    psi = lambda ref, cur: np.sum((cur - ref) * np.log(cur / ref))
    ref, cur = np.array([0.5, 0.3, 0.2]), np.array([0.4, 0.4, 0.2])
    print(round(psi(ref, cur), 4), round(psi(cur, ref), 4))
    1. A0.0511 0.0511
    2. B0.0511 -0.0511
    3. C0.0511 0.0487
    4. D0.1 0.1
    Show answer

    Answer: A (0.0511 0.0511)

    Each term (cur - ref) * ln(cur / ref) keeps its sign when both factors flip, so PSI is symmetric, unlike KL divergence. 0.05 is under the usual 0.1 threshold, so this would read as stable.

  11. 11.

    A category had no rows in the reference window but 20% of rows today. Without any smoothing, what is the PSI?

    mid
    ref = np.array([0.5, 0.5, 0.0])
    cur = np.array([0.4, 0.4, 0.2])
    print(np.sum((cur - ref) * np.log(cur / ref)))
    1. A0.2
    2. Bnan
    3. Cinf
    4. D0.0
    Show answer

    Answer: C (inf)

    The new bin contributes 0.2 * ln(0.2 / 0), which is infinite. Implementations clip shares to a small epsilon or add an "other" bucket so a brand new category produces a large but finite value.

  12. 12.

    Two samples of 1,000,000 values differ in mean by 0.01 standard deviations. What does ks_2samp report?

    hard
    from scipy.stats import ks_2samp
    a = rng.normal(0, 1, 1_000_000)
    b = rng.normal(0.01, 1, 1_000_000)
    print(ks_2samp(a, b))
    1. AA large statistic and a large p-value
    2. BA tiny statistic (around 0.004) and a p-value far below 0.05
    3. CA tiny statistic and a p-value near 1
    4. DAn error: the samples are too large
    Show answer

    Answer: B (A tiny statistic (around 0.004) and a p-value far below 0.05)

    With huge samples, even a negligible shift is statistically significant: the statistic is around 0.004 while p is typically between 1e-5 and 1e-13, depending on the sample. Alerting on p below 0.05 would page someone every day; alert on the statistic, PSI or another effect size instead.

  13. 13.

    A feature's PSI against the training data is 0.15. Using the common rule of thumb, how should it be read?

    easy
    1. ANo meaningful change
    2. BSevere shift: retrain immediately
    3. CModerate shift: investigate
    4. DPSI above 0.1 means the model is broken
    Show answer

    Answer: C (Moderate shift: investigate)

    The usual bands are below 0.1 stable, 0.1 to 0.25 moderate, above 0.25 significant. Even a large PSI calls for investigation rather than automatic retraining, because drift on an unimportant feature may not hurt predictions.

  14. 14.

    Transaction features look exactly like last quarter, but fraudsters now use a pattern the model learned was safe. What is this?

    easy
    1. AData drift
    2. BTraining-serving skew
    3. CConcept drift
    4. DLabel leakage
    Show answer

    Answer: C (Concept drift)

    The inputs P(X) are unchanged but the relationship to the label P(y given X) has changed, which is concept drift. Input drift monitors will stay quiet, so you need label-based metrics or proxies such as chargeback rates.

  15. 15.

    A new recommender runs in shadow mode for two weeks. Which question can shadow mode not answer?

    mid
    1. AIs its p99 latency within budget on real traffic?
    2. BDoes it receive the features it was trained on?
    3. CHow often does it disagree with the current model?
    4. DWould users click more on its recommendations?
    Show answer

    Answer: D (Would users click more on its recommendations?)

    Shadow outputs are logged but never shown, so user reactions to them cannot be observed. Latency, feature health and disagreement are exactly what shadow mode is good for; behavioural impact needs a canary or A/B test.

  16. 16.

    An A/B test was configured 50/50 but has 52,000 users in control and 48,000 in treatment. What should you conclude?

    hard
    from scipy.stats import chisquare
    print(chisquare([52000, 48000]).pvalue)   # 1.1e-36
    1. AThe assignment or logging is broken; do not trust the results yet
    2. BFine: a 2% imbalance is normal noise
    3. CTreatment is worse, because users left it
    4. DRebalance by dropping 4,000 control users
    Show answer

    Answer: A (The assignment or logging is broken; do not trust the results yet)

    This is a sample ratio mismatch: with 100,000 users such an imbalance has a p-value around 1e-36 under a true 50/50 split. Something is filtering users unevenly (crashes, redirects, bot filtering, logging), which biases every metric. Find the cause before reading the results.

  17. 17.

    In Google Cloud's MLOps maturity model, what is the defining feature of level 1?

    mid
    1. AModels are trained manually in notebooks
    2. BPipeline code changes are built, tested and deployed by CI/CD
    3. CEvery model runs on GPUs
    4. DThe training pipeline is automated, enabling continuous training
    Show answer

    Answer: D (The training pipeline is automated, enabling continuous training)

    Level 0 is a manual process, level 1 automates the ML pipeline so models retrain on new data with automated validation, and level 2 adds CI/CD for the pipeline code itself. Notebook training is level 0; automated pipeline deployment is level 2.

  18. 18.

    You enable dynamic batching on a GPU model server and raise the maximum queue delay from 2 ms to 50 ms. What is the most likely effect?

    hard
    1. AHigher throughput and higher tail latency
    2. BLower throughput and lower latency
    3. CNo change: batching only affects training
    4. DLower GPU memory use and identical latency
    Show answer

    Answer: A (Higher throughput and higher tail latency)

    Waiting longer lets the server build bigger batches, which uses the GPU more efficiently and raises throughput, but each request may wait up to 50 ms in the queue first, so p99 grows. The delay cap is the knob that trades one for the other.

  19. 19.

    Two prompt versions are registered, then the alias is set. Which template does the application load?

    mid
    mlflow.genai.register_prompt(name="support-reply", template="Answer politely: {{question}}")
    mlflow.genai.register_prompt(name="support-reply", template="Cite the policy: {{question}}")
    mlflow.genai.set_prompt_alias("support-reply", alias="production", version=1)
    
    p = mlflow.genai.load_prompt("prompts:/support-reply@production")
    print(p.version)
    1. A2, the newest version
    2. BBoth, concatenated
    3. C1, the version the alias points to
    4. DIt raises: two versions exist
    Show answer

    Answer: C (1, the version the alias points to)

    Prompt aliases work like model aliases: @production resolves to the version it was set to, version 1, regardless of newer registrations. Releasing version 2 or rolling back is an alias move, not a code deploy.

  20. 20.

    A pipeline retrains nightly and automatically deploys the new model. An upstream feature started arriving as all nulls yesterday. What is the main risk?

    mid
    1. ANone: retraining adapts the model to the new data
    2. BThe pipeline fails to start because of the nulls
    3. COnly training time increases
    4. DThe model learns from broken data and the regression ships automatically
    Show answer

    Answer: D (The model learns from broken data and the regression ships automatically)

    Retraining cannot fix broken inputs; it bakes them into the next model. Without data validation before training and a champion comparison before deployment, automation just ships the regression faster. Schema and null-rate checks should stop the run.

esc