Ch. 31 · MLOps

LLMOps: Prompt Versioning and Evaluation Gates in CI

Treat prompts as versioned releases: register and alias them in MLflow, trace every call, and block regressions with an evaluation gate in CI.

~7 min readadvancedupdated Oct 6, 2026

“How do you manage prompts in production, and how do you know a change did not make the application worse?” is the LLMOps question that now appears in MLOps and ML platform interviews. Interviewers want to hear that you apply the same discipline to prompts and model choices that you would to code and trained models: versioning, review, automated evaluation before release, tracing in production, and fast rollback. Candidates who answer “we keep prompts in a config file and test them manually” usually get the follow-up “what happened the last time someone edited one?”.

Before you start

You should know what a prompt template is, roughly how an LLM-backed feature works (a template filled with user input, sent to a model, output parsed), and what a model registry alias is. The code uses MLflow 3.16 on Python 3.14 with a local SQLite tracking store. To keep the examples runnable without an API key, a small fake_llm function stands in for the model call; in a real pipeline it would call your provider. Outputs in comments are from real runs.

The short answer

Prompts change application behaviour as much as code does, so they belong in a prompt registry with immutable versions, commit messages and aliases such as production. The application loads prompts:/support-reply@production, so promotion and rollback are alias moves rather than deploys. The thing you release is the full configuration: prompt version, model name and version, temperature, tools and retrieval settings. Every candidate change runs through an evaluation gate in CI: a versioned dataset of real and adversarial inputs, scored with deterministic checks, reference metrics and calibrated LLM judges, compared against production with tolerances. Production traces record which version answered each request, and failures flow back into the evaluation set.

How it works

The MLflow Prompt Registry stores templates with {{variable}} placeholders. Each registration creates a new immutable version, and an alias points at one of them:

import mlflow

mlflow.set_tracking_uri("sqlite:///mlflow.db")

v1 = mlflow.genai.register_prompt(
    name="support-reply",
    template="Answer the customer politely: {{question}}",
    commit_message="initial")
v2 = mlflow.genai.register_prompt(
    name="support-reply",
    template="Reply as JSON with keys answer and citation. "
             "Cite the policy section. Question: {{question}}",
    commit_message="structured output with citations")
mlflow.genai.set_prompt_alias("support-reply", alias="production", version=v1.version)

prompt = mlflow.genai.load_prompt("prompts:/support-reply@production")
print(prompt.version, prompt.format(question="Where is my refund?"))
# 1 Answer the customer politely: Where is my refund?
python

Registering version 2 did not change what production loads; only moving the alias does. This is the same property that makes model registry aliases safe, and it means a bad prompt can be rolled back in seconds without touching application code.

Evaluation uses mlflow.genai.evaluate, which runs a predict_fn over a dataset and applies scorers. A scorer can be a plain Python function decorated with @scorer, or one of MLflow’s built-in LLM judges such as Correctness, RelevanceToQuery, RetrievalGroundedness, Safety and Guidelines. Deterministic scorers are cheap, exact and should come first:

import json
from mlflow.genai.scorers import scorer

@scorer
def valid_json(outputs) -> bool:
    try:
        return isinstance(json.loads(outputs), dict)
    except ValueError:
        return False

@scorer
def has_citation(outputs) -> bool:
    return "policy" in outputs.lower()

@scorer
def short_enough(outputs) -> bool:
    return len(outputs) <= 400
python

Step-by-step walkthrough

Step 1: Build a versioned evaluation dataset

Collect real, anonymised inputs: the most common requests, known past failures, edge cases (empty input, other languages, very long messages) and adversarial or policy-sensitive cases such as prompt injection attempts. Attach expected answers or grading guidelines where possible. Version the dataset like code, because a score is only comparable to another score on the same data. A few hundred well-chosen cases beat thousands of random ones.

Step 2: Run candidate and production through the same evaluation

def run_eval(prompt_uri: str) -> dict:
    prompt = mlflow.genai.load_prompt(prompt_uri)
    def predict_fn(question):
        return fake_llm(prompt.format(question=question))
    result = mlflow.genai.evaluate(data=eval_data, predict_fn=predict_fn,
                                   scorers=[valid_json, has_citation, short_enough])
    return dict(sorted((k.removesuffix("/mean"), round(float(v), 2))
                       for k, v in result.metrics.items()))

prod = run_eval("prompts:/support-reply@production")
cand = run_eval(f"prompts:/support-reply/{v2.version}")
print(prod)   # {'has_citation': 0.0, 'short_enough': 1.0, 'valid_json': 0.0}
print(cand)   # {'has_citation': 1.0, 'short_enough': 1.0, 'valid_json': 1.0}
python

Each evaluation is logged as an MLflow run with per-example results, so a reviewer can open the failing cases instead of trusting an average.

Step 3: Fail the build on regressions

regressions = [k for k in prod if cand[k] < prod[k] - 0.05]
print("gate:", "pass" if not regressions else f"fail {regressions}")   # gate: pass
python

Safety checks should be absolute (any failure blocks), quality scores relative to production with a tolerance set from measured variance. Outputs are non-deterministic, so use temperature 0 where the product allows, repeat borderline cases, and track cost and latency per request beside quality.

Step 4: Trace production and close the loop

Instrument the application so each request records the prompt version, model, parameters, retrieved documents, output, latency and token cost (@mlflow.trace or the autologging integrations). When a user reports a bad answer, the trace says exactly which configuration produced it. Bad answers, once triaged, become new cases in the evaluation dataset, so the gate gets stricter over time.

Worked scenario

A support assistant used prompt version 2 above with a mid-sized model. To cut cost, an engineer changed the model name in a config file to a cheaper model, reasoning that “the prompt did not change”. Within a day the downstream parser was failing on some replies: the new model sometimes prefixed its JSON with a friendly sentence. Simulating that behaviour and running the gate shows that the change would have been caught:

def cheaper_llm(prompt_text: str) -> str:
    body = json.dumps({"answer": "Refunds take 5 to 7 days.", "citation": "policy 4.2"})
    return body if "refund" in prompt_text.lower() else "Sure! Here is the JSON: " + body

fake_llm = cheaper_llm
new_model = run_eval(f"prompts:/support-reply/{v2.version}")
print(new_model)   # {'has_citation': 1.0, 'short_enough': 1.0, 'valid_json': 0.33}
regressions = [k for k in cand if new_model[k] < cand[k] - 0.05]
print("gate:", "pass" if not regressions else f"fail {regressions}")
# gate: fail ['valid_json']
python

The team changed two things. The model name, temperature and prompt version became one versioned configuration, so any change to any part of it triggered the evaluation workflow. And the parser got a defensive fallback, extracting the first JSON object from the reply and logging a format-violation metric, because no gate catches everything. With the cheaper model they then iterated on the prompt (an explicit “output only JSON” instruction and a one-shot example) until the gate passed, and kept the savings.

Common mistake

  • Hard-coding prompts in application code. Every change needs a deploy, and nobody can tell which version produced an answer.
  • Versioning the prompt but not the model. Providers update and retire models; a model change is a release.
  • Relying only on LLM-as-judge scores. Judges are biased (towards longer answers, towards their own style) and must be checked against human labels before you trust them.
  • Evaluating on a handful of happy-path examples. The gate is only as good as its dataset; include past failures and adversarial cases.
  • Treating one run as the truth. Non-determinism means small score differences can be noise; set tolerances from repeated runs.

Verify the behavior

Three checks show the setup works. Register a new prompt version and confirm load_prompt("prompts:/support-reply@production").version is unchanged until the alias moves; it stayed at 1 in our run. Feed the gate a deliberately broken candidate (the cheaper model above) and confirm the CI job fails; ours reported fail ['valid_json']. Finally, move the alias back to the previous version and confirm the application serves the old template without a redeploy.

Follow-up questions

  • How do you A/B test two prompt versions? Assign users by a stable hash, serve each group a different alias or version, and compare task success, user feedback and cost on live traffic after both have passed the offline gate.
  • How do you make an LLM judge trustworthy? Have humans label a sample, measure agreement with the judge, refine the judge’s instructions, and re-check whenever the judge model changes.
  • What would you do when a provider deprecates your model? Treat the replacement like any candidate: run the gate, adjust prompts, canary it, and keep the configuration versioned so you can compare.
  • How do you evaluate retrieval-augmented generation? Score retrieval (did the right documents come back) separately from generation (is the answer grounded in them), so failures point at the right component.

Interview exercise

Your team ships an LLM feature that summarises customer calls for agents. Prompts are edited directly in the codebase, there is no evaluation set, and last month a prompt tweak caused summaries to omit refund amounts for a week before anyone noticed. Design the release process you would introduce.

Answer and reasoning

Move prompts into a registry and load them by alias, so each summary can be traced to a prompt version and rollback is an alias move. Make the release unit the whole configuration (prompt version, model, temperature). Build an evaluation set from real, anonymised calls, deliberately including calls with refunds, cancellations and complaints, with reference facts each summary must contain. The gate checks deterministic facts first (does the summary contain the refund amount when the call mentions one), then judge-based scores for faithfulness and brevity validated against agent ratings, and blocks on any drop beyond tolerance. In production, trace every summary with its configuration, sample a few for human review each day, and track a “missing amount” detector as a live metric. The refund omission would have failed the deterministic check before release.

Continue learning

More in MLOps

esc