“How do you manage prompts in production, and how do you know a change did not make the application worse?” is the LLMOps question that now appears in MLOps and ML platform interviews. Interviewers want to hear that you apply the same discipline to prompts and model choices that you would to code and trained models: versioning, review, automated evaluation before release, tracing in production, and fast rollback. Candidates who answer “we keep prompts in a config file and test them manually” usually get the follow-up “what happened the last time someone edited one?”.
Before you start
You should know what a prompt template is, roughly how an LLM-backed feature works (a template filled with user input, sent to a model, output parsed), and what a model registry alias is. The code uses MLflow 3.16 on Python 3.14 with a local SQLite tracking store. To keep the examples runnable without an API key, a small fake_llm function stands in for the model call; in a real pipeline it would call your provider. Outputs in comments are from real runs.
The short answer
Prompts change application behaviour as much as code does, so they belong in a prompt registry with immutable versions, commit messages and aliases such as production. The application loads prompts:/support-reply@production, so promotion and rollback are alias moves rather than deploys. The thing you release is the full configuration: prompt version, model name and version, temperature, tools and retrieval settings. Every candidate change runs through an evaluation gate in CI: a versioned dataset of real and adversarial inputs, scored with deterministic checks, reference metrics and calibrated LLM judges, compared against production with tolerances. Production traces record which version answered each request, and failures flow back into the evaluation set.
How it works
The MLflow Prompt Registry stores templates with {{variable}} placeholders. Each registration creates a new immutable version, and an alias points at one of them:
import mlflow
mlflow.set_tracking_uri("sqlite:///mlflow.db")
v1 = mlflow.genai.register_prompt(
name="support-reply",
template="Answer the customer politely: {{question}}",
commit_message="initial")
v2 = mlflow.genai.register_prompt(
name="support-reply",
template="Reply as JSON with keys answer and citation. "
"Cite the policy section. Question: {{question}}",
commit_message="structured output with citations")
mlflow.genai.set_prompt_alias("support-reply", alias="production", version=v1.version)
prompt = mlflow.genai.load_prompt("prompts:/support-reply@production")
print(prompt.version, prompt.format(question="Where is my refund?"))
# 1 Answer the customer politely: Where is my refund?Registering version 2 did not change what production loads; only moving the alias does. This is the same property that makes model registry aliases safe, and it means a bad prompt can be rolled back in seconds without touching application code.
Evaluation uses mlflow.genai.evaluate, which runs a predict_fn over a dataset and applies scorers. A scorer can be a plain Python function decorated with @scorer, or one of MLflow’s built-in LLM judges such as Correctness, RelevanceToQuery, RetrievalGroundedness, Safety and Guidelines. Deterministic scorers are cheap, exact and should come first:
import json
from mlflow.genai.scorers import scorer
@scorer
def valid_json(outputs) -> bool:
try:
return isinstance(json.loads(outputs), dict)
except ValueError:
return False
@scorer
def has_citation(outputs) -> bool:
return "policy" in outputs.lower()
@scorer
def short_enough(outputs) -> bool:
return len(outputs) <= 400Step-by-step walkthrough
Step 1: Build a versioned evaluation dataset
Collect real, anonymised inputs: the most common requests, known past failures, edge cases (empty input, other languages, very long messages) and adversarial or policy-sensitive cases such as prompt injection attempts. Attach expected answers or grading guidelines where possible. Version the dataset like code, because a score is only comparable to another score on the same data. A few hundred well-chosen cases beat thousands of random ones.
Step 2: Run candidate and production through the same evaluation
def run_eval(prompt_uri: str) -> dict:
prompt = mlflow.genai.load_prompt(prompt_uri)
def predict_fn(question):
return fake_llm(prompt.format(question=question))
result = mlflow.genai.evaluate(data=eval_data, predict_fn=predict_fn,
scorers=[valid_json, has_citation, short_enough])
return dict(sorted((k.removesuffix("/mean"), round(float(v), 2))
for k, v in result.metrics.items()))
prod = run_eval("prompts:/support-reply@production")
cand = run_eval(f"prompts:/support-reply/{v2.version}")
print(prod) # {'has_citation': 0.0, 'short_enough': 1.0, 'valid_json': 0.0}
print(cand) # {'has_citation': 1.0, 'short_enough': 1.0, 'valid_json': 1.0}Each evaluation is logged as an MLflow run with per-example results, so a reviewer can open the failing cases instead of trusting an average.
Step 3: Fail the build on regressions
regressions = [k for k in prod if cand[k] < prod[k] - 0.05]
print("gate:", "pass" if not regressions else f"fail {regressions}") # gate: passSafety checks should be absolute (any failure blocks), quality scores relative to production with a tolerance set from measured variance. Outputs are non-deterministic, so use temperature 0 where the product allows, repeat borderline cases, and track cost and latency per request beside quality.
Step 4: Trace production and close the loop
Instrument the application so each request records the prompt version, model, parameters, retrieved documents, output, latency and token cost (@mlflow.trace or the autologging integrations). When a user reports a bad answer, the trace says exactly which configuration produced it. Bad answers, once triaged, become new cases in the evaluation dataset, so the gate gets stricter over time.
Worked scenario
A support assistant used prompt version 2 above with a mid-sized model. To cut cost, an engineer changed the model name in a config file to a cheaper model, reasoning that “the prompt did not change”. Within a day the downstream parser was failing on some replies: the new model sometimes prefixed its JSON with a friendly sentence. Simulating that behaviour and running the gate shows that the change would have been caught:
def cheaper_llm(prompt_text: str) -> str:
body = json.dumps({"answer": "Refunds take 5 to 7 days.", "citation": "policy 4.2"})
return body if "refund" in prompt_text.lower() else "Sure! Here is the JSON: " + body
fake_llm = cheaper_llm
new_model = run_eval(f"prompts:/support-reply/{v2.version}")
print(new_model) # {'has_citation': 1.0, 'short_enough': 1.0, 'valid_json': 0.33}
regressions = [k for k in cand if new_model[k] < cand[k] - 0.05]
print("gate:", "pass" if not regressions else f"fail {regressions}")
# gate: fail ['valid_json']The team changed two things. The model name, temperature and prompt version became one versioned configuration, so any change to any part of it triggered the evaluation workflow. And the parser got a defensive fallback, extracting the first JSON object from the reply and logging a format-violation metric, because no gate catches everything. With the cheaper model they then iterated on the prompt (an explicit “output only JSON” instruction and a one-shot example) until the gate passed, and kept the savings.
Common mistake
- Hard-coding prompts in application code. Every change needs a deploy, and nobody can tell which version produced an answer.
- Versioning the prompt but not the model. Providers update and retire models; a model change is a release.
- Relying only on LLM-as-judge scores. Judges are biased (towards longer answers, towards their own style) and must be checked against human labels before you trust them.
- Evaluating on a handful of happy-path examples. The gate is only as good as its dataset; include past failures and adversarial cases.
- Treating one run as the truth. Non-determinism means small score differences can be noise; set tolerances from repeated runs.
Verify the behavior
Three checks show the setup works. Register a new prompt version and confirm load_prompt("prompts:/support-reply@production").version is unchanged until the alias moves; it stayed at 1 in our run. Feed the gate a deliberately broken candidate (the cheaper model above) and confirm the CI job fails; ours reported fail ['valid_json']. Finally, move the alias back to the previous version and confirm the application serves the old template without a redeploy.
Follow-up questions
- How do you A/B test two prompt versions? Assign users by a stable hash, serve each group a different alias or version, and compare task success, user feedback and cost on live traffic after both have passed the offline gate.
- How do you make an LLM judge trustworthy? Have humans label a sample, measure agreement with the judge, refine the judge’s instructions, and re-check whenever the judge model changes.
- What would you do when a provider deprecates your model? Treat the replacement like any candidate: run the gate, adjust prompts, canary it, and keep the configuration versioned so you can compare.
- How do you evaluate retrieval-augmented generation? Score retrieval (did the right documents come back) separately from generation (is the answer grounded in them), so failures point at the right component.
Interview exercise
Your team ships an LLM feature that summarises customer calls for agents. Prompts are edited directly in the codebase, there is no evaluation set, and last month a prompt tweak caused summaries to omit refund amounts for a week before anyone noticed. Design the release process you would introduce.
Answer and reasoning
Move prompts into a registry and load them by alias, so each summary can be traced to a prompt version and rollback is an alias move. Make the release unit the whole configuration (prompt version, model, temperature). Build an evaluation set from real, anonymised calls, deliberately including calls with refunds, cancellations and complaints, with reference facts each summary must contain. The gate checks deterministic facts first (does the summary contain the refund amount when the call mentions one), then judge-based scores for faithfulness and brevity validated against agent ratings, and blocks on any drop beyond tolerance. In production, trace every summary with its configuration, sample a few for human review each day, and track a “missing amount” detector as a live metric. The refund omission would have failed the deterministic check before release.
Continue learning
- Practise in the MLOps chapter and the MLOps MCQs.
- Go deeper on judges, datasets and metrics in evaluating LLM apps, and see the same alias pattern for trained models in MLflow tracking and model registry.
- Reference: the MLflow Prompt Registry guide and MLflow evaluation and monitoring for GenAI.