“How do you know your LLM feature works, and how do you know a prompt change did not break it?” Interviewers ask this because it is the question that separates demos from products. Classic unit tests assume one correct output; LLM outputs vary in wording, can be partially right, and drift when the model, prompt or retrieval changes. A strong candidate describes evaluation as an engineering system: a curated dataset, automated graders, a judge that has itself been checked, metrics for each pipeline stage, and a gate that blocks regressions before release.
Before you start
You should know basic Python (dataclasses, regular expressions) and the idea of a regression test. Familiarity with retrieval-augmented generation helps, because RAG systems need both retrieval and answer metrics. Examples run on Python 3.12+ with the standard library. Model outputs are recorded strings, so nothing calls a paid API; in a real suite, the runner would call your application exactly as production does.
The short answer
Build a golden dataset of realistic inputs, including edge cases and questions the system must refuse, each with what a correct answer must contain or avoid. Run the whole application over it and grade with code-based checks first (required facts, valid JSON, citations that exist, correct abstention), adding a model-based judge with a written rubric only for qualities code cannot check, such as tone or completeness. Calibrate the judge against human labels. For RAG, measure retrieval (recall@k, MRR) separately from answers (faithfulness, correctness). Run the suite in CI on every prompt, model or retrieval change, compare with the baseline per case, and treat small differences on small datasets as noise.
How it works
An eval case states the input and the observable properties of a good output. Graders return pass or fail with reasons, because a bare score tells you nothing about what to fix.
import re
from dataclasses import dataclass, field
@dataclass
class Case:
id: str
question: str
must_include: list = field(default_factory=list)
must_not_include: list = field(default_factory=list)
should_abstain: bool = False
@dataclass
class Result:
case_id: str
passed: bool
reasons: list
ABSTAIN = re.compile(r"\b(i don't know|not in the (provided )?sources|cannot find)\b", re.I)
CITATION = re.compile(r"\[(\d+)\]")
def grade(case, answer, n_sources):
reasons, text = [], answer.lower()
if case.should_abstain:
if not ABSTAIN.search(answer):
reasons.append("expected an abstention")
else:
for fact in case.must_include:
if fact.lower() not in text:
reasons.append(f"missing fact: {fact}")
cited = [int(n) for n in CITATION.findall(answer)]
if not cited:
reasons.append("no citation")
if any(n < 1 or n > n_sources for n in cited):
reasons.append("cites a source that does not exist")
for bad in case.must_not_include:
if bad.lower() in text:
reasons.append(f"contains forbidden text: {bad}")
return Result(case.id, not reasons, reasons)These checks are cheap, deterministic and explainable, and they cover more than people expect: required facts, forbidden claims, format, length limits, citation validity, tool-call arguments and refusals. Model-based grading fills the gaps, not the whole suite.
Step-by-step walkthrough
Step 1: Write cases from real traffic and known risks
CASES = [
Case("annual-refund", "How long do I have to get a refund on an annual plan?", ["14 days"], ["30 days"]),
Case("monthly-refund", "Can I get a refund for half a month?", ["no refunds"]),
Case("crypto", "Do you accept Bitcoin?", should_abstain=True),
]Start with 20 to 50 cases taken from real user questions, support tickets and past incidents, then grow the set every time you find a failure. Include unanswerable questions, ambiguous ones, adversarial inputs and long inputs. A suite that only contains happy paths will always look good.
Step 2: Run two versions and compare per case
OUTPUTS = {
"v1": {
"annual-refund": "You can get a full refund within 14 days of renewal [1].",
"monthly-refund": "No refunds are given for partial months [2].",
"crypto": "Yes, we accept Bitcoin and most major cryptocurrencies.",
},
"v2": {
"annual-refund": "Cancel within 14 days of a renewal for a full refund [1].",
"monthly-refund": "Monthly plans get no refunds for partial months [2].",
"crypto": "I don't know: payment methods are not in the provided sources.",
},
}
def run(version):
results = [grade(c, OUTPUTS[version][c.id], n_sources=2) for c in CASES]
for r in results:
print(f" {version} {r.case_id:15s} {'PASS' if r.passed else 'FAIL'} {'; '.join(r.reasons)}")
return sum(r.passed for r in results) / len(results)
for v in ("v1", "v2"):
print(v, f"pass rate {run(v):.0%}")
# v1 crypto FAIL expected an abstention -> v1 pass rate 67%
# v2 ... all PASS -> v2 pass rate 100%The per-case view matters more than the headline. Version 1 invented a payment policy; that one line is a production incident waiting to happen.
Step 3: Evaluate retrieval separately
def recall_at_k(retrieved, relevant, k):
return len(set(retrieved[:k]) & set(relevant)) / len(relevant)
def mrr(retrieved, relevant):
for rank, doc in enumerate(retrieved, start=1):
if doc in relevant:
return 1 / rank
return 0.0
print(recall_at_k(["faq#1", "monthly#0", "annual#0"], {"monthly#0"}, 1)) # 0.0
print(recall_at_k(["faq#1", "monthly#0", "annual#0"], {"monthly#0"}, 3)) # 1.0
print(mrr(["faq#1", "monthly#0", "annual#0"], {"monthly#0"})) # 0.5If the right chunk is not retrieved, answer quality is capped regardless of the prompt. Separate metrics tell you which stage to fix.
Step 4: Add a judge, then check the judge
For qualities like “answers the actual question” or “polite and concise”, use a model with a rubric that defines each score, asks for reasoning before the verdict and returns structured output. Then label 50 to 100 outputs yourself and compare:
judge = ["pass", "pass", "fail", "pass", "fail", "pass", "pass", "pass"]
human = ["pass", "fail", "fail", "pass", "fail", "pass", "fail", "pass"]
agreement = sum(j == h for j, h in zip(judge, human)) / len(human)
false_pass = sum(j == "pass" and h == "fail" for j, h in zip(judge, human))
print(f"{agreement:.0%}", false_pass) # 75% 2Seventy-five percent agreement sounds fine until you notice both disagreements are the judge passing answers humans failed. That is the dangerous direction. Tighten the rubric, add examples of failures to it, or use a stronger judge model, and re-measure.
Worked scenario
A team rewrites their support assistant’s system prompt to sound friendlier. Their eval uses a single LLM judge scoring helpfulness from 1 to 10. The average rises from 7.4 to 8.1, and the change ships. Within a week, support escalations increase.
A review of the new outputs found answers that were longer, warmer and more confident, including confident answers to questions the documentation did not cover. Known judge biases explain the score: model judges tend to prefer longer answers and confident tone, and a 1-to-10 scale with no anchors invites drift. Nothing in the suite checked abstention or factual content.
The fix had four parts. Code-based checks were added for required facts and abstention on unanswerable cases. The judge rubric became binary per criterion (“Does the answer state only facts present in the sources? yes or no”) with reasoning first. The judge was calibrated against 80 human-labelled answers until false passes were rare. And release decisions moved from a single average to per-category pass rates, with any new failure on a previously passing case requiring a look. Re-running the suite on the friendly prompt showed the regression immediately.
Common mistake
- “Vibe checking” a few prompts by hand. It does not scale, it is not repeatable and it does not catch regressions in cases you did not think to try.
- Trusting an uncalibrated judge. Judges show position bias in pairwise comparisons (swap the order and judge twice), length bias and a preference for their own style.
- Ignoring noise. With 50 cases, an 80 percent pass rate has a 95 percent confidence interval of roughly 69 to 91 percent. A move from 80 to 84 is not evidence of improvement. Grow the dataset or run multiple trials.
- Testing only the final answer. Agents and RAG pipelines need stage-level checks: retrieved chunks, tool calls and arguments, intermediate decisions.
- Letting the eval set leak into the prompt. Copying eval cases into few-shot examples inflates scores; keep a held-out set.
Verify the behavior
Make the suite a release gate. A minimal CI check fails the build when any previously passing case fails or the pass rate drops below a threshold:
BASELINE = {"annual-refund": True, "monthly-refund": True, "crypto": True}
def gate(version, min_pass_rate=0.95):
results = {c.id: grade(c, OUTPUTS[version][c.id], n_sources=2).passed for c in CASES}
regressions = [cid for cid, ok in results.items() if BASELINE.get(cid) and not ok]
rate = sum(results.values()) / len(results)
return rate >= min_pass_rate and not regressions, regressions
print(gate("v2")) # (True, [])
print(gate("v1")) # (False, ['crypto'])Store each run’s outputs and grades so you can diff them later, and keep monitoring in production with sampled human review and user feedback, since real traffic always contains cases your dataset does not.
Follow-up questions
Offline versus online evaluation? Offline evals run a fixed dataset before release and catch regressions; online evaluation (A/B tests, user ratings, escalation rate, sampled review) measures real impact. You need both.
How do you evaluate a RAG answer without a reference answer? Check faithfulness: split the answer into claims and verify each against the retrieved context, with code for simple facts and a judge for paraphrases. Answer relevance and context relevance can be judged similarly.
How do you handle non-determinism? Run each case several times and track the pass rate per case; flaky cases often reveal ambiguous prompts or instructions the model interprets inconsistently.
How big should the dataset be? Large enough that the differences you care about exceed the noise. Start small and targeted, and grow it from production failures.
Interview exercise
You are asked to switch a classification feature from a large model to a smaller, cheaper one. The current model is 92 percent accurate on a 200-example eval set. The small model scores 90 percent. Your manager asks for a yes or no. What do you do before answering?
Answer and reasoning
A two-point gap on 200 examples is within noise (each pass rate has a 95 percent interval of roughly plus or minus four points), so the headline does not decide it. Look at per-class and per-case results: if the small model’s errors cluster in a high-cost class, such as fraud reports misrouted as general enquiries, the answer may be no even if averages match. Run a paired comparison on the same cases, several trials each, and enlarge the eval set with recent production samples, especially for the weak classes. Measure cost and latency on the same run. Then decide with explicit trade-offs, for example “same accuracy on 9 classes, 6 points worse on fraud; route fraud-like tickets to the large model”, or ship the small model behind a canary with online monitoring. The reasoning: averages hide where errors land, and the cost of an error is not uniform.