Ch. 30 · AI & LLM Engineering

How to Reduce LLM Hallucinations in Production

Why LLMs hallucinate and the layered fixes that work: grounding, permission to abstain, citations, structured output and automatic claim checks.

~8 min readintermediateupdated Oct 6, 2026

“How would you reduce hallucinations in an LLM application?” is asked in nearly every AI engineering interview, often as a follow-up to a RAG design question. The trap is to give a single answer (“use RAG”, “set temperature to zero”, “tell it not to make things up”). Interviewers want a layered answer: why hallucinations happen at all, which design choices make them less likely, how you detect the ones that still get through, and how you measure whether your changes worked. This guide builds those layers in code around one realistic failure.

Before you start

You should know what retrieval-augmented generation is, roughly how a prompt with retrieved sources looks, and basic Python including regular expressions and json. The examples run on Python 3.12+ with the standard library. They focus on checks you run in your own code before an answer reaches a user, which work with any model provider.

The short answer

A model hallucinates because it generates the most plausible continuation, not a looked-up fact; when it lacks the information, a fluent guess is often more probable than “I don’t know”. Reduce it in layers: ground the model in retrieved, relevant sources; give it explicit permission and a format to abstain; require citations to numbered sources; use tools for things models do badly, such as arithmetic, dates and live data; use structured output so answers are machine-checkable; then verify after generation (citation validity, numbers and quotes present in the cited source, a judge model for paraphrased claims) and abstain or escalate when checks fail. Measure the rate with evals, because no single technique removes hallucinations entirely.

How it works

It helps to separate two kinds of error. A factuality error contradicts the world (“the Eiffel Tower is in Rome”). A faithfulness error contradicts or goes beyond the provided sources, even if it happens to be true elsewhere. In a RAG product, faithfulness is the property you can engineer and test, because you know exactly what the model was given.

Hallucinations get more likely when the answer is not in the context, the question is ambiguous, the prompt pressures the model to always answer, the task needs exact recall of numbers or names, or the context is so long that the relevant passage is buried. Each mitigation targets one of these causes:

Cause Mitigation
Answer not in context Better retrieval; abstain when retrieval is weak
Pressure to answer Explicit “say you don’t know” path, rewarded in evals
Unverifiable output Citations, quotes, structured fields
Arithmetic, dates, live data Tool calls instead of recall
Long, noisy context Fewer, reranked chunks; extract relevant quotes first

Step-by-step walkthrough

Step 1: Abstain before generating when retrieval is weak

If the best retrieval score is below a threshold you calibrated on labelled questions, do not ask the model to answer from irrelevant text:

def answer_or_abstain(scores, threshold):
    return "answer" if scores and max(scores) >= threshold else "abstain"

print(answer_or_abstain([0.31, 0.28], threshold=0.5))  # abstain
print(answer_or_abstain([0.82, 0.40], threshold=0.5))  # answer
python

The threshold is model- and corpus-specific; pick it by plotting answer accuracy against score on real questions. Abstaining here also saves a model call.

Step 2: Prompt for grounded, citable answers

A grounded prompt states the rules plainly: use only the numbered sources, cite each claim, quote exact figures, and reply in a defined way when the sources do not contain the answer. For long sources, asking the model to first extract the relevant quotes and then answer from those quotes keeps it anchored to the text. Lower or default temperature suits factual tasks, but temperature alone does not prevent confident fabrication.

Step 3: Request structured output and validate it

import json

def parse_answer(raw, n_sources):
    try:
        data = json.loads(raw)
    except json.JSONDecodeError as exc:
        return None, f"invalid JSON: {exc.msg}"
    if set(data) != {"answer", "citations", "answerable"}:
        return None, f"unexpected keys {sorted(data)}"
    if not isinstance(data["answerable"], bool):
        return None, "answerable must be a boolean"
    if data["answerable"] and not data["citations"]:
        return None, "answerable answers need at least one citation"
    if any(not isinstance(c, int) or not 1 <= c <= n_sources for c in data["citations"]):
        return None, "citation out of range"
    return data, None

print(parse_answer('{"answer": "Within 14 days of renewal.", "citations": [1], "answerable": true}', 2))
# ({'answer': 'Within 14 days of renewal.', 'citations': [1], 'answerable': True}, None)
print(parse_answer('{"answer": "Within 30 days.", "citations": [], "answerable": true}', 2))
# (None, 'answerable answers need at least one citation')
print(parse_answer('{"answer": "Within 30 days.", "citations": [3], "answerable": true}', 2))
# (None, 'citation out of range')
print(parse_answer('{"answer": "Not covered by the sources.", "citations": [], "answerable": false}', 2))
# ({'answer': 'Not covered by the sources.', 'citations': [], 'answerable': False}, None)
python

Many providers can constrain output to a JSON Schema, which removes the parse failures. It does not stop the model from putting a wrong number inside a perfectly valid field, which is why the next step exists.

Step 4: Check claims against the cited sources

import re

SOURCES = {
    1: "A full refund is available if you cancel within 14 days of a renewal.",
    2: "No refunds are given for partial months.",
}

def split_claims(answer):
    return [s.strip() for s in re.split(r"(?<=[.!?])\s+", answer) if s.strip()]

def check_claim(claim, sources):
    cited = [int(n) for n in re.findall(r"\[(\d+)\]", claim)]
    if not cited:
        return "unsupported: no citation"
    if any(n not in sources for n in cited):
        return "invalid: cites a missing source"
    numbers = re.findall(r"\d+(?:\.\d+)?", re.sub(r"\[\d+\]", "", claim))
    cited_text = " ".join(sources[n] for n in cited)
    unsupported = [n for n in numbers if n not in re.findall(r"\d+(?:\.\d+)?", cited_text)]
    if unsupported:
        return f"suspect: {unsupported} not in cited source"
    return "ok"
python

This check is crude by design: it is fast, free and catches the most damaging class of error, wrong numbers. Claims that pass can go to a judge model that answers “is this claim supported by this source, yes or no” for paraphrases the regex cannot verify.

Worked scenario

A subscription company’s assistant answered “Annual plans can be refunded within 30 days of renewal [1]. Partial months are not refunded [2]. Refunds arrive in 5 business days.” The policy says 14 days, and nothing in the documentation mentions refund timing. Thirty days is a common refund window in general, so it was a very plausible guess; five business days is plausible too. Both were confidently wrong, and one carried a citation, which made it look verified.

Running the checker on that answer shows exactly what a verification layer would have caught:

answer = ("Annual plans can be refunded within 30 days of renewal [1]. "
          "Partial months are not refunded [2]. Refunds arrive in 5 business days.")
for claim in split_claims(answer):
    print(f"{check_claim(claim, SOURCES):40s} | {claim}")
# suspect: ['30'] not in cited source      | Annual plans can be refunded within 30 days of renewal [1].
# ok                                       | Partial months are not refunded [2].
# unsupported: no citation                 | Refunds arrive in 5 business days.
python

The team changed the pipeline in three ways. The prompt now requires quoting the exact figure from the source and offers a defined “not covered” answer. Answers pass through parse_answer and check_claim; a suspect or unsupported claim triggers one regeneration with the failure reason, and a second failure returns a safe fallback that links to the policy page. And every incident like this became an eval case, so the hallucination rate on known traps is tracked on each release.

Common mistake

  • “RAG eliminates hallucinations.” It reduces them when retrieval works. When the answer is missing from the retrieved chunks, the model may fill the gap from training data.
  • “Temperature 0 stops hallucinations.” It makes the most likely answer more consistent, including a consistently wrong one.
  • “Citations prove correctness.” Models can cite a source that does not support the claim. Verify the citation, not just its presence.
  • “Structured output means correct output.” It guarantees shape, not truth.
  • Penalizing “I don’t know” in evals. If the eval rewards any answer over abstention, prompt tuning will teach the system to guess.
  • Asking the model to compute. Dates, totals and conversions should be tool calls or code.

Verify the behavior

Encode the traps as tests so a prompt or model change cannot silently reintroduce them:

def test_wrong_number_is_flagged():
    assert check_claim("Refunds within 30 days [1].", SOURCES).startswith("suspect")

def test_correct_number_passes():
    assert check_claim("Refunds within 14 days of a renewal [1].", SOURCES) == "ok"

def test_uncited_claim_is_flagged():
    assert check_claim("Refunds arrive in 5 business days.", SOURCES).startswith("unsupported")

def test_abstention_is_valid_output():
    data, error = parse_answer('{"answer": "Not covered.", "citations": [], "answerable": false}', 2)
    assert error is None and data["answerable"] is False

for t in (test_wrong_number_is_flagged, test_correct_number_passes, test_uncited_claim_is_flagged, test_abstention_is_valid_output):
    t()
print("ok")
python

In the eval suite, track two rates over time: unsupported claims per answer on answerable questions, and correct abstention on unanswerable ones. Improving one at the expense of the other is easy; improving both is the goal.

Follow-up questions

Why do models hallucinate at all? They are trained to predict likely text, and training and evaluation often reward a confident answer over an abstention, so guessing is learned behaviour. They also have a knowledge cutoff and no built-in way to know what they do not know.

How do you detect hallucinations without a reference answer? Check faithfulness against the provided context: claim-level verification with code and a judge, or sample several answers and treat disagreement between them as a warning sign.

Would fine-tuning fix it? Fine-tuning can teach format and when to abstain in a narrow domain, but it is a poor way to add facts, and it can make a model more confident without making it more correct. Grounding is the primary tool for facts.

What do you do in high-stakes domains? Narrow the scope, show sources next to every answer, require human review for consequential outputs, and log enough to audit any answer later.

Interview exercise

A legal-research assistant summarizes case law from a retrieval system. Lawyers report that it occasionally cites cases that do not exist. Retrieval logs show the fabricated case names never appear in the retrieved documents. Design the fix and explain how you would prove it worked.

Answer and reasoning

Because the fabricated names never appear in the retrieved context, this is a faithfulness failure in generation, which is checkable. Require structured output in which every cited case is a reference to a retrieved document id rather than free text, and render case names from your database using that id, never from model text. Run a post-generation check that every case name or citation pattern in the summary matches a retrieved document; any mismatch blocks the answer and triggers a regeneration with the error, then a fallback that shows the retrieved sources without a summary. Add an explicit instruction and format for “no relevant cases found”. To prove it, build an eval set from the reported incidents plus queries designed to have no good matches, and measure the fabricated-citation rate before and after; the target is zero for the id-based design, since fabricated ids are rejected by code. The reasoning: when a mistake is catastrophic, remove the model’s ability to express it rather than asking it to be careful.

Continue learning

More in AI & LLM Engineering

read ✓AI & LLM Engineering · hard

Fine-Tuning vs RAG: When to Use Each in LLM Apps

Fine-tuning changes how a model behaves; RAG changes what it knows at request time. A decision guide with data prep, costs and failure cases.

~9 min readread →
esc