Ch. 30 · AI & LLM Engineering

Fine-Tuning vs RAG: When to Use Each in LLM Apps

Fine-tuning changes how a model behaves; RAG changes what it knows at request time. A decision guide with data prep, costs and failure cases.

~9 min readadvancedupdated Oct 6, 2026

“Should we fine-tune a model on our documents or build RAG?” is a favourite senior-level interview question because the wrong answer is expensive and common. Teams fine-tune to “teach the model our data”, then discover it still invents facts, cannot cite anything, and is out of date the day after training. Interviewers want a decision framework: what each technique actually changes, which problems each solves, what it costs to build and run, and how you would prove the choice with an evaluation instead of a hunch.

Before you start

You should know what RAG is (retrieve relevant documents and put them in the prompt) and have a rough idea of what training a model means: adjusting its weights from examples. Familiarity with evaluation sets is important, since every decision here is settled by measurement. The examples use Python 3.12+ and the standard library: a dataset validator for a chat-format fine-tuning file and a cost comparison where you supply the prices, because prices change and differ by provider.

The short answer

RAG changes what the model knows at request time; fine-tuning changes how the model behaves by updating its weights. Choose RAG when the knowledge is large, changes often, must be cited, or must be filtered by user permissions. Choose fine-tuning when a well-prompted model has the knowledge but not the behaviour: a strict output format, a house style, a narrow classification task, or when you want a smaller, cheaper model to match a larger one on one task, with shorter prompts. Start with prompting and few-shot examples, add retrieval for knowledge gaps, and fine-tune only when evals show a behaviour gap that prompting cannot close. The two combine well: a model fine-tuned to use retrieved context and cite it, fed by a RAG pipeline.

How it works

A fine-tuning job trains on input and output pairs. Supervised fine-tuning (SFT) shows the model example conversations and adjusts weights so it is more likely to produce the target output for similar inputs. Preference methods such as DPO train on pairs of better and worse answers. Full fine-tuning updates every weight; parameter-efficient methods such as LoRA freeze the base model and train small low-rank adapter matrices, which cuts memory and lets you keep one base model with many small adapters. QLoRA does the same on a quantized base model to fit training on smaller GPUs. Hosted APIs hide these details and accept a JSONL file of chat examples.

What training is good at is learning patterns: output structure, tone, which label goes with which kind of input, how long answers should be. It is unreliable at storing facts you can retrieve exactly and update later. Facts seen a few times in training are recalled imprecisely, mixed with what the model already believed, and cannot be removed when they change. RAG keeps facts outside the model, where you can update, delete, cite and permission them.

Need RAG Fine-tuning
Fresh or frequently changing knowledge Yes: re-index No: retrain
Citations and provenance Yes No
Per-user access control Yes: filter retrieval No
Consistent format, tone or labels Partly, via prompt Yes
Shorter prompts, lower latency per call No: adds context Yes
Smaller model matching a larger one on one task No Yes (distillation)

Step-by-step walkthrough

Step 1: Establish a baseline with prompting and an eval

Before any training, write the best prompt you can, add a handful of few-shot examples, and measure on a held-out eval set. Then classify the failures. If the model lacks information (wrong product details, unknown internal terms), that is a knowledge gap: add retrieval. If it has the information but formats it wrongly, ignores the style or confuses similar labels, that is a behaviour gap and a fine-tuning candidate.

Step 2: Prepare and validate training data

A support-ticket classifier is a classic fine-tuning task. The data must look exactly like production traffic, including the system prompt, and must be clean:

import hashlib, json, random, re

LABELS = {"billing", "bug", "how-to", "account"}
EMAIL = re.compile(r"[\w.+-]+@[\w-]+\.[\w.]+")
SYSTEM = {"role": "system", "content": "Classify the support ticket. Reply with one label."}

EXAMPLES = [
    {"messages": [SYSTEM, {"role": "user", "content": "I was charged twice for March."},
                  {"role": "assistant", "content": "billing"}]},
    {"messages": [SYSTEM, {"role": "user", "content": "The app crashes when I upload a photo."},
                  {"role": "assistant", "content": "bug"}]},
    {"messages": [SYSTEM, {"role": "user", "content": "I was charged twice for March."},
                  {"role": "assistant", "content": "billing"}]},
    {"messages": [SYSTEM, {"role": "user", "content": "Email me at jane.doe@example.com about the refund"},
                  {"role": "assistant", "content": "billing"}]},
    {"messages": [{"role": "user", "content": "How do I export my data?"},
                  {"role": "assistant", "content": "Feature"}]},
]

def problems(example):
    msgs, found = example.get("messages", []), []
    if not msgs or msgs[-1]["role"] != "assistant":
        found.append("last message must be the assistant target")
    if msgs and msgs[0]["role"] != "system":
        found.append("missing the system prompt used in production")
    if msgs and msgs[-1]["content"] not in LABELS:
        found.append(f"unknown label {msgs[-1]['content']!r}")
    if any(EMAIL.search(m["content"]) for m in msgs):
        found.append("contains an email address")
    return found

def clean(examples):
    seen, kept = set(), []
    for i, ex in enumerate(examples):
        key = hashlib.sha256(json.dumps(ex, sort_keys=True).encode()).hexdigest()
        if key in seen:
            print(f"example {i}: duplicate, dropped"); continue
        seen.add(key)
        if issues := problems(ex):
            print(f"example {i}: {'; '.join(issues)}"); continue
        kept.append(ex)
    return kept

kept = clean(EXAMPLES)
# example 2: duplicate, dropped
# example 3: contains an email address
# example 4: missing the system prompt used in production; unknown label 'Feature'
print(len(kept))  # 2
python

Duplicates over-weight some inputs, personal data should not be baked into weights, and an invalid label teaches the model a class that does not exist. Real datasets need the same checks at scale, plus a held-out split the model never trains on.

Step 3: Train, then evaluate against the baseline

Shuffle, split into training and validation sets, train, and run the same eval as in Step 1 on the tuned model. Compare per class, not just overall accuracy, and also re-run a few general tasks the application still relies on, because tuning on a narrow task can degrade others.

random.Random(0).shuffle(kept)
split = max(1, int(len(kept) * 0.8))
train, validation = kept[:split], kept[split:]
print(len(train), len(validation))  # 1 1 with this toy data; use hundreds of examples in practice
python

Step 4: Compare the running costs

def monthly_cost(requests, input_tokens, output_tokens, in_price, out_price):
    """Prices are per million tokens; pass your provider's current numbers."""
    return requests * (input_tokens * in_price + output_tokens * out_price) / 1_000_000

few_shot = monthly_cost(1_000_000, input_tokens=3_500, output_tokens=5, in_price=1.0, out_price=4.0)
tuned = monthly_cost(1_000_000, input_tokens=150, output_tokens=5, in_price=1.5, out_price=6.0)
print(f"few-shot prompt: ${few_shot:,.0f}/month, fine-tuned short prompt: ${tuned:,.0f}/month")
# few-shot prompt: $3,520/month, fine-tuned short prompt: $255/month
python

With illustrative prices, the fine-tuned model costs more per token yet far less per request, because dozens of few-shot examples no longer travel with every call. Add training cost and the engineering time to maintain the dataset. Also compare with prompt caching, which can make a long, fixed few-shot prefix much cheaper without any training.

Worked scenario

A retailer fine-tuned a model on 20,000 product pages so its assistant would “know the catalogue”. Early demos impressed. Three months later, problems piled up: the assistant quoted last season’s prices, recommended discontinued products, invented specifications for products that resembled ones in the training set, and could not say where an answer came from. Each catalogue update meant another training run, and the legal team asked for a way to remove a product’s data, which the weights could not provide.

The redesign split the problem by type. Product facts moved into a RAG pipeline over the live catalogue with metadata for availability and region, so updates took effect on re-index and every answer linked to the product page. The fine-tuning budget went to behaviour: a small set of curated conversations teaching the model to answer from provided product data, ask a clarifying question when several products matched, format comparisons as tables and say when a product was unavailable. Evals tracked factual accuracy against the catalogue and format compliance separately. Accuracy rose and the retraining treadmill stopped, because facts no longer lived in the weights.

Common mistake

  • “Fine-tune to add knowledge.” It is an unreliable way to store facts and offers no citations, updates or deletions.
  • Fine-tuning before trying prompting and retrieval. Many apparent needs for training disappear with a better prompt, examples or retrieval.
  • Training on unrepresentative data. If the training system prompt, formatting or input distribution differs from production, the gains do not transfer.
  • No held-out evaluation. Training loss going down does not mean the task got better.
  • Forgetting maintenance. When the base model is deprecated or upgraded, the fine-tune must be redone and re-evaluated.
  • Ignoring cheaper levers. A smaller model with good prompts, prompt caching or batching may solve the cost problem without training.

Verify the behavior

Make the dataset rules executable so every new training file is checked before a job starts:

def test_dataset_rules():
    bad_label = {"messages": [SYSTEM, {"role": "user", "content": "x"}, {"role": "assistant", "content": "spam"}]}
    no_system = {"messages": [{"role": "user", "content": "x"}, {"role": "assistant", "content": "bug"}]}
    good = {"messages": [SYSTEM, {"role": "user", "content": "Login fails"}, {"role": "assistant", "content": "account"}]}
    assert "unknown label 'spam'" in problems(bad_label)
    assert "missing the system prompt used in production" in problems(no_system)
    assert problems(good) == []

test_dataset_rules(); print("ok")
python

The decisive verification is the comparison eval: baseline prompt, prompt plus retrieval, and fine-tuned model, on the same held-out cases, with accuracy per class, cost per request and latency side by side.

Follow-up questions

What is LoRA and why is it popular? Low-rank adaptation trains small adapter matrices added to frozen weights. It needs far less memory and storage than full fine-tuning, and adapters can be swapped per task on one base model.

What is distillation? Using a large model’s outputs to fine-tune a smaller model on one task, so the small model approaches the large model’s quality at lower cost and latency.

How much data do you need? It depends on the task; narrow format or classification tasks can improve with dozens to hundreds of high-quality examples, and harder tasks need more. Quality, coverage of edge cases and consistency of labels matter more than raw volume.

Can you combine them? Yes, and it is common: retrieval supplies facts, and fine-tuning teaches the model to use retrieved context, cite it and follow your output format.

Interview exercise

A bank wants an assistant that answers employees’ questions about 3,000 internal policy documents, which change weekly, in a strict format with policy references. Each employee may only see policies for their region and role. The team proposes fine-tuning an open-weight model on the documents. Give your recommendation and justify it.

Answer and reasoning

Recommend RAG as the core, not fine-tuning on the documents. Weekly changes would mean weekly retraining with no way to remove superseded policies from the weights; policy references require citing a specific document and version, which retrieval provides directly; and region and role restrictions must be enforced by filtering retrieval with the employee’s entitlements, which weights cannot do. Fine-tuning could still help with behaviour: teaching the strict answer format and how to respond when policies conflict or are missing, ideally on synthetic or redacted examples with no policy content memorized. Prove it with an eval of real employee questions per region, measuring correctness, citation accuracy and leakage across permission boundaries (which must be zero), comparing prompt-only RAG with RAG plus a format fine-tune. The reasoning: map each requirement to the mechanism that can actually satisfy it, then let the eval decide the remainder.

Continue learning

More in AI & LLM Engineering

esc