Ch. 30 · AI & LLM Engineering

How to Design a RAG Pipeline: Chunking, Retrieval and Prompts

A step-by-step RAG design for interviews: ingestion, structure-aware chunking, hybrid retrieval, grounded prompts with citations and evaluation.

~9 min readintermediateupdated Oct 6, 2026

“Design a chatbot that answers questions about our internal documentation” is the most common system design prompt in AI engineering interviews, and the expected architecture is retrieval-augmented generation (RAG). The interviewer is not checking whether you can draw “documents, vector DB, LLM” in three boxes. They want to hear how documents become chunks, how you retrieve the right ones, how the prompt keeps the model grounded, what happens when the answer is not in the corpus, and how you would know the system works. This guide walks through each stage with runnable code and one real failure.

Before you start

You should know what an embedding is and what nearest-neighbour search does (see the embeddings guide linked below), and be comfortable reading Python. The code uses Python 3.12+ and only the standard library. Retrieval uses BM25, a classic keyword ranking function, so the example runs anywhere; in production you would combine it with vector search. The language model call is left out on purpose: everything that decides answer quality happens before it.

The short answer

RAG has two pipelines. Offline ingestion loads documents, cleans them, splits them into chunks, attaches metadata (source, section, permissions, date), embeds them and writes them to a keyword index and a vector index. Online answering takes a question, optionally rewrites it, retrieves candidates from both indexes, merges and reranks them, puts the best few into a prompt that tells the model to answer only from those sources and cite them, and returns the answer with links. The model never “learns” the documents; it reads them at request time, so updates take effect as soon as the index is refreshed. Most quality problems trace back to retrieval, and most retrieval problems trace back to chunking.

How it works

Every stage makes a decision you should be able to defend:

  • Parsing. PDFs, HTML and slides must become clean text that keeps headings, lists and tables. Garbage in the text becomes garbage in the embeddings.
  • Chunking. Chunks must be small enough to be specific and to fit several in the prompt, but large enough to contain a complete answer. Splitting along headings, paragraphs and sentences beats splitting every N tokens.
  • Indexing. Store chunk text, embeddings and metadata. Access-control metadata matters: retrieval must filter by what the user is allowed to see, or the model will happily quote a document the user should never read.
  • Retrieval. Hybrid retrieval (keyword plus vector) with a reranker over the top 20 to 50 candidates is a strong default.
  • Generation. The prompt separates instructions from sources, numbers the sources, requires citations and defines what to do when the sources are silent.

The code below is the core of the example: two chunkers, a BM25 scorer and a prompt builder.

import math, re
from collections import Counter

POLICY = """# Billing policy

## Annual plans
These renew once a year on the purchase date. We send a reminder email seven days before each renewal. To cancel, open Settings and choose Billing. A full refund is available if you cancel within 14 days of a renewal.

## Monthly plans
These renew every month. Cancelling stops the next renewal, and access continues until the paid month ends. No refunds are given for partial months."""

def words(text):
    return re.findall(r"[a-z0-9]+", text.lower())

def fixed_chunks(text, size, overlap=0):
    tokens = text.split()
    step = size - overlap
    return [" ".join(tokens[i:i + size]) for i in range(0, len(tokens), step)]

def section_chunks(markdown, max_words=80):
    chunks, doc_title = [], ""
    for block in re.split(r"\n(?=#)", markdown):
        heading, _, body = block.partition("\n")
        if heading.startswith("# "):
            doc_title = heading[2:].strip()
            continue
        title = heading.lstrip("#").strip()
        current = []
        for sentence in re.split(r"(?<=[.!?])\s+", body.strip()):
            if current and len(" ".join(current + [sentence]).split()) > max_words:
                chunks.append(f"{doc_title} > {title}: " + " ".join(current)); current = []
            current.append(sentence)
        if current:
            chunks.append(f"{doc_title} > {title}: " + " ".join(current))
    return chunks
python

section_chunks never cuts a sentence, never mixes two sections, and prefixes every chunk with its document and section title. That prefix is cheap and powerful: a chunk that says “These renew once a year” is meaningless on its own, but “Billing policy > Annual plans: These renew once a year” is searchable. Generating a short, chunk-specific context with an LLM at ingestion time takes the same idea further.

Step-by-step walkthrough

Step 1: Index the chunks with BM25

class BM25:
    def __init__(self, docs, k1=1.5, b=0.75):
        self.docs = [words(d) for d in docs]
        self.avg = sum(map(len, self.docs)) / len(self.docs)
        self.df = Counter(t for d in self.docs for t in set(d))
        self.k1, self.b, self.n = k1, b, len(docs)

    def score(self, query, i):
        doc, tf = self.docs[i], Counter(self.docs[i])
        total = 0.0
        for t in words(query):
            if t not in tf:
                continue
            idf = math.log(1 + (self.n - self.df[t] + 0.5) / (self.df[t] + 0.5))
            total += idf * tf[t] * (self.k1 + 1) / (tf[t] + self.k1 * (1 - self.b + self.b * len(doc) / self.avg))
        return total

    def top(self, query, k):
        return sorted(range(self.n), key=lambda i: self.score(query, i), reverse=True)[:k]
python

BM25 rewards query terms that are rare across the corpus (idf), with diminishing returns for repeated terms (k1) and a penalty for long chunks (b). It is the keyword half of hybrid retrieval and catches exact terms such as product names and error codes that embeddings blur.

Step 2: Retrieve with each chunking strategy

question = "How long do I have to get a refund on an annual plan?"
strategies = [("fixed-30", fixed_chunks(POLICY, 30)),
              ("fixed-30-overlap-10", fixed_chunks(POLICY, 30, 10)),
              ("sections", section_chunks(POLICY))]
for name, chunks in strategies:
    best = chunks[BM25(chunks).top(question, 1)[0]]
    print(name, len(chunks), "chunks; top-1 contains '14 days':", "14 days" in best)
# fixed-30 3 chunks; top-1 contains '14 days': False
# fixed-30-overlap-10 4 chunks; top-1 contains '14 days': False
# sections 2 chunks; top-1 contains '14 days': True
python

Step 3: Build a grounded prompt

def build_prompt(question, chunks):
    sources = "\n".join(f"[{i + 1}] {c}" for i, c in enumerate(chunks))
    return (
        "Answer using only the sources below. Cite sources like [1]. "
        "If the sources do not contain the answer, say you don't know.\n\n"
        f"Sources:\n{sources}\n\nQuestion: {question}"
    )

print(build_prompt(question, [section_chunks(POLICY)[0]]))
python

The instruction does three jobs: it limits the model to the sources, it makes every claim traceable through a citation number, and it gives the model a legitimate way out. Without that last clause, a model asked about something absent from the sources tends to answer from its training data, which is exactly the hallucination RAG was supposed to prevent. Put the question after the sources, and keep instructions in the system prompt when your API has one.

Step 4: Decide how many chunks to pass

Retrieval returns a ranked list; the prompt gets the top few. Too few and the answer may be missing; too many and you pay for tokens, add latency and bury the relevant passage among distractors. Start with 3 to 8 reranked chunks, then tune using the evaluation in the next sections rather than intuition.

Worked scenario

The billing assistant above shipped with fixed 30-word chunks, the default in the team’s first prototype. Customers asking “How long do I have to get a refund on an annual plan?” received answers like “You’ll get a reminder seven days before renewal” or “I don’t know”. The source document clearly says 14 days.

Printing the chunks showed the cause. The fixed splitter produced:

  1. “# Billing policy ## Annual plans These renew once a year on the purchase date. We send a reminder email seven days before each renewal. To cancel, open Settings and”
  2. “choose Billing. A full refund is available if you cancel within 14 days of a renewal. ## Monthly plans These renew every month. …”

The 14-day rule lives in a chunk that starts mid-sentence, has no “annual” in it, and is glued to the monthly plan section. The first chunk matches “annual plan” strongly, so it wins retrieval, and it contains everything except the answer. Adding a 10-word overlap did not help, because the problem is not a sentence cut at a boundary; it is that the answer was separated from the words that identify its subject.

Switching to section_chunks produced one chunk per section, each prefixed with “Billing policy > Annual plans” or “Billing policy > Monthly plans”. The annual chunk now contains both the subject and the 14-day rule, and the model answers “within 14 days of a renewal [1]”. The team then added the regression test below so a future change to chunk size could not reintroduce the bug.

Common mistake

  • Blaming the model first. When a RAG answer is wrong, inspect the retrieved chunks before touching the prompt. If the answer is not in them, no prompt will fix it.
  • One global chunk size. FAQs, API references, contracts and chat logs have different structures; chunk each by its structure.
  • Dropping metadata. Without source, section and permission metadata you cannot cite, filter or enforce access control.
  • Vector search only. Exact identifiers and rare terms need keyword retrieval; use hybrid search.
  • No abstain path. The prompt must say what to do when the sources do not contain the answer, and your evaluation must test it.
  • Stale indexes. Deleted or updated documents must be removed or re-embedded; schedule incremental re-indexing and track document versions.

Verify the behavior

Retrieval deserves its own tests, separate from answer quality. A small golden set of questions with the chunk that must be retrieved is enough to start:

GOLDEN = [
    ("How long do I have to get a refund on an annual plan?", "14 days"),
    ("Do I get money back for half a month?", "No refunds"),
]

def test_retrieval_finds_the_answer_chunk():
    chunks = section_chunks(POLICY)
    index = BM25(chunks)
    for question, must_contain in GOLDEN:
        top = [chunks[i] for i in index.top(question, 2)]
        assert any(must_contain in c for c in top), question

test_retrieval_finds_the_answer_chunk(); print("ok")
python

At scale, track recall@k and mean reciprocal rank over a few hundred real questions, and evaluate answers separately for faithfulness to the sources and for correctness.

Follow-up questions

How would you handle a question that needs facts from two documents? Retrieve enough candidates to cover both, or decompose the question into sub-queries, retrieve for each and combine. Agentic retrieval, where the model issues its own searches, helps here but costs latency.

What is query rewriting? Turning a conversational follow-up such as “what about monthly?” into a standalone query (“refund policy for monthly plans”) using the conversation history, so retrieval sees the full intent.

How do you enforce permissions? Store access-control lists as chunk metadata and filter at query time with the user’s identity, inside the retrieval call. Never rely on the prompt to hide documents.

When is long context better than RAG? When the corpus is small and stable enough to fit in the window, sending everything (ideally with prompt caching) can be simpler and more accurate. RAG wins on large, changing or permissioned corpora.

Interview exercise

Your RAG assistant over 50,000 engineering wiki pages has good answers in testing, but in production 30 percent of answers are rated unhelpful. Logs show retrieval returns chunks from pages last edited in 2019 that describe deprecated systems, while newer pages exist. How would you diagnose and fix this?

Answer and reasoning

First confirm it is a retrieval problem: sample the unhelpful answers and check whether the current page was in the top-k candidates at all, or present but outranked. If it was outranked, older pages probably match better because they are longer and use the exact old terminology. Fixes, in order: add last_modified and an “archived” flag to metadata and filter out archived content; add a recency signal in reranking for content where freshness matters; deduplicate near-identical pages so stale copies do not crowd out the current one; and give the owners a process to archive deprecated pages, because no ranking trick beats removing wrong content. Then add these failing questions to the golden set so the improvement is measured and cannot regress. The reasoning: diagnose by stage, fix the data before tuning the model, and turn every incident into an evaluation case.

Continue learning

More in AI & LLM Engineering

read ✓AI & LLM Engineering · hard

Fine-Tuning vs RAG: When to Use Each in LLM Apps

Fine-tuning changes how a model behaves; RAG changes what it knows at request time. A decision guide with data prep, costs and failure cases.

~9 min readread →
esc