“Design a chatbot that answers questions about our internal documentation” is the most common system design prompt in AI engineering interviews, and the expected architecture is retrieval-augmented generation (RAG). The interviewer is not checking whether you can draw “documents, vector DB, LLM” in three boxes. They want to hear how documents become chunks, how you retrieve the right ones, how the prompt keeps the model grounded, what happens when the answer is not in the corpus, and how you would know the system works. This guide walks through each stage with runnable code and one real failure.
Before you start
You should know what an embedding is and what nearest-neighbour search does (see the embeddings guide linked below), and be comfortable reading Python. The code uses Python 3.12+ and only the standard library. Retrieval uses BM25, a classic keyword ranking function, so the example runs anywhere; in production you would combine it with vector search. The language model call is left out on purpose: everything that decides answer quality happens before it.
The short answer
RAG has two pipelines. Offline ingestion loads documents, cleans them, splits them into chunks, attaches metadata (source, section, permissions, date), embeds them and writes them to a keyword index and a vector index. Online answering takes a question, optionally rewrites it, retrieves candidates from both indexes, merges and reranks them, puts the best few into a prompt that tells the model to answer only from those sources and cite them, and returns the answer with links. The model never “learns” the documents; it reads them at request time, so updates take effect as soon as the index is refreshed. Most quality problems trace back to retrieval, and most retrieval problems trace back to chunking.
How it works
Every stage makes a decision you should be able to defend:
- Parsing. PDFs, HTML and slides must become clean text that keeps headings, lists and tables. Garbage in the text becomes garbage in the embeddings.
- Chunking. Chunks must be small enough to be specific and to fit several in the prompt, but large enough to contain a complete answer. Splitting along headings, paragraphs and sentences beats splitting every N tokens.
- Indexing. Store chunk text, embeddings and metadata. Access-control metadata matters: retrieval must filter by what the user is allowed to see, or the model will happily quote a document the user should never read.
- Retrieval. Hybrid retrieval (keyword plus vector) with a reranker over the top 20 to 50 candidates is a strong default.
- Generation. The prompt separates instructions from sources, numbers the sources, requires citations and defines what to do when the sources are silent.
The code below is the core of the example: two chunkers, a BM25 scorer and a prompt builder.
import math, re
from collections import Counter
POLICY = """# Billing policy
## Annual plans
These renew once a year on the purchase date. We send a reminder email seven days before each renewal. To cancel, open Settings and choose Billing. A full refund is available if you cancel within 14 days of a renewal.
## Monthly plans
These renew every month. Cancelling stops the next renewal, and access continues until the paid month ends. No refunds are given for partial months."""
def words(text):
return re.findall(r"[a-z0-9]+", text.lower())
def fixed_chunks(text, size, overlap=0):
tokens = text.split()
step = size - overlap
return [" ".join(tokens[i:i + size]) for i in range(0, len(tokens), step)]
def section_chunks(markdown, max_words=80):
chunks, doc_title = [], ""
for block in re.split(r"\n(?=#)", markdown):
heading, _, body = block.partition("\n")
if heading.startswith("# "):
doc_title = heading[2:].strip()
continue
title = heading.lstrip("#").strip()
current = []
for sentence in re.split(r"(?<=[.!?])\s+", body.strip()):
if current and len(" ".join(current + [sentence]).split()) > max_words:
chunks.append(f"{doc_title} > {title}: " + " ".join(current)); current = []
current.append(sentence)
if current:
chunks.append(f"{doc_title} > {title}: " + " ".join(current))
return chunkssection_chunks never cuts a sentence, never mixes two sections, and prefixes every chunk with its document and section title. That prefix is cheap and powerful: a chunk that says “These renew once a year” is meaningless on its own, but “Billing policy > Annual plans: These renew once a year” is searchable. Generating a short, chunk-specific context with an LLM at ingestion time takes the same idea further.
Step-by-step walkthrough
Step 1: Index the chunks with BM25
class BM25:
def __init__(self, docs, k1=1.5, b=0.75):
self.docs = [words(d) for d in docs]
self.avg = sum(map(len, self.docs)) / len(self.docs)
self.df = Counter(t for d in self.docs for t in set(d))
self.k1, self.b, self.n = k1, b, len(docs)
def score(self, query, i):
doc, tf = self.docs[i], Counter(self.docs[i])
total = 0.0
for t in words(query):
if t not in tf:
continue
idf = math.log(1 + (self.n - self.df[t] + 0.5) / (self.df[t] + 0.5))
total += idf * tf[t] * (self.k1 + 1) / (tf[t] + self.k1 * (1 - self.b + self.b * len(doc) / self.avg))
return total
def top(self, query, k):
return sorted(range(self.n), key=lambda i: self.score(query, i), reverse=True)[:k]BM25 rewards query terms that are rare across the corpus (idf), with diminishing returns for repeated terms (k1) and a penalty for long chunks (b). It is the keyword half of hybrid retrieval and catches exact terms such as product names and error codes that embeddings blur.
Step 2: Retrieve with each chunking strategy
question = "How long do I have to get a refund on an annual plan?"
strategies = [("fixed-30", fixed_chunks(POLICY, 30)),
("fixed-30-overlap-10", fixed_chunks(POLICY, 30, 10)),
("sections", section_chunks(POLICY))]
for name, chunks in strategies:
best = chunks[BM25(chunks).top(question, 1)[0]]
print(name, len(chunks), "chunks; top-1 contains '14 days':", "14 days" in best)
# fixed-30 3 chunks; top-1 contains '14 days': False
# fixed-30-overlap-10 4 chunks; top-1 contains '14 days': False
# sections 2 chunks; top-1 contains '14 days': TrueStep 3: Build a grounded prompt
def build_prompt(question, chunks):
sources = "\n".join(f"[{i + 1}] {c}" for i, c in enumerate(chunks))
return (
"Answer using only the sources below. Cite sources like [1]. "
"If the sources do not contain the answer, say you don't know.\n\n"
f"Sources:\n{sources}\n\nQuestion: {question}"
)
print(build_prompt(question, [section_chunks(POLICY)[0]]))The instruction does three jobs: it limits the model to the sources, it makes every claim traceable through a citation number, and it gives the model a legitimate way out. Without that last clause, a model asked about something absent from the sources tends to answer from its training data, which is exactly the hallucination RAG was supposed to prevent. Put the question after the sources, and keep instructions in the system prompt when your API has one.
Step 4: Decide how many chunks to pass
Retrieval returns a ranked list; the prompt gets the top few. Too few and the answer may be missing; too many and you pay for tokens, add latency and bury the relevant passage among distractors. Start with 3 to 8 reranked chunks, then tune using the evaluation in the next sections rather than intuition.
Worked scenario
The billing assistant above shipped with fixed 30-word chunks, the default in the team’s first prototype. Customers asking “How long do I have to get a refund on an annual plan?” received answers like “You’ll get a reminder seven days before renewal” or “I don’t know”. The source document clearly says 14 days.
Printing the chunks showed the cause. The fixed splitter produced:
- “# Billing policy ## Annual plans These renew once a year on the purchase date. We send a reminder email seven days before each renewal. To cancel, open Settings and”
- “choose Billing. A full refund is available if you cancel within 14 days of a renewal. ## Monthly plans These renew every month. …”
The 14-day rule lives in a chunk that starts mid-sentence, has no “annual” in it, and is glued to the monthly plan section. The first chunk matches “annual plan” strongly, so it wins retrieval, and it contains everything except the answer. Adding a 10-word overlap did not help, because the problem is not a sentence cut at a boundary; it is that the answer was separated from the words that identify its subject.
Switching to section_chunks produced one chunk per section, each prefixed with “Billing policy > Annual plans” or “Billing policy > Monthly plans”. The annual chunk now contains both the subject and the 14-day rule, and the model answers “within 14 days of a renewal [1]”. The team then added the regression test below so a future change to chunk size could not reintroduce the bug.
Common mistake
- Blaming the model first. When a RAG answer is wrong, inspect the retrieved chunks before touching the prompt. If the answer is not in them, no prompt will fix it.
- One global chunk size. FAQs, API references, contracts and chat logs have different structures; chunk each by its structure.
- Dropping metadata. Without source, section and permission metadata you cannot cite, filter or enforce access control.
- Vector search only. Exact identifiers and rare terms need keyword retrieval; use hybrid search.
- No abstain path. The prompt must say what to do when the sources do not contain the answer, and your evaluation must test it.
- Stale indexes. Deleted or updated documents must be removed or re-embedded; schedule incremental re-indexing and track document versions.
Verify the behavior
Retrieval deserves its own tests, separate from answer quality. A small golden set of questions with the chunk that must be retrieved is enough to start:
GOLDEN = [
("How long do I have to get a refund on an annual plan?", "14 days"),
("Do I get money back for half a month?", "No refunds"),
]
def test_retrieval_finds_the_answer_chunk():
chunks = section_chunks(POLICY)
index = BM25(chunks)
for question, must_contain in GOLDEN:
top = [chunks[i] for i in index.top(question, 2)]
assert any(must_contain in c for c in top), question
test_retrieval_finds_the_answer_chunk(); print("ok")At scale, track recall@k and mean reciprocal rank over a few hundred real questions, and evaluate answers separately for faithfulness to the sources and for correctness.
Follow-up questions
How would you handle a question that needs facts from two documents? Retrieve enough candidates to cover both, or decompose the question into sub-queries, retrieve for each and combine. Agentic retrieval, where the model issues its own searches, helps here but costs latency.
What is query rewriting? Turning a conversational follow-up such as “what about monthly?” into a standalone query (“refund policy for monthly plans”) using the conversation history, so retrieval sees the full intent.
How do you enforce permissions? Store access-control lists as chunk metadata and filter at query time with the user’s identity, inside the retrieval call. Never rely on the prompt to hide documents.
When is long context better than RAG? When the corpus is small and stable enough to fit in the window, sending everything (ideally with prompt caching) can be simpler and more accurate. RAG wins on large, changing or permissioned corpora.
Interview exercise
Your RAG assistant over 50,000 engineering wiki pages has good answers in testing, but in production 30 percent of answers are rated unhelpful. Logs show retrieval returns chunks from pages last edited in 2019 that describe deprecated systems, while newer pages exist. How would you diagnose and fix this?
Answer and reasoning
First confirm it is a retrieval problem: sample the unhelpful answers and check whether the current page was in the top-k candidates at all, or present but outranked. If it was outranked, older pages probably match better because they are longer and use the exact old terminology. Fixes, in order: add last_modified and an “archived” flag to metadata and filter out archived content; add a recency signal in reranking for content where freshness matters; deduplicate near-identical pages so stale copies do not crowd out the current one; and give the owners a process to archive deprecated pages, because no ranking trick beats removing wrong content. Then add these failing questions to the golden set so the improvement is measured and cannot regress. The reasoning: diagnose by stage, fix the data before tuning the model, and turn every incident into an evaluation case.