Embeddings and Vector Search: AI Engineer Interview Guide
What embeddings are, how cosine similarity and approximate nearest neighbour indexes work, and when to add keyword search and reranking.
AI engineering interviews: how LLMs work, prompting, embeddings and vector search, RAG design, tool use and agents, evaluation, cost, latency and safety.
Official reference: Anthropic: Building effective agents
What embeddings are, how cosine similarity and approximate nearest neighbour indexes work, and when to add keyword search and reranking.
Build an eval suite for an LLM app: golden datasets, code-based checks, LLM-as-judge calibration, retrieval metrics and regression gates in CI.
Fine-tuning changes how a model behaves; RAG changes what it knows at request time. A decision guide with data prep, costs and failure cases.
How LLM tool use works: schemas, the call-execute-return loop, validation, error results, step limits and when an agent beats a workflow.
Direct and indirect prompt injection, why prompts alone cannot stop it, and layered defenses: tool policy, confirmation and output filtering.
A step-by-step RAG design for interviews: ingestion, structure-aware chunking, hybrid retrieval, grounded prompts with citations and evaluation.
Why LLMs hallucinate and the layered fixes that work: grounding, permission to abstain, citations, structured output and automatic claim checks.
How LLMs turn text into tokens, why the context window is a shared budget, and what temperature, top-k and top-p actually do to the next token.
A token is the unit a model actually reads and writes: usually a subword piece produced by a tokenizer such as byte-pair encoding. Common words are often one token, rare words and identifiers split into several, and code, numbers and non-English text usually need more tokens per character.
The model's vocabulary, its input and its output are all sequences of token ids, so the work it does (and the memory it needs) scales with tokens. That is why context windows, rate limits and prices are measured in tokens.
Practical consequences:
Likely follow-up: Why do models struggle to count the letters in a word? · Why might the same prompt cost more in Japanese than in English?
The context window is the maximum number of tokens the model can attend to in one request. Everything counts: the system prompt, tool definitions, conversation history, retrieved documents and the tokens the model generates. Many APIs also cap output length separately.
When a request would exceed the window, the API rejects it, or your application must drop or compress content. Naive truncation from the front silently removes the system prompt or early facts, which shows up as the model "forgetting" its instructions.
I manage it as a budget: pin the system prompt, keep recent turns verbatim, summarize older turns, retrieve only the most relevant chunks, and reserve room for the output. Even within the limit, long contexts cost more, add latency and can bury relevant facts in the middle, so bigger windows do not remove the need to choose what goes in.
Likely follow-up: How would you implement conversation summarization? · Is a 1M-token window a replacement for RAG?
At each step the model produces a score for every token; softmax turns the scores into probabilities and the decoder picks one. Temperature divides the scores before softmax: below 1 sharpens the distribution toward the top token, above 1 flattens it so unlikely tokens get chosen more often. Top-p (nucleus sampling) keeps only the smallest set of tokens whose probabilities sum to p, cutting off the tail; top-k keeps a fixed number.
For extraction, classification or code where there is one right answer, I use a low temperature or the model's default if it does not accept sampling parameters. For brainstorming or varied copy, a higher temperature gives diversity. I usually tune one of temperature or top-p, not both.
Temperature 0 is close to greedy decoding but not a determinism guarantee: serving infrastructure can still produce different outputs.
Likely follow-up: Why might two calls at temperature 0 return different text? · Why do some reasoning models reject sampling parameters?
An LLM is a transformer trained to predict the next token. Generation is a loop: the prompt is tokenized, the model processes all prompt tokens in parallel and produces a probability distribution for the next token, the decoder picks one, appends it, and the model runs again for the following token until it emits a stop token, hits a stop sequence or reaches the output limit.
Self-attention lets each position weigh every earlier token, which is how the model uses context. A KV cache stores intermediate results for tokens already processed so each new step only computes the new token.
This explains the operational facts: input is processed quickly in one pass, output is produced one token at a time, so latency is roughly time to first token plus output length divided by generation speed, and output tokens are usually priced higher. The model is predicting plausible text, not looking up facts.
Likely follow-up: What is time to first token and what affects it? · Why does a longer prompt increase memory use on the server?
A good prompt reads like a clear brief to a capable new colleague. It states the task and audience, gives the context the model needs, defines the output format exactly, says what to do in edge cases (including "I don't know"), and separates instructions from data with clear structure such as headings or labelled sections. Stable instructions go in the system prompt; per-request data goes in the user message.
Few-shot examples help when the format or judgement is easier to show than to describe, such as a labelling scheme or a house style. They should be diverse, cover tricky cases and match the real input distribution, because the model copies patterns, including accidental ones.
Most importantly, prompts are code: version them, test them against an eval set, and change one thing at a time.
Likely follow-up: How do few-shot examples cause unwanted bias? · Where should retrieved documents go relative to the question?
Chain-of-thought means having the model work through intermediate steps before the final answer. With standard models you can prompt for it ("think step by step", or a scratchpad section), which improves multi-step maths, logic and planning, at the cost of more output tokens and latency.
Reasoning models are trained to do this internally: they generate reasoning tokens before answering, often controlled by an effort or budget setting rather than by the prompt. That changes practice:
I measure whether the extra reasoning actually improves my eval before paying for it.
Likely follow-up: When would reasoning hurt a product? · How would you cap cost on a reasoning-heavy route?
An embedding model maps text (or images) to a fixed-length vector so that inputs with similar meaning are close together, typically measured by cosine similarity. Unlike keyword search, "my card was charged twice" can match "duplicate payment" even with no shared words.
For semantic search, I chunk documents, embed each chunk once with the same model, and store vectors with metadata. At query time I embed the query with the same model and return the nearest vectors, usually through an approximate nearest neighbour index.
Points I watch:
Likely follow-up: Why normalize vectors before storing them? · How do you choose an embedding model?
Exact search compares the query with every vector, which is linear in the corpus size: fine for thousands of vectors, too slow and costly at tens of millions with high query rates. ANN indexes inspect a small fraction and accept occasionally missing a true neighbour.
HNSW builds a layered proximity graph and searches by greedily walking toward the query; ef_search trades recall for latency. It has excellent recall and speed but keeps the graph and vectors in memory and builds slowly. IVF clusters vectors with k-means and scans only the n_probe nearest clusters; it is cheaper in memory, especially combined with product quantization, at some recall cost.
The key is measuring: compute exact neighbours for a sample of real queries offline, then tune parameters until recall@k meets the target at acceptable p99 latency. Recall depends heavily on how clustered your data is.
Likely follow-up: What happens to recall when you add a strict metadata filter? · How would you reduce memory for 100 million vectors?
Hybrid search runs a keyword retriever (usually BM25) and a vector retriever on the same query and merges the results. They fail differently: embeddings capture meaning and paraphrase but blur exact strings such as part numbers, error codes and names, while BM25 nails exact terms but misses synonyms.
Scores from the two systems are on different scales, so the usual merge is reciprocal rank fusion: each document scores the sum of 1 / (k + rank) over the lists it appears in, with k around 60. It needs no calibration and rewards documents both systems rank highly. Weighted score blending is an alternative once scores are normalized and tuned.
After fusion I typically rerank the top 20 to 50 candidates with a cross-encoder and pass the best few to the model. I validate the gain with recall@k on a golden query set.
Likely follow-up: Why not just add the two scores? · When would keyword search alone be enough?
Chunks should be small enough to be specific and to fit several in the prompt, but complete enough to contain an answer together with the words that identify its subject. Fixed-size splitting every N tokens is easy but cuts sentences, separates facts from their headings and mixes sections.
I prefer structure-aware chunking: split on headings, then paragraphs, then sentences, with a size cap; keep tables and code blocks intact; and prefix each chunk with its document and section title so a chunk like "These renew yearly" becomes searchable. Overlap helps with sentences cut at boundaries but does not fix lost context. Adding a short generated summary of where the chunk sits in the document improves retrieval further.
Different sources need different strategies: FAQs by question, API docs by endpoint, contracts by clause. I choose by measuring recall@k on real questions, not by a default size.
Likely follow-up: When does overlap help and when does it not? · How would you chunk a 300-page PDF with tables?
RAG has two pipelines. Ingestion runs offline: load sources, parse them to clean text that keeps structure, chunk, attach metadata (source, section, permissions, version, date), embed, and write to a vector index and a keyword index. It must handle updates and deletions incrementally.
Answering runs per request: rewrite the question into a standalone query using conversation history, retrieve candidates from both indexes filtered by the user's permissions, fuse and rerank, then build a prompt that numbers the top chunks, tells the model to answer only from them, cite them, and say when the answer is not there. The response is returned with source links, and logged with the retrieved ids.
Around it I add evals for retrieval and answers, caching, monitoring of unanswered questions, and feedback capture. RAG keeps knowledge outside the model, so updates take effect on re-index without retraining.
Likely follow-up: Where do you enforce document permissions? · How do you handle a document that is deleted?
First-stage retrievers, both embeddings and BM25, score the query and each document independently, which is fast enough to search millions of chunks but loses nuance. A reranker, usually a cross-encoder, reads the query and one candidate together and outputs a relevance score, so it can judge whether the passage actually answers the question.
Because it runs a model per candidate, it is too slow for the whole corpus, so it only reorders the top 20 to 100 candidates from the first stage, and the best few go to the LLM. The cost is added latency (tens to hundreds of milliseconds depending on model and candidate count) and compute.
The gain is usually largest when the first stage returns the right chunk but not at the top, which shows up as good recall@50 but poor recall@5. I confirm with metrics before adding it.
Likely follow-up: Could you use an LLM as the reranker? · How many candidates would you rerank?
I evaluate the stages separately, because each fails differently.
I run the suite in CI on every change to chunking, embeddings, prompts or models, compare per case with the baseline, and check error bars before trusting small differences. In production I add sampled human review, user feedback and monitoring of low-retrieval-score queries, then feed failures back into the golden set.
Likely follow-up: How do you build the golden set when you have no labelled data? · How would you detect retrieval drift after a re-index?
I trace it stage by stage using the logged request.
Then I fix the earliest failing stage, add the case to the eval set, and check that the fix does not regress other cases.
Likely follow-up: What would you log per request to make this possible? · How do you handle two sources that contradict each other?
You send the model tool definitions: a name, a description of when to use it, and a JSON Schema for the arguments. When the model decides a tool is needed, it returns a structured tool call (name, arguments, call id) instead of plain text, and the API signals this with a stop or finish reason.
Your code executes it. The model never runs anything; it only proposes. So the application validates arguments against the schema and business rules (does this user own this order?), executes the function, and returns a tool result linked to the call id, including an explicit error result if it failed. Then you call the model again with the result and it continues or answers.
A model may request several independent calls at once; I run them concurrently and return all results together. Strict schema modes reduce malformed arguments but do not replace authorization checks.
Likely follow-up: What makes a good tool description? · How do you handle a tool that takes 30 seconds?
In a workflow, your code defines the path: prompt chaining, routing to specialized prompts, parallel calls whose results are combined, or an evaluator that checks and requests a revision. The model fills in steps, but the control flow is fixed and testable.
In an agent, the model directs the process: it decides which tools to call, in what order and when it is done, in a loop until it answers or hits a limit.
I start with the simplest thing that works: a single well-prompted call, then a workflow. An agent is justified when the steps genuinely cannot be known in advance (open-ended research, debugging, multi-step tasks over a changing environment), the task is valuable enough to absorb higher latency and cost, and errors can be caught or reversed. Agents need step limits, evals of whole trajectories and stronger guardrails.
Likely follow-up: Give an example task where an agent clearly beats a workflow. · How would you evaluate an agent's trajectory?
I assume the model will sometimes be wrong or manipulated, and design so that mistakes are contained.
These turn "the model must behave" into "the system stays safe when it does not".
Likely follow-up: How would you stop an agent from looping forever? · Where would you put the approval step in the loop?
In order of strength:
Even with constraints, I validate in code with a schema library such as Pydantic, because the schema cannot express every rule: value ranges, cross-field consistency, citations that point to real sources. I also check the stop reason, since an output truncated at the token limit is invalid however good the schema.
Structured output fixes shape, not truth: a valid field can still contain a wrong value. On validation failure I retry once with the error message, then fall back. Keeping schemas small and flat with clear field descriptions also improves quality.
Likely follow-up: How would you handle optional fields the model invents? · Why might constrained decoding reduce answer quality?
Models generate the most plausible continuation, not a retrieved fact. When the information is missing, ambiguous or buried, a fluent guess is often more likely than "I don't know", and training tends to reward confident answers.
I reduce it in layers:
No single technique removes it; temperature 0 and RAG alone are not enough.
Likely follow-up: How do you verify a citation automatically? · Would fine-tuning fix hallucinations?
An eval needs three parts: a fixed dataset, a way to run the real system on it, and graders.
Because outputs vary, I run several trials per case for important suites and look at confidence intervals before treating a small change as real.
Likely follow-up: How many cases are enough? · How would you evaluate a multi-turn conversation?
A judge model is useful for qualities code cannot check, such as whether an answer addresses the question, is faithful to sources in paraphrase, or follows a tone guideline. It scales far better than human review.
But judges have known biases: preference for longer and more confident answers, position bias in pairwise comparisons, and preference for text in their own style. To make one trustworthy:
I keep code-based checks for anything objective and treat the judge as one signal among several.
Likely follow-up: Should the judge be the same model as the one being evaluated? · How do you measure judge agreement?
First I measure where time goes: time to first token (queueing plus prompt processing), generation time (output tokens divided by speed), and everything around the model such as retrieval, reranking and tool calls.
Then, roughly in order:
I verify each change against both a latency percentile and the quality eval.
Likely follow-up: Why is output length usually the biggest lever? · How would you hide latency in a chat UI?
I start with a token profile: input versus output tokens per request, by route, and how many calls one user task takes. Then the levers, cheapest first:
I judge cost per completed task, not per request, because a cheaper call that needs retries can cost more. Every change is checked against the quality eval before rollout.
Likely follow-up: What breaks prompt caching silently? · When is a batch API the wrong choice?
Providers can store the processed state of a prompt prefix so that later requests starting with the same prefix skip reprocessing it. Cache hits are billed at a steep discount and reduce time to first token. Depending on the provider, caching is automatic above a minimum length or enabled with explicit markers, and entries expire after a short idle period unless refreshed.
It is an exact prefix match, so any change early in the prompt invalidates everything after it. Design accordingly:
I verify with the usage fields the API returns for cached tokens; a hit rate of zero usually means a hidden change in the prefix.
Likely follow-up: How is this different from caching whole responses? · Why can a timestamp in the system prompt be expensive?
Provider limits usually apply to requests per minute and tokens per minute, and exceeding them returns HTTP 429, often with a Retry-After header. Server errors and overload responses also happen.
My approach:
Retry-After, with a capped number of attempts; do not retry 400-class validation errorsLikely follow-up: Why add jitter to backoff? · How would you share a token-per-minute limit across many workers?
Prompt injection is input that makes the model follow an attacker's instructions instead of yours. Direct injection comes from the user. Indirect injection hides instructions in content the model reads: web pages, emails, documents, tool results. It works because instructions and data share one channel, and there is no escaping function for natural language.
So I design for a model that may be compromised:
The dangerous combination is private data, untrusted content and an external channel together.
Likely follow-up: Why is a stronger system prompt not a fix? · How can a Markdown image leak data?
The common ones:
I treat the model as a component that may repeat anything it was shown.
Likely follow-up: How would you test for cross-tenant leakage? · What is PII redaction and where would you apply it?
They solve different problems. RAG changes what the model knows at request time: right for large or frequently changing knowledge, citations and per-user permissions. Fine-tuning changes how the model behaves: right for a strict output format, a house style, a narrow classification task, shorter prompts, or making a smaller model match a larger one on one task.
My order is: best prompt plus few-shot examples, measured on an eval; retrieval for knowledge gaps; fine-tuning only when evals show a behaviour gap prompting cannot close, or when the cost and latency of long prompts justify it.
Fine-tuning to add facts is the classic mistake: recall is imprecise, there are no citations, updates mean retraining and you cannot delete data from the weights. It also has upkeep: curated data, held-out evals and retraining when the base model changes. The two combine well.
Likely follow-up: How would you prove a fine-tune was worth it? · What is distillation?
Full fine-tuning updates every weight, which needs memory for weights, gradients and optimizer state for the whole model, and produces a full-size copy per task. LoRA (low-rank adaptation) freezes the base model and learns small low-rank matrices added to selected weight matrices, typically in the attention layers. Only those adapters are trained and stored.
Benefits:
QLoRA trains LoRA adapters on top of a quantized base model, so larger models fit on smaller GPUs. The trade-offs are hyperparameters to tune (rank, which layers, learning rate) and sometimes slightly lower quality than full fine-tuning on hard tasks. As with any tuning, a held-out eval decides.
Likely follow-up: What does the rank hyperparameter control? · Can you merge a LoRA adapter into the base weights?
The API is stateless: the model only knows what you send each time, so memory is something the application builds.
Within a conversation:
Across sessions:
Risks to manage: summaries lose details, stale memories contradict new information (store timestamps and let new facts override), and memory is personal data, so it needs access control and retention rules. I evaluate with long scripted conversations that check whether key facts survive.
Likely follow-up: How would you test that summarization keeps important facts? · What should never be stored as memory?
No questions match that filter.
Prefer multiple choice? All 20 AI & LLM Engineering MCQs with answers →