Ch. 30

AI & LLM Engineering interview questions & answers

AI engineering interviews: how LLMs work, prompting, embeddings and vector search, RAG design, tool use and agents, evaluation, cost, latency and safety.

30 interview questions20 quiz questions8 notes
your progress0%

Notes in this chapter

Filter all notes →
read ✓AI & LLM Engineering · hard

Fine-Tuning vs RAG: When to Use Each in LLM Apps

Fine-tuning changes how a model behaves; RAG changes what it knows at request time. A decision guide with data prep, costs and failure cases.

~9 min readread →
read ✓AI & LLM Engineering · mid

LLM Function Calling and Agent Loops Explained

How LLM tool use works: schemas, the call-execute-return loop, validation, error results, step limits and when an agent beats a workflow.

~8 min readread →
read ✓AI & LLM Engineering · mid

How to Reduce LLM Hallucinations in Production

Why LLMs hallucinate and the layered fixes that work: grounding, permission to abstain, citations, structured output and automatic claim checks.

~8 min readread →

30 AI & LLM Engineering interview questions study by subtopic

30 questions
  1. 1.What is a token, and why do LLM limits and prices use tokens instead of words or characters?easy

    A token is the unit a model actually reads and writes: usually a subword piece produced by a tokenizer such as byte-pair encoding. Common words are often one token, rare words and identifiers split into several, and code, numbers and non-English text usually need more tokens per character.

    The model's vocabulary, its input and its output are all sequences of token ids, so the work it does (and the memory it needs) scales with tokens. That is why context windows, rate limits and prices are measured in tokens.

    Practical consequences:

    • the same text costs a different number of tokens on different models
    • a characters-divided-by-four estimate is only a rough guide for English
    • count with the provider's tokenizer or token-counting endpoint before enforcing limits
    What interviewers listen for
    • Subword units from a trained tokenizer
    • Token count depends on the model's tokenizer
    • Limits, latency and price scale with tokens
    • Count with the real tokenizer, not a heuristic

    Likely follow-up: Why do models struggle to count the letters in a word? · Why might the same prompt cost more in Japanese than in English?

  2. 2.What is a context window, and what happens when a conversation grows beyond it?easy

    The context window is the maximum number of tokens the model can attend to in one request. Everything counts: the system prompt, tool definitions, conversation history, retrieved documents and the tokens the model generates. Many APIs also cap output length separately.

    When a request would exceed the window, the API rejects it, or your application must drop or compress content. Naive truncation from the front silently removes the system prompt or early facts, which shows up as the model "forgetting" its instructions.

    I manage it as a budget: pin the system prompt, keep recent turns verbatim, summarize older turns, retrieve only the most relevant chunks, and reserve room for the output. Even within the limit, long contexts cost more, add latency and can bury relevant facts in the middle, so bigger windows do not remove the need to choose what goes in.

    What interviewers listen for
    • Input and output share the window
    • Overflow means rejection or truncation you control
    • Treat it as a token budget
    • Long context still costs money and attention

    Likely follow-up: How would you implement conversation summarization? · Is a 1M-token window a replacement for RAG?

  3. 3.What do temperature and top-p control, and what values would you use for extraction versus brainstorming?easy

    At each step the model produces a score for every token; softmax turns the scores into probabilities and the decoder picks one. Temperature divides the scores before softmax: below 1 sharpens the distribution toward the top token, above 1 flattens it so unlikely tokens get chosen more often. Top-p (nucleus sampling) keeps only the smallest set of tokens whose probabilities sum to p, cutting off the tail; top-k keeps a fixed number.

    For extraction, classification or code where there is one right answer, I use a low temperature or the model's default if it does not accept sampling parameters. For brainstorming or varied copy, a higher temperature gives diversity. I usually tune one of temperature or top-p, not both.

    Temperature 0 is close to greedy decoding but not a determinism guarantee: serving infrastructure can still produce different outputs.

    What interviewers listen for
    • Temperature rescales the distribution, ranking unchanged
    • Top-p and top-k truncate the tail
    • Low for single-answer tasks, higher for variety
    • Temperature 0 is not guaranteed deterministic

    Likely follow-up: Why might two calls at temperature 0 return different text? · Why do some reasoning models reject sampling parameters?

  4. 4.At an engineer's level, how does an LLM generate a response?easy

    An LLM is a transformer trained to predict the next token. Generation is a loop: the prompt is tokenized, the model processes all prompt tokens in parallel and produces a probability distribution for the next token, the decoder picks one, appends it, and the model runs again for the following token until it emits a stop token, hits a stop sequence or reaches the output limit.

    Self-attention lets each position weigh every earlier token, which is how the model uses context. A KV cache stores intermediate results for tokens already processed so each new step only computes the new token.

    This explains the operational facts: input is processed quickly in one pass, output is produced one token at a time, so latency is roughly time to first token plus output length divided by generation speed, and output tokens are usually priced higher. The model is predicting plausible text, not looking up facts.

    What interviewers listen for
    • Next-token prediction in a loop
    • Prompt processed in parallel, output sequentially
    • KV cache avoids recomputation
    • Plausible text, not database lookup

    Likely follow-up: What is time to first token and what affects it? · Why does a longer prompt increase memory use on the server?

  5. 5.What makes a good production prompt, and when do you add few-shot examples?easy

    A good prompt reads like a clear brief to a capable new colleague. It states the task and audience, gives the context the model needs, defines the output format exactly, says what to do in edge cases (including "I don't know"), and separates instructions from data with clear structure such as headings or labelled sections. Stable instructions go in the system prompt; per-request data goes in the user message.

    Few-shot examples help when the format or judgement is easier to show than to describe, such as a labelling scheme or a house style. They should be diverse, cover tricky cases and match the real input distribution, because the model copies patterns, including accidental ones.

    Most importantly, prompts are code: version them, test them against an eval set, and change one thing at a time.

    What interviewers listen for
    • Clear task, context, format and edge cases
    • Separate instructions from data
    • Few-shot when showing beats telling
    • Version and evaluate prompts like code

    Likely follow-up: How do few-shot examples cause unwanted bias? · Where should retrieved documents go relative to the question?

  6. 6.What is chain-of-thought, and how do reasoning models change how you prompt and budget?mid

    Chain-of-thought means having the model work through intermediate steps before the final answer. With standard models you can prompt for it ("think step by step", or a scratchpad section), which improves multi-step maths, logic and planning, at the cost of more output tokens and latency.

    Reasoning models are trained to do this internally: they generate reasoning tokens before answering, often controlled by an effort or budget setting rather than by the prompt. That changes practice:

    • give the goal and constraints clearly instead of scripting each step
    • expect higher and more variable latency and cost, since reasoning tokens are usually billed as output
    • tune effort per route: high for hard problems, low for simple classification
    • do not parse or show internal reasoning as if it were a guaranteed explanation

    I measure whether the extra reasoning actually improves my eval before paying for it.

    What interviewers listen for
    • Intermediate steps improve multi-step tasks
    • Reasoning models think before answering
    • Reasoning tokens cost money and time
    • Tune effort per task and measure

    Likely follow-up: When would reasoning hurt a product? · How would you cap cost on a reasoning-heavy route?

  7. 7.What are embeddings, and how are they used for semantic search?easy

    An embedding model maps text (or images) to a fixed-length vector so that inputs with similar meaning are close together, typically measured by cosine similarity. Unlike keyword search, "my card was charged twice" can match "duplicate payment" even with no shared words.

    For semantic search, I chunk documents, embed each chunk once with the same model, and store vectors with metadata. At query time I embed the query with the same model and return the nearest vectors, usually through an approximate nearest neighbour index.

    Points I watch:

    • query and document vectors must come from the same model and version; changing models means re-embedding
    • some models expect different prefixes or task types for queries and passages
    • similarity scores are not probabilities and their scale differs by model
    • embeddings blur exact identifiers, so I usually add keyword search
    What interviewers listen for
    • Meaning becomes geometry
    • Same model for queries and documents
    • Nearest-neighbour search over chunk vectors
    • Weak on exact strings and identifiers

    Likely follow-up: Why normalize vectors before storing them? · How do you choose an embedding model?

  8. 8.Why do vector databases use approximate nearest neighbour indexes, and how do HNSW and IVF differ?mid

    Exact search compares the query with every vector, which is linear in the corpus size: fine for thousands of vectors, too slow and costly at tens of millions with high query rates. ANN indexes inspect a small fraction and accept occasionally missing a true neighbour.

    HNSW builds a layered proximity graph and searches by greedily walking toward the query; ef_search trades recall for latency. It has excellent recall and speed but keeps the graph and vectors in memory and builds slowly. IVF clusters vectors with k-means and scans only the n_probe nearest clusters; it is cheaper in memory, especially combined with product quantization, at some recall cost.

    The key is measuring: compute exact neighbours for a sample of real queries offline, then tune parameters until recall@k meets the target at acceptable p99 latency. Recall depends heavily on how clustered your data is.

    What interviewers listen for
    • Exact search is linear in corpus size
    • HNSW: graph, fast, memory-heavy
    • IVF: clusters plus n_probe, often with PQ
    • Measure recall against exact search on your data

    Likely follow-up: What happens to recall when you add a strict metadata filter? · How would you reduce memory for 100 million vectors?

  9. 10.How do you choose a chunking strategy for RAG?mid

    Chunks should be small enough to be specific and to fit several in the prompt, but complete enough to contain an answer together with the words that identify its subject. Fixed-size splitting every N tokens is easy but cuts sentences, separates facts from their headings and mixes sections.

    I prefer structure-aware chunking: split on headings, then paragraphs, then sentences, with a size cap; keep tables and code blocks intact; and prefix each chunk with its document and section title so a chunk like "These renew yearly" becomes searchable. Overlap helps with sentences cut at boundaries but does not fix lost context. Adding a short generated summary of where the chunk sits in the document improves retrieval further.

    Different sources need different strategies: FAQs by question, API docs by endpoint, contracts by clause. I choose by measuring recall@k on real questions, not by a default size.

    What interviewers listen for
    • Specific but self-contained chunks
    • Split along document structure
    • Keep headings and context with the text
    • Choose by retrieval metrics per source type

    Likely follow-up: When does overlap help and when does it not? · How would you chunk a 300-page PDF with tables?

  10. 11.Walk me through the architecture of a retrieval-augmented generation system.mid

    RAG has two pipelines. Ingestion runs offline: load sources, parse them to clean text that keeps structure, chunk, attach metadata (source, section, permissions, version, date), embed, and write to a vector index and a keyword index. It must handle updates and deletions incrementally.

    Answering runs per request: rewrite the question into a standalone query using conversation history, retrieve candidates from both indexes filtered by the user's permissions, fuse and rerank, then build a prompt that numbers the top chunks, tells the model to answer only from them, cite them, and say when the answer is not there. The response is returned with source links, and logged with the retrieved ids.

    Around it I add evals for retrieval and answers, caching, monitoring of unanswered questions, and feedback capture. RAG keeps knowledge outside the model, so updates take effect on re-index without retraining.

    What interviewers listen for
    • Offline ingestion and online answering
    • Metadata for citations and permissions
    • Hybrid retrieval, rerank, grounded prompt
    • Evals and monitoring around both stages

    Likely follow-up: Where do you enforce document permissions? · How do you handle a document that is deleted?

  11. 12.What does a reranker add to a retrieval pipeline, and what does it cost?mid

    First-stage retrievers, both embeddings and BM25, score the query and each document independently, which is fast enough to search millions of chunks but loses nuance. A reranker, usually a cross-encoder, reads the query and one candidate together and outputs a relevance score, so it can judge whether the passage actually answers the question.

    Because it runs a model per candidate, it is too slow for the whole corpus, so it only reorders the top 20 to 100 candidates from the first stage, and the best few go to the LLM. The cost is added latency (tens to hundreds of milliseconds depending on model and candidate count) and compute.

    The gain is usually largest when the first stage returns the right chunk but not at the top, which shows up as good recall@50 but poor recall@5. I confirm with metrics before adding it.

    What interviewers listen for
    • Cross-encoder scores query and passage jointly
    • Only reorders the top candidates
    • Adds latency per candidate
    • Helps when recall@k is high but precision is low

    Likely follow-up: Could you use an LLM as the reranker? · How many candidates would you rerank?

  12. 13.How would you evaluate a RAG system end to end?hard

    I evaluate the stages separately, because each fails differently.

    • Retrieval: a golden set of real questions labelled with the chunks that contain the answer; measure recall@k and mean reciprocal rank. If the answer is not retrieved, nothing downstream can fix it.
    • Faithfulness: is every claim in the answer supported by the retrieved context? Check numbers and quotes with code, paraphrases with a calibrated judge.
    • Correctness and relevance: does it answer the question, judged against reference answers or a rubric.
    • Abstention: unanswerable questions must produce "I don't know", not a guess.

    I run the suite in CI on every change to chunking, embeddings, prompts or models, compare per case with the baseline, and check error bars before trusting small differences. In production I add sampled human review, user feedback and monitoring of low-retrieval-score queries, then feed failures back into the golden set.

    What interviewers listen for
    • Separate retrieval from generation metrics
    • Recall@k and MRR for retrieval
    • Faithfulness and abstention for answers
    • CI gate plus production feedback loop

    Likely follow-up: How do you build the golden set when you have no labelled data? · How would you detect retrieval drift after a re-index?

  13. 14.A user reports a wrong answer from your RAG assistant. How do you debug it?hard

    I trace it stage by stage using the logged request.

    • Was the right content in the corpus? If the document is missing, stale or duplicated by an outdated copy, it is a data problem.
    • Was it retrieved? Check the retrieved chunk ids and scores. If the answer chunk is absent, look at the query rewrite, chunking (was the fact separated from its subject?), keyword versus vector behaviour and metadata filters.
    • Was it in the prompt? It may have been retrieved but cut by the top-k or token budget.
    • Did the model use it? If the answer chunk was in the prompt and the answer is still wrong, it is a generation problem: prompt clarity, conflicting sources, buried position or the model guessing instead of abstaining.

    Then I fix the earliest failing stage, add the case to the eval set, and check that the fix does not regress other cases.

    What interviewers listen for
    • Use logs to trace one request
    • Corpus, retrieval, prompt, generation in order
    • Fix the earliest failing stage
    • Turn the incident into an eval case

    Likely follow-up: What would you log per request to make this possible? · How do you handle two sources that contradict each other?

  14. 15.How does function calling (tool use) work, and who actually executes the function?mid

    You send the model tool definitions: a name, a description of when to use it, and a JSON Schema for the arguments. When the model decides a tool is needed, it returns a structured tool call (name, arguments, call id) instead of plain text, and the API signals this with a stop or finish reason.

    Your code executes it. The model never runs anything; it only proposes. So the application validates arguments against the schema and business rules (does this user own this order?), executes the function, and returns a tool result linked to the call id, including an explicit error result if it failed. Then you call the model again with the result and it continues or answers.

    A model may request several independent calls at once; I run them concurrently and return all results together. Strict schema modes reduce malformed arguments but do not replace authorization checks.

    What interviewers listen for
    • Tools defined by name, description, schema
    • Model proposes; application executes
    • Validate and authorize every call
    • Return results and errors linked to the call id

    Likely follow-up: What makes a good tool description? · How do you handle a tool that takes 30 seconds?

  15. 16.What is the difference between an LLM workflow and an agent, and how do you choose?mid

    In a workflow, your code defines the path: prompt chaining, routing to specialized prompts, parallel calls whose results are combined, or an evaluator that checks and requests a revision. The model fills in steps, but the control flow is fixed and testable.

    In an agent, the model directs the process: it decides which tools to call, in what order and when it is done, in a loop until it answers or hits a limit.

    I start with the simplest thing that works: a single well-prompted call, then a workflow. An agent is justified when the steps genuinely cannot be known in advance (open-ended research, debugging, multi-step tasks over a changing environment), the task is valuable enough to absorb higher latency and cost, and errors can be caught or reversed. Agents need step limits, evals of whole trajectories and stronger guardrails.

    What interviewers listen for
    • Workflow: code controls the path
    • Agent: model controls the loop
    • Start simple and add autonomy only when needed
    • Agents cost more and need stronger guardrails

    Likely follow-up: Give an example task where an agent clearly beats a workflow. · How would you evaluate an agent's trajectory?

  16. 17.What guardrails would you put around an agent that can take real actions?hard

    I assume the model will sometimes be wrong or manipulated, and design so that mistakes are contained.

    • Least privilege: only the tools the task needs, credentials scoped to the current user, read-only where possible.
    • Validation and authorization in code for every call, not in the prompt.
    • Human confirmation for irreversible or external actions (payments, emails, deletions), showing the exact action.
    • Idempotency keys on side effects so retries cannot duplicate them.
    • Budgets: maximum steps, tokens, cost and wall-clock time, with a handoff when exceeded.
    • Error results returned to the model instead of crashes or empty strings.
    • Sandboxing for code execution and network allowlists.
    • Observability: log every model call, tool call and result, and replay failures as evals.

    These turn "the model must behave" into "the system stays safe when it does not".

    What interviewers listen for
    • Least privilege and scoped credentials
    • Confirmation for irreversible actions
    • Idempotency and budgets
    • Logging and replayable evals

    Likely follow-up: How would you stop an agent from looping forever? · Where would you put the approval step in the loop?

  17. 18.How do you get reliable JSON out of an LLM?mid

    In order of strength:

    • describe the schema and give an example in the prompt, which works most of the time but not always
    • use a JSON mode that guarantees syntactically valid JSON
    • use schema-constrained output (structured outputs or strict tool schemas), where decoding is constrained so the result matches your JSON Schema

    Even with constraints, I validate in code with a schema library such as Pydantic, because the schema cannot express every rule: value ranges, cross-field consistency, citations that point to real sources. I also check the stop reason, since an output truncated at the token limit is invalid however good the schema.

    Structured output fixes shape, not truth: a valid field can still contain a wrong value. On validation failure I retry once with the error message, then fall back. Keeping schemas small and flat with clear field descriptions also improves quality.

    What interviewers listen for
    • Prompted schema, JSON mode, constrained decoding
    • Validate in code anyway
    • Check for truncation
    • Valid shape does not mean correct content

    Likely follow-up: How would you handle optional fields the model invents? · Why might constrained decoding reduce answer quality?

  18. 19.Why do LLMs hallucinate, and how do you reduce it in a product?mid

    Models generate the most plausible continuation, not a retrieved fact. When the information is missing, ambiguous or buried, a fluent guess is often more likely than "I don't know", and training tends to reward confident answers.

    I reduce it in layers:

    • ground answers in retrieved, relevant sources, and abstain when retrieval scores are weak
    • give permission to abstain with a defined format, and reward abstention in evals
    • require citations to numbered sources, and quote exact figures
    • use tools for arithmetic, dates and live data
    • structured output so claims are machine-checkable
    • verify after generation: citations exist, numbers appear in the cited source, a judge checks paraphrases; on failure regenerate or fall back
    • measure unsupported-claim and abstention rates on an eval set

    No single technique removes it; temperature 0 and RAG alone are not enough.

    What interviewers listen for
    • Plausible text, not lookup
    • Ground and allow abstention
    • Citations plus verification
    • Measure with evals

    Likely follow-up: How do you verify a citation automatically? · Would fine-tuning fix hallucinations?

  19. 20.How do you design an evaluation suite for an LLM feature?mid

    An eval needs three parts: a fixed dataset, a way to run the real system on it, and graders.

    • Dataset: start with 20 to 50 cases from real traffic, past incidents and known risks, including edge cases, adversarial inputs and questions the system must refuse. Grow it with every production failure, and keep a held-out set that never becomes few-shot examples.
    • Graders: code-based checks wherever possible (required facts, forbidden claims, valid JSON, citations that exist, correct tool calls, abstention), and a model-based judge with a written rubric only for qualities code cannot check.
    • Process: run on every prompt, model, retrieval or tool change; compare per case with the baseline; block releases on regressions.

    Because outputs vary, I run several trials per case for important suites and look at confidence intervals before treating a small change as real.

    What interviewers listen for
    • Dataset, runner and graders
    • Real cases plus edge cases and refusals
    • Code checks first, judges second
    • Per-case comparison and noise awareness

    Likely follow-up: How many cases are enough? · How would you evaluate a multi-turn conversation?

  20. 21.When is LLM-as-judge appropriate, and how do you make it trustworthy?hard

    A judge model is useful for qualities code cannot check, such as whether an answer addresses the question, is faithful to sources in paraphrase, or follows a tone guideline. It scales far better than human review.

    But judges have known biases: preference for longer and more confident answers, position bias in pairwise comparisons, and preference for text in their own style. To make one trustworthy:

    • write a rubric with specific, ideally binary criteria and examples of failures
    • ask for reasoning before the verdict, and structured output
    • in pairwise mode, judge both orders
    • calibrate against 50 to 100 human-labelled outputs, paying attention to false passes, not just overall agreement
    • re-calibrate when the judge model or rubric changes

    I keep code-based checks for anything objective and treat the judge as one signal among several.

    What interviewers listen for
    • Useful for subjective or paraphrase checks
    • Length, position and self-preference biases
    • Specific rubric, reasoning first
    • Calibrate against human labels

    Likely follow-up: Should the judge be the same model as the one being evaluated? · How do you measure judge agreement?

  21. 22.An LLM endpoint is too slow. How would you reduce latency?mid

    First I measure where time goes: time to first token (queueing plus prompt processing), generation time (output tokens divided by speed), and everything around the model such as retrieval, reranking and tool calls.

    Then, roughly in order:

    • stream the response so users see tokens immediately
    • generate fewer tokens: tighter instructions, lower output limits, concise formats, lower reasoning effort where quality holds
    • send fewer input tokens: trim history, retrieve fewer and better chunks
    • prompt caching for a long stable prefix, which cuts time to first token
    • a smaller or faster model for routes where evals show it is good enough, or routing easy requests to it
    • parallelize independent calls and tool executions
    • cache whole responses for repeated identical requests

    I verify each change against both a latency percentile and the quality eval.

    What interviewers listen for
    • Measure TTFT, generation and surrounding work
    • Stream and shorten outputs
    • Trim input and use prompt caching
    • Smaller models and parallelism, checked with evals

    Likely follow-up: Why is output length usually the biggest lever? · How would you hide latency in a chat UI?

  22. 23.How would you cut the cost of an LLM feature without hurting quality?mid

    I start with a token profile: input versus output tokens per request, by route, and how many calls one user task takes. Then the levers, cheapest first:

    • prompt caching for stable prefixes (system prompt, tools, documents), which are billed at a large discount on cache hits
    • batch APIs for non-urgent work, often around half price in exchange for asynchronous completion
    • input hygiene: shorter history, fewer retrieved chunks, no unused tool definitions
    • output hygiene: concise formats and sensible limits
    • model routing: a smaller model for easy requests, the large one for hard ones, decided by evals
    • response caching for repeated questions
    • fewer calls per task in agent loops

    I judge cost per completed task, not per request, because a cheaper call that needs retries can cost more. Every change is checked against the quality eval before rollout.

    What interviewers listen for
    • Profile tokens per route first
    • Caching and batching are near-free wins
    • Route by difficulty, verified by evals
    • Measure cost per completed task

    Likely follow-up: What breaks prompt caching silently? · When is a batch API the wrong choice?

  23. 24.How does prompt caching work, and how do you design prompts to benefit from it?mid

    Providers can store the processed state of a prompt prefix so that later requests starting with the same prefix skip reprocessing it. Cache hits are billed at a steep discount and reduce time to first token. Depending on the provider, caching is automatic above a minimum length or enabled with explicit markers, and entries expire after a short idle period unless refreshed.

    It is an exact prefix match, so any change early in the prompt invalidates everything after it. Design accordingly:

    • put stable content first: system prompt, tool definitions, reference documents, few-shot examples
    • put variable content last: the user's question, retrieved chunks, timestamps
    • keep serialization deterministic (sorted keys, fixed tool order)
    • avoid per-request values such as dates or request ids in the system prompt

    I verify with the usage fields the API returns for cached tokens; a hit rate of zero usually means a hidden change in the prefix.

    What interviewers listen for
    • Reuses processed prefix state
    • Exact prefix match
    • Stable first, variable last
    • Verify with cached-token usage fields

    Likely follow-up: How is this different from caching whole responses? · Why can a timestamp in the system prompt be expensive?

  24. 25.How should a service handle LLM API rate limits, timeouts and outages?mid

    Provider limits usually apply to requests per minute and tokens per minute, and exceeding them returns HTTP 429, often with a Retry-After header. Server errors and overload responses also happen.

    My approach:

    • retry 429, 5xx and connection errors with exponential backoff and jitter, honouring Retry-After, with a capped number of attempts; do not retry 400-class validation errors
    • set timeouts that fit the expected output length, and stream long responses
    • put a queue in front of bulk work and control concurrency so you stay under token limits instead of bouncing off them
    • make side effects idempotent so a retried request cannot duplicate them
    • have a fallback: another model or region for critical paths, or a graceful degraded response
    • monitor error rates, latency and limit headroom, and request higher limits before launches
    What interviewers listen for
    • 429 with Retry-After, backoff with jitter
    • Retry only retryable errors
    • Queue and concurrency control
    • Idempotency and fallbacks

    Likely follow-up: Why add jitter to backoff? · How would you share a token-per-minute limit across many workers?

  25. 26.What is prompt injection, and how do you defend against it?hard

    Prompt injection is input that makes the model follow an attacker's instructions instead of yours. Direct injection comes from the user. Indirect injection hides instructions in content the model reads: web pages, emails, documents, tool results. It works because instructions and data share one channel, and there is no escaping function for natural language.

    So I design for a model that may be compromised:

    • least privilege for tools and data
    • authorization and allowlists in code before any side effect
    • human confirmation for sensitive or external actions when untrusted content is in context
    • treat output as untrusted: sanitize before rendering (remote images and links can exfiltrate data), never execute it blindly
    • mark untrusted content with delimiters and instructions, and run injection classifiers, as risk reducers, not boundaries
    • keep secrets out of prompts, and red-team with an injection corpus

    The dangerous combination is private data, untrusted content and an external channel together.

    What interviewers listen for
    • Instructions and data share one channel
    • Indirect injection via documents and tools
    • Limit damage with code-enforced controls
    • Output is untrusted too

    Likely follow-up: Why is a stronger system prompt not a fix? · How can a Markdown image leak data?

  26. 27.What are the main data leakage risks in an LLM application, and how do you mitigate them?hard

    The common ones:

    • Cross-user leakage in RAG: retrieval returns documents the user may not see. Mitigate by filtering retrieval with the user's entitlements inside the query, never by asking the model to hide things.
    • System prompt extraction: assume prompts can be revealed, so keep secrets, keys and internal URLs out of them.
    • Exfiltration through output: links or images that carry data to external hosts; sanitize rendered output.
    • Over-privileged tools: an agent with a broad service account can read far more than the user; scope credentials per user.
    • Logs and traces: prompts often contain personal data; redact, restrict access and set retention.
    • Provider data handling: check retention and training-use terms, regional processing and zero-retention options for sensitive data.
    • Training data: never fine-tune on data you would not want reproduced.

    I treat the model as a component that may repeat anything it was shown.

    What interviewers listen for
    • Permission-filtered retrieval
    • No secrets in prompts
    • Sanitize output and scope tools
    • Redact logs and check provider terms

    Likely follow-up: How would you test for cross-tenant leakage? · What is PII redaction and where would you apply it?

  27. 28.When would you fine-tune a model instead of using RAG or better prompts?hard

    They solve different problems. RAG changes what the model knows at request time: right for large or frequently changing knowledge, citations and per-user permissions. Fine-tuning changes how the model behaves: right for a strict output format, a house style, a narrow classification task, shorter prompts, or making a smaller model match a larger one on one task.

    My order is: best prompt plus few-shot examples, measured on an eval; retrieval for knowledge gaps; fine-tuning only when evals show a behaviour gap prompting cannot close, or when the cost and latency of long prompts justify it.

    Fine-tuning to add facts is the classic mistake: recall is imprecise, there are no citations, updates mean retraining and you cannot delete data from the weights. It also has upkeep: curated data, held-out evals and retraining when the base model changes. The two combine well.

    What interviewers listen for
    • RAG for knowledge, fine-tuning for behaviour
    • Prompting and retrieval first
    • Fine-tuning is a poor fact store
    • Data, evals and retraining upkeep

    Likely follow-up: How would you prove a fine-tune was worth it? · What is distillation?

  28. 29.What is LoRA, and why is parameter-efficient fine-tuning popular?hard

    Full fine-tuning updates every weight, which needs memory for weights, gradients and optimizer state for the whole model, and produces a full-size copy per task. LoRA (low-rank adaptation) freezes the base model and learns small low-rank matrices added to selected weight matrices, typically in the attention layers. Only those adapters are trained and stored.

    Benefits:

    • far less GPU memory and faster training
    • adapters are small files, so one base model can serve many tasks by swapping adapters
    • less risk of degrading the base model's general abilities than full fine-tuning

    QLoRA trains LoRA adapters on top of a quantized base model, so larger models fit on smaller GPUs. The trade-offs are hyperparameters to tune (rank, which layers, learning rate) and sometimes slightly lower quality than full fine-tuning on hard tasks. As with any tuning, a held-out eval decides.

    What interviewers listen for
    • Freeze base weights, train low-rank adapters
    • Much lower memory and storage
    • Swap adapters per task
    • QLoRA adds a quantized base

    Likely follow-up: What does the rank hyperparameter control? · Can you merge a LoRA adapter into the base weights?

  29. 30.How do you give a chatbot memory across a long conversation or across sessions?mid

    The API is stateless: the model only knows what you send each time, so memory is something the application builds.

    Within a conversation:

    • keep the system prompt pinned and the most recent turns verbatim
    • summarize older turns into a running summary when the token budget gets tight
    • drop or compress bulky tool results once they have been used

    Across sessions:

    • extract durable facts and preferences into a store (with user consent and the ability to view and delete them)
    • retrieve relevant memories per request, like RAG over the user's history, instead of injecting everything

    Risks to manage: summaries lose details, stale memories contradict new information (store timestamps and let new facts override), and memory is personal data, so it needs access control and retention rules. I evaluate with long scripted conversations that check whether key facts survive.

    What interviewers listen for
    • Stateless API, application-managed memory
    • Recent turns verbatim plus summaries
    • Persistent memories retrieved on demand
    • Consent, staleness and privacy

    Likely follow-up: How would you test that summarization keeps important facts? · What should never be stored as memory?

Prefer multiple choice? All 20 AI & LLM Engineering MCQs with answers →

esc