The vendor-neutral facts an AI engineering round tends to touch, from what a token is to why an agent needs a step limit.
How an LLM generates text
- Token: subword unit from a trained tokenizer (BPE, WordPiece, SentencePiece). Limits, latency and price are all counted in tokens; counts differ by model. “About 4 characters per token” is a rough English-only heuristic.
- Generation loop: prompt processed in parallel → logits for every vocabulary token → softmax → pick one → append → repeat until a stop token, stop sequence or output limit.
- Context window: maximum tokens per request, shared by system prompt, tools, history, retrieved text and output. Many APIs also cap output separately.
- Latency ≈ time to first token (queue + prompt processing) + output tokens ÷ generation speed. Output tokens are usually the bigger lever and the pricier ones.
- KV cache: stored attention keys/values for processed tokens, so each step only computes the new token. Long contexts need lots of accelerator memory.
- Models predict plausible text; they do not look facts up and have a training cutoff.
Decoding parameters
| Setting | Effect | Use |
|---|---|---|
| temperature < 1 | sharpens distribution, ranking unchanged | extraction, classification, code |
| temperature > 1 | flattens it, tail tokens more likely | brainstorming, variety |
| top-k | keep the k most likely tokens | crude tail cut |
| top-p (nucleus) | keep smallest set with cumulative prob ≥ p | adaptive tail cut |
| stop sequences | end generation at a marker | few-shot formats, save tokens |
| max output tokens | hard cap; truncates mid-output | size from the largest expected output |
Gotcha
Temperature 0 approximates greedy decoding but is not a determinism guarantee. Some reasoning models reject or ignore sampling parameters entirely; check the model docs. Tune temperature or top-p, not both.
Prompting
- Clear task, audience, context, exact output format and edge-case behaviour (including “say you don’t know”).
- Stable instructions in the system prompt; per-request data in the user turn; label sections so instructions and data are distinct.
- Few-shot examples when showing beats telling; make them diverse and representative (the model copies accidental patterns too).
- Chain-of-thought helps multi-step tasks on standard models; reasoning models think internally and are tuned with an effort or budget setting.
- Prompts are code: version them, evaluate them, change one thing at a time.
Embeddings & vector search
- Embedding = fixed-length vector; similar meaning → nearby vectors. Same model (and version) for queries and documents; a model change means re-embedding.
- Normalize once: then cosine = dot product and squared Euclidean =
2 - 2·cos, so all three rank identically. - Scores are not probabilities; calibrate thresholds on labelled data.
| Index | How | Trade-off |
|---|---|---|
| Flat (exact) | scan every vector | perfect recall, linear cost |
| HNSW | layered proximity graph, ef_search |
high recall + speed, memory-heavy, slow build |
| IVF | k-means lists, scan n_probe nearest |
cheaper, recall depends on n_probe and data |
| PQ / scalar quantization | compress vectors | less memory, some accuracy loss |
- Measure recall@k against exact search on real queries before tuning parameters.
- Filter inside the search (tenant, permissions, date); post-filtering top-k can leave nothing.
- Hybrid search: BM25 + vectors, merged with reciprocal rank fusion, then a cross-encoder reranker over the top 20 to 50.
def rrf(rankings, k=60):
scores = {}
for ranking in rankings:
for rank, doc in enumerate(ranking, start=1):
scores[doc] = scores.get(doc, 0) + 1 / (k + rank)
return sorted(scores, key=scores.get, reverse=True)RAG pipeline
| Stage | Key decisions |
|---|---|
| Parse | keep headings, lists, tables; drop boilerplate |
| Chunk | split on structure (heading → paragraph → sentence), size cap, prefix with doc + section title |
| Index | text + vector + metadata (source, section, ACL, version, date); incremental updates and deletes |
| Query | rewrite follow-ups into standalone queries |
| Retrieve | hybrid, permission-filtered, rerank, top 3 to 8 into the prompt |
| Generate | answer only from numbered sources, cite them, defined “I don’t know” path |
| Evaluate | retrieval (recall@k, MRR) separately from answers (faithfulness, correctness, abstention) |
Interview tip
When a RAG answer is wrong, look at the retrieved chunks first. If the answer is not in them, no prompt change will fix it.
Tool use & agents
- Model proposes a call (name + JSON args + id); your code validates, authorizes, executes and returns a result linked to the id. Return failures as error results.
- Parallel calls: run concurrently, return all results together.
- Workflow (code controls the path: chaining, routing, parallelization, evaluator loops) vs agent (model controls the loop). Start simple; add autonomy only when steps cannot be known in advance.
- Guardrails: least privilege, scoped credentials, max steps/tokens/time, idempotency keys, human confirmation for irreversible actions, sandboxed execution, full logging.
- MCP (Model Context Protocol): open standard for exposing tools and resources to LLM hosts.
for step in range(max_steps):
reply = call_model(messages, tools)
if not reply.tool_calls:
return reply.text
messages += [reply, *[run_tool(c) for c in reply.tool_calls]]
return handoff_to_human()Structured output & hallucinations
- Strength: schema in prompt < JSON mode (valid JSON) < schema-constrained decoding (matches JSON Schema). Still validate values in code and check the stop reason for truncation.
- Structured output fixes shape, not truth.
- Hallucination layers: ground in sources → abstain when retrieval is weak → require citations and exact quotes → tools for maths, dates and live data → verify claims after generation (numbers present in cited source, judge for paraphrase) → measure unsupported-claim and abstention rates.
Evals
- Eval = dataset + runner + graders. Start with 20 to 50 real cases including refusals and adversarial inputs; grow from production failures; keep a held-out set.
- Code checks first (facts, format, citations, tool args, abstention); LLM-as-judge only for subjective qualities, with a specific rubric, reasoning before verdict, both orders for pairwise, and calibration against human labels.
- Judge biases: length, confidence, position, self-preference.
- Noise: 80% on 50 cases has a 95% interval of roughly 69 to 91%. Compare per case; run multiple trials.
Latency & cost levers
| Lever | Helps | Note |
|---|---|---|
| Streaming | perceived latency | users read while the rest generates |
| Shorter outputs | latency + cost | biggest single lever |
| Prompt caching | TTFT + input cost | exact prefix match: stable content first, variable last |
| Batch API | cost (often ~half price) | asynchronous, hours not seconds |
| Smaller model / routing | cost + latency | verify per route with evals |
| Fewer, better chunks | cost + quality | rerank instead of stuffing |
| Response cache | repeated questions | key on normalized input |
| Backoff with jitter | reliability | retry 429/5xx, honour Retry-After, never retry 400s |
Security
- Prompt injection: instructions hidden in input; indirect via web pages, emails, documents, tool results. No escaping exists for natural language, so limit what a compromised model can do.
- Dangerous trio: private data + untrusted content + an external channel. Remove a leg or put a human in the path.
- Controls: allowlists and authorization in code, confirmation for side effects, scoped credentials, sanitize rendered output (remote images and links can exfiltrate), classifiers as signals, red-team corpus.
- Leakage: permission-filtered retrieval, no secrets in system prompts, redacted logs with retention limits, provider retention and training terms checked.
Fine-tuning vs RAG
| Need | RAG | Fine-tune |
|---|---|---|
| Changing knowledge, citations, permissions | yes | no |
| Format, tone, narrow labels | partly | yes |
| Shorter prompts, smaller model (distillation) | no | yes |
- Order: prompt + few-shot → retrieval → fine-tune only for a measured behaviour gap.
- LoRA: freeze base, train small low-rank adapters; QLoRA: same on a quantized base.
- Fine-tuning is a poor fact store; retrain and re-evaluate when the base model changes.
Practice with the AI & LLM engineering interview questions.