AI & LLM Engineering · cheat sheet

AI & LLM Engineering

Tokens, sampling, prompting, embeddings, chunking, RAG, tool use, structured output, evals, cost and security: the LLM facts AI engineer interviews probe.

The vendor-neutral facts an AI engineering round tends to touch, from what a token is to why an agent needs a step limit.

How an LLM generates text

  • Token: subword unit from a trained tokenizer (BPE, WordPiece, SentencePiece). Limits, latency and price are all counted in tokens; counts differ by model. “About 4 characters per token” is a rough English-only heuristic.
  • Generation loop: prompt processed in parallel → logits for every vocabulary token → softmax → pick one → append → repeat until a stop token, stop sequence or output limit.
  • Context window: maximum tokens per request, shared by system prompt, tools, history, retrieved text and output. Many APIs also cap output separately.
  • Latency ≈ time to first token (queue + prompt processing) + output tokens ÷ generation speed. Output tokens are usually the bigger lever and the pricier ones.
  • KV cache: stored attention keys/values for processed tokens, so each step only computes the new token. Long contexts need lots of accelerator memory.
  • Models predict plausible text; they do not look facts up and have a training cutoff.

Decoding parameters

Setting Effect Use
temperature < 1 sharpens distribution, ranking unchanged extraction, classification, code
temperature > 1 flattens it, tail tokens more likely brainstorming, variety
top-k keep the k most likely tokens crude tail cut
top-p (nucleus) keep smallest set with cumulative prob ≥ p adaptive tail cut
stop sequences end generation at a marker few-shot formats, save tokens
max output tokens hard cap; truncates mid-output size from the largest expected output

Gotcha

Temperature 0 approximates greedy decoding but is not a determinism guarantee. Some reasoning models reject or ignore sampling parameters entirely; check the model docs. Tune temperature or top-p, not both.

Prompting

  • Clear task, audience, context, exact output format and edge-case behaviour (including “say you don’t know”).
  • Stable instructions in the system prompt; per-request data in the user turn; label sections so instructions and data are distinct.
  • Few-shot examples when showing beats telling; make them diverse and representative (the model copies accidental patterns too).
  • Chain-of-thought helps multi-step tasks on standard models; reasoning models think internally and are tuned with an effort or budget setting.
  • Prompts are code: version them, evaluate them, change one thing at a time.
  • Embedding = fixed-length vector; similar meaning → nearby vectors. Same model (and version) for queries and documents; a model change means re-embedding.
  • Normalize once: then cosine = dot product and squared Euclidean = 2 - 2·cos, so all three rank identically.
  • Scores are not probabilities; calibrate thresholds on labelled data.
Index How Trade-off
Flat (exact) scan every vector perfect recall, linear cost
HNSW layered proximity graph, ef_search high recall + speed, memory-heavy, slow build
IVF k-means lists, scan n_probe nearest cheaper, recall depends on n_probe and data
PQ / scalar quantization compress vectors less memory, some accuracy loss
  • Measure recall@k against exact search on real queries before tuning parameters.
  • Filter inside the search (tenant, permissions, date); post-filtering top-k can leave nothing.
  • Hybrid search: BM25 + vectors, merged with reciprocal rank fusion, then a cross-encoder reranker over the top 20 to 50.
def rrf(rankings, k=60):
    scores = {}
    for ranking in rankings:
        for rank, doc in enumerate(ranking, start=1):
            scores[doc] = scores.get(doc, 0) + 1 / (k + rank)
    return sorted(scores, key=scores.get, reverse=True)
python

RAG pipeline

Stage Key decisions
Parse keep headings, lists, tables; drop boilerplate
Chunk split on structure (heading → paragraph → sentence), size cap, prefix with doc + section title
Index text + vector + metadata (source, section, ACL, version, date); incremental updates and deletes
Query rewrite follow-ups into standalone queries
Retrieve hybrid, permission-filtered, rerank, top 3 to 8 into the prompt
Generate answer only from numbered sources, cite them, defined “I don’t know” path
Evaluate retrieval (recall@k, MRR) separately from answers (faithfulness, correctness, abstention)

Interview tip

When a RAG answer is wrong, look at the retrieved chunks first. If the answer is not in them, no prompt change will fix it.

Tool use & agents

  • Model proposes a call (name + JSON args + id); your code validates, authorizes, executes and returns a result linked to the id. Return failures as error results.
  • Parallel calls: run concurrently, return all results together.
  • Workflow (code controls the path: chaining, routing, parallelization, evaluator loops) vs agent (model controls the loop). Start simple; add autonomy only when steps cannot be known in advance.
  • Guardrails: least privilege, scoped credentials, max steps/tokens/time, idempotency keys, human confirmation for irreversible actions, sandboxed execution, full logging.
  • MCP (Model Context Protocol): open standard for exposing tools and resources to LLM hosts.
for step in range(max_steps):
    reply = call_model(messages, tools)
    if not reply.tool_calls:
        return reply.text
    messages += [reply, *[run_tool(c) for c in reply.tool_calls]]
return handoff_to_human()
python

Structured output & hallucinations

  • Strength: schema in prompt < JSON mode (valid JSON) < schema-constrained decoding (matches JSON Schema). Still validate values in code and check the stop reason for truncation.
  • Structured output fixes shape, not truth.
  • Hallucination layers: ground in sources → abstain when retrieval is weak → require citations and exact quotes → tools for maths, dates and live data → verify claims after generation (numbers present in cited source, judge for paraphrase) → measure unsupported-claim and abstention rates.

Evals

  • Eval = dataset + runner + graders. Start with 20 to 50 real cases including refusals and adversarial inputs; grow from production failures; keep a held-out set.
  • Code checks first (facts, format, citations, tool args, abstention); LLM-as-judge only for subjective qualities, with a specific rubric, reasoning before verdict, both orders for pairwise, and calibration against human labels.
  • Judge biases: length, confidence, position, self-preference.
  • Noise: 80% on 50 cases has a 95% interval of roughly 69 to 91%. Compare per case; run multiple trials.

Latency & cost levers

Lever Helps Note
Streaming perceived latency users read while the rest generates
Shorter outputs latency + cost biggest single lever
Prompt caching TTFT + input cost exact prefix match: stable content first, variable last
Batch API cost (often ~half price) asynchronous, hours not seconds
Smaller model / routing cost + latency verify per route with evals
Fewer, better chunks cost + quality rerank instead of stuffing
Response cache repeated questions key on normalized input
Backoff with jitter reliability retry 429/5xx, honour Retry-After, never retry 400s

Security

  • Prompt injection: instructions hidden in input; indirect via web pages, emails, documents, tool results. No escaping exists for natural language, so limit what a compromised model can do.
  • Dangerous trio: private data + untrusted content + an external channel. Remove a leg or put a human in the path.
  • Controls: allowlists and authorization in code, confirmation for side effects, scoped credentials, sanitize rendered output (remote images and links can exfiltrate), classifiers as signals, red-team corpus.
  • Leakage: permission-filtered retrieval, no secrets in system prompts, redacted logs with retention limits, provider retention and training terms checked.

Fine-tuning vs RAG

Need RAG Fine-tune
Changing knowledge, citations, permissions yes no
Format, tone, narrow labels partly yes
Shorter prompts, smaller model (distillation) no yes
  • Order: prompt + few-shot → retrieval → fine-tune only for a measured behaviour gap.
  • LoRA: freeze base, train small low-rank adapters; QLoRA: same on a quantized base.
  • Fine-tuning is a poor fact store; retrain and re-evaluate when the base model changes.

Practice with the AI & LLM engineering interview questions.

esc