Ch. 30 · AI & LLM Engineering

LLM Tokens, Temperature and Top-p Sampling Explained

How LLMs turn text into tokens, why the context window is a shared budget, and what temperature, top-k and top-p actually do to the next token.

~8 min readbeginnerupdated Oct 6, 2026

“What is a token, and what does temperature do?” sounds like a warm-up, but interviewers for AI engineering roles use it to find out whether you can reason about an LLM as a component with measurable inputs and outputs. Candidates who answer “temperature makes it more creative” stop there. Candidates who explain that the model produces a probability distribution over its vocabulary, and that temperature reshapes that distribution before a token is drawn, can then explain truncated JSON, surprising bills, flaky tests and why “set temperature to 0” does not make an application deterministic.

Before you start

You need basic Python (lists, functions, the math and random modules) and a rough idea of what a probability distribution is. No machine learning background is required. The examples run on Python 3.12 or newer with the standard library only; nothing calls a paid API. Where provider behaviour matters, the article describes the common pattern and tells you to confirm details in your provider’s documentation, because limits and parameter support differ by model.

The short answer

A language model reads and writes tokens: subword pieces produced by a tokenizer such as byte-pair encoding. At each step the model outputs a score (logit) for every token in its vocabulary; a softmax turns those scores into probabilities, one token is chosen, appended to the input, and the loop repeats. Temperature divides the logits before the softmax: below 1 it sharpens the distribution toward the top token, above 1 it flattens it. Top-k keeps only the k most likely tokens and top-p keeps the smallest set whose probabilities add up to p. The context window is the maximum number of tokens the model can attend to, and the prompt and the generated output share it.

How it works

A tokenizer is trained once, before the model, on a large corpus. Byte-pair encoding starts with single characters (or bytes), counts which adjacent pair appears most often, merges it into a new symbol, and repeats until it has the vocabulary size it wants. Frequent words become a single token; rare words are split into known pieces. This toy trainer shows the idea:

from collections import Counter

def train_bpe(words, merges):
    corpus = Counter(tuple(w) for w in words)
    rules = []
    for _ in range(merges):
        pairs = Counter()
        for symbols, freq in corpus.items():
            for a, b in zip(symbols, symbols[1:]):
                pairs[(a, b)] += freq
        if not pairs:
            break
        best = max(pairs, key=pairs.get)
        rules.append(best)
        corpus = Counter({merge(s, best): f for s, f in corpus.items()})
    return rules

def merge(symbols, pair):
    out, i = [], 0
    while i < len(symbols):
        if i + 1 < len(symbols) and (symbols[i], symbols[i + 1]) == pair:
            out.append(symbols[i] + symbols[i + 1]); i += 2
        else:
            out.append(symbols[i]); i += 1
    return tuple(out)

def tokenize(word, rules):
    symbols = tuple(word)
    for rule in rules:
        symbols = merge(symbols, rule)
    return list(symbols)

words = ["lower"] * 5 + ["lowest"] * 3 + ["newer"] * 6 + ["wider"] * 2
rules = train_bpe(words, 6)
print(rules)  # [('w', 'e'), ('we', 'r'), ('l', 'o'), ('n', 'e'), ('ne', 'wer'), ('lo', 'wer')]
for w in ("lower", "newest", "slower", "xylophone"):
    print(w, tokenize(w, rules))
# lower ['lower']
# newest ['ne', 'we', 's', 't']
# slower ['s', 'lower']
# xylophone ['x', 'y', 'lo', 'p', 'h', 'o', 'ne']
python

Three consequences follow. Token counts depend on the tokenizer, so the same text costs a different number of tokens on different models. Text unlike the training corpus (code, non-English languages, long numbers, unusual identifiers) splits into more tokens. And the model never sees letters directly, which is why counting characters or reversing strings is surprisingly hard for it.

Generation is a loop. The model produces logits for every vocabulary entry, the decoder turns them into probabilities and picks one token, and the chosen token is appended to the sequence for the next step. The decoding settings live entirely in that “pick one” step:

import math, random

def softmax(logits, temperature=1.0):
    scaled = [x / temperature for x in logits]
    m = max(scaled)                       # subtract the max for numerical stability
    exps = [math.exp(x - m) for x in scaled]
    total = sum(exps)
    return [e / total for e in exps]

def top_k_filter(probs, k):
    keep = sorted(range(len(probs)), key=lambda i: probs[i], reverse=True)[:k]
    total = sum(probs[i] for i in keep)
    return [probs[i] / total if i in keep else 0.0 for i in range(len(probs))]

def top_p_filter(probs, p):
    kept, cumulative = set(), 0.0
    for i in sorted(range(len(probs)), key=lambda i: probs[i], reverse=True):
        kept.add(i)
        cumulative += probs[i]
        if cumulative >= p:
            break
    total = sum(probs[i] for i in kept)
    return [probs[i] / total if i in kept else 0.0 for i in range(len(probs))]
python

Step-by-step walkthrough

Step 1: See temperature reshape one distribution

Take a prompt like “The capital of France is” and four candidate tokens with made-up logits:

vocab = ["Paris", "Lyon", "London", "banana"]
logits = [4.0, 2.0, 1.5, -1.0]
for t in (0.2, 1.0, 2.0):
    print(t, [round(p, 3) for p in softmax(logits, t)])
# 0.2 [1.0, 0.0, 0.0, 0.0]
# 1.0 [0.817, 0.111, 0.067, 0.006]
# 2.0 [0.576, 0.212, 0.165, 0.047]
python

At 0.2 the top token holds essentially all the probability. At 2.0 even “banana” gets almost 5 percent. Temperature never changes the ranking of tokens; it changes how much probability the tail receives. That is why high temperature produces more varied text and also more nonsense.

Step 2: Cut the tail with top-k and top-p

probs = softmax(logits, 1.0)
print([round(p, 3) for p in top_k_filter(probs, 2)])    # [0.881, 0.119, 0.0, 0.0]
print([round(p, 3) for p in top_p_filter(probs, 0.9)])  # [0.881, 0.119, 0.0, 0.0]
python

Top-k always keeps a fixed number of candidates, which is too many when the model is confident and too few when it is not. Top-p adapts: when one token dominates, the nucleus might be a single token; when many continuations are plausible, it widens. Here both happen to keep the same two tokens. Open-source decoders such as Hugging Face’s generate apply temperature before the cutoffs, and hosted APIs generally advise tuning temperature or top-p, not both.

Step 3: Sample many times and compare with greedy

rng = random.Random(7)
counts = {w: 0 for w in vocab}
for _ in range(1000):
    i = rng.choices(range(len(vocab)), weights=softmax(logits, 1.0), k=1)[0]
    counts[vocab[i]] += 1
print(counts)  # {'Paris': 810, 'Lyon': 111, 'London': 75, 'banana': 4}
python

Greedy decoding always takes the argmax (“Paris”). Sampling at temperature 1 picks a wrong city about one time in five. In a long answer those small per-token chances compound: one unlucky token early on changes everything generated after it.

Step 4: Budget the context window

def fits(context_window, prompt_tokens, max_output_tokens):
    return prompt_tokens + max_output_tokens <= context_window

print(fits(128_000, 120_000, 8_000))  # True
print(fits(128_000, 121_000, 8_000))  # False
python

The prompt includes everything you send: system prompt, tool definitions, conversation history and retrieved documents. Many providers also cap output separately from the window. Count with the provider’s tokenizer or token-counting endpoint, not with a characters-divided-by-four heuristic, which is only a rough guide for English prose.

Worked scenario

A team extracts invoice fields into JSON with max_tokens set to 300 and temperature 0.7. Most invoices work. Long invoices with forty line items fail with JSONDecodeError, and the same invoice sometimes produces a different vendor name on a retry.

Two separate causes are mixed together. The parse errors come from truncation: the response hit the output limit mid-object, and the API reported it (a stop or finish reason such as max_tokens or length) but the code never checked. The inconsistent vendor names come from sampling at 0.7 on a task with one right answer.

The fix: check the stop reason before parsing and treat truncation as an error with its own handling; size the output budget from the largest expected output; use structured output or a schema-constrained mode if the provider has one; lower temperature (or leave the model default if the model does not accept sampling parameters) for extraction. Long invoices can also be split so each call extracts a bounded number of rows.

Common mistake

  • “A token is a word.” It is a subword unit; one word can be several tokens and common words with a leading space are often one.
  • “Temperature 0 is deterministic.” It approximates greedy decoding, but batching, floating-point non-associativity on GPUs and serving changes can still change outputs. Design tests that tolerate variation.
  • “The context window is for my input.” Output tokens come out of the same budget, and a long conversation silently pushes old turns out unless you manage history.
  • “Higher temperature makes the model smarter or more creative.” It only makes lower-probability tokens more likely.
  • “Every model accepts temperature and top-p.” Some recent reasoning models reject or ignore sampling parameters; check the model’s documentation.

Verify the behavior

Turn the claims into assertions and run them with python3 -m pytest or plain python3:

def test_temperature_keeps_ranking():
    for t in (0.3, 1.0, 3.0):
        p = softmax(logits, t)
        assert sorted(range(4), key=lambda i: -p[i]) == [0, 1, 2, 3]

def test_lower_temperature_sharpens():
    assert softmax(logits, 0.5)[0] > softmax(logits, 1.0)[0] > softmax(logits, 2.0)[0]

def test_top_p_keeps_only_the_nucleus():
    assert top_p_filter(softmax(logits), 0.5) == [1.0, 0.0, 0.0, 0.0]

test_temperature_keeps_ranking(); test_lower_temperature_sharpens(); test_top_p_keeps_only_the_nucleus()
print("ok")
python

Follow-up questions

Why are output tokens usually more expensive and slower than input tokens? Input tokens are processed in parallel in one forward pass; output tokens are generated one at a time, each needing its own step. Latency is roughly time to first token plus output tokens divided by generation speed.

What is the KV cache? During generation the model stores the attention keys and values for tokens it has already processed, so each new step only computes the new token. It is why long contexts use a lot of accelerator memory.

Does information position in a long prompt matter? Yes. Research such as “Lost in the Middle” found that models use facts at the start and end of long contexts more reliably than facts buried in the middle. Put instructions and the most relevant material where the model will use it, and test with your own data.

When would you use a stop sequence? When the output has a natural terminator, such as the end of a single answer in a few-shot format, so generation stops early and you pay for fewer tokens.

Interview exercise

Your chatbot works for short conversations, but after about forty turns users report that it “forgets” its persona and occasionally returns empty replies. The system prompt is 3,000 tokens, each turn averages 600 tokens, the model’s window is 32,000 tokens and you request up to 4,000 output tokens. Explain what is happening and propose a fix.

Answer and reasoning

Forty turns at 600 tokens is 24,000 tokens; add the 3,000-token system prompt and the 4,000 reserved for output and you reach 31,000, right at the 32,000 limit. Beyond that, either the request is rejected or the application’s truncation code drops content. If the code trims from the front of the message list, it may drop the system prompt, which explains the lost persona. Empty or cut-off replies fit an output budget squeezed by the remaining space. The fix is to treat the window as a budget: pin the system prompt, keep the most recent turns verbatim, summarize older turns into a compact memory, count tokens with the real tokenizer before each call, and check the stop reason so truncation is detected rather than shown to users. The reasoning is simple arithmetic, which is exactly what interviewers want to see.

Continue learning

More in AI & LLM Engineering

read ✓AI & LLM Engineering · hard

Fine-Tuning vs RAG: When to Use Each in LLM Apps

Fine-tuning changes how a model behaves; RAG changes what it knows at request time. A decision guide with data prep, costs and failure cases.

~9 min readread →
esc