Zum Inhalt springen

From Prompt to Answer

Knowledge

In the beginner path you met analogies: the conference room for attention, the desk for the context window. Analogies are good entry points, but they hide what actually happens. Here is the real path.

Every token of your prompt passes through the same five stations. The model repeats this path for every single token it produces -- a sentence of 40 tokens means 40 complete passes.

Click a station -- you will see what happens to your text there.
Transformer block · 12×
AttentionShape of the data: 9 × 768
In plain words
Now every token may look at all the previous ones. The model derives three roles from each token: a question (query), a description (key) and a content (value). Where question and description match, a lot of that content flows back.
Technically
Scaled dot-product attention with a causal mask, split across 12 heads of 64 dimensions each. A softmax over the scores produces the weights -- exactly the ones in the attention visualisation.
The numbers come from GPT-2 small: 124M parameters, 12 blocks, 12 attention heads, 768 dimensions, a vocabulary of 50,257 tokens. Larger models have more blocks, more heads and wider vectors -- the path through the model stays the same.

iWhy GPT-2 as the example?

The numbers above come from GPT-2 small -- a model from 2019 with 124 million parameters. It is small enough to understand completely and architecturally identical to today's models: the models of 2026 have more blocks, more heads and wider vectors, but the same layout. Understand GPT-2 and you understand the rest.

Understanding

Query, key, value -- the three roles of every token

Attention sounds mysterious but it is a lookup operation. From every token vector the model derives three different versions of the same token through three learned matrices:

  • Query: "What am I looking for right now?"
  • Key: "What can I be found by?"
  • Value: "This is what I pass on once I have been found."

The rest is simple: a token's query is matched against the keys of all tokens before it (a dot product). High agreement means a high score. The scores are divided by the square root of the dimension -- which keeps the numbers in a range where softmax does not collapse into extremes -- and then turned into weights that add up to 1. The result is a weighted sum of the values.

The crucial part: these weights are not learned. What is learned are only the three matrices that produce query, key and value. The weights themselves are recomputed on every pass, depending on what the text actually says.

Multi-head: several perspectives at once

A single set of query, key and value could only capture one kind of relationship. So the model does the same thing several times in parallel: GPT-2 small uses 12 heads of 64 dimensions each. Every head has its own matrices and therefore learns to attend to something different -- references, neighbourhood, sentence structure. At the end the results are stitched back together into 768 dimensions.

Switch between the heads below and compare how differently the same attention is distributed:

Choose a sentence
Attention head
Tracks what a word refers back to
Click a word -- you will see what it pays attention to.
Greyed-out words come later in the sentence. While predicting, a token may only ever look left -- to the right it is blind.
"she" looks almost entirely at "cat" -- not at "mat". This is exactly how a Transformer resolves references.
Attention distribution — «she»
cat62 %
mat14 %
she(itself)8 %
because6 %
the3 %
on3 %
Note: these values are set for teaching, not read out of a real model. They do follow the rules a real Transformer obeys: only look left, and every row adds up to 100%.

*What happens in real models

The division of labour between heads is not a theoretical construct -- in real models you can find heads that attend almost exclusively to the immediately preceding token, and others that reliably link pronouns to what they refer to. At the same time, by no means every head is this cleanly interpretable. The picture here is idealised so that the principle becomes visible.

Causal masking -- and why inference is slow

Before softmax runs, all scores pointing to tokens right of the current one are set to minus infinity. After the softmax they are exactly zero. That is the causal mask.

It has a practical consequence you notice every day:

  • During training the model can process all positions of a text at once -- the mask makes sure no position peeks ahead. This is why training parallelises so well.
  • During inference it cannot. Every new token depends on the previous one. That is why the answer arrives token by token, and why a long text takes linearly longer.

iKV cache

Without a trick, the model would have to recompute the keys and values of all previous tokens for every new token. Exactly that is cached -- the KV cache. It is the reason the first token of an answer takes noticeably longer than the ones that follow, and why large context costs memory above all. More on this in the "Context Strategies" section.

The compute cost of attention grows quadratically with sequence length: twice as many tokens means four times as many score computations. That is the real reason very large context windows are technically demanding.

The last row decides

After the final block there is still a matrix with one row per token. Only the last row matters for the prediction: it is projected to the size of the vocabulary -- 50,257 numbers, the logits. Softmax turns them into probabilities, and only then do the sampling parameters kick in:

  • Temperature divides the logits before the softmax. Small values sharpen the distribution (the favourite almost always wins), large values flatten it.
  • Top-k keeps only the k most likely candidates and renormalises.
  • Top-p (nucleus sampling) keeps as many candidates as it takes for their probability to add up to p. The difference to top-k: the count adapts. For an unambiguous prediction one candidate remains, for an open slot dozens.
Choose an example
The model has read:
The sky is ???
Sharpens or flattens the distribution before the dice are rolled.
Only the k most likely tokens stay in the race.
Only as many tokens stay as it takes for their probability to add up to p.
Probability for the next token
blue71.1 %
today12.3 %
grey8.5 %
full3.5 %
clear2.1 %
overcast1.5 %
not0.9 %
green< 0.1 %
Balanced: the favourite usually wins, but there is room for variation.
Note: a real model scores all ~50,000 tokens of its vocabulary at every step. What you see here are the eight strongest candidates, normalised to 100%. The softmax behind it is real.

*For further exploration

The Transformer Explainer by the Polo Club (Georgia Tech) runs a real GPT-2 small directly in the browser and shows all intermediate results -- including the actual Q/K/V matrices. The visualisations here are deliberately simplified for teaching; if you want the real numbers, that is the place to go.

Apply

Three decisions follow from this process that you can make deliberately in practice:

1. Choose temperature by task, not by feel. Extraction, classification, code, structured JSON: temperature towards 0. As soon as you need an unambiguous, reproducible answer, any randomness is a risk.

2. Vary either temperature or top-p -- not both at once. Both parameters act on the same distribution. Turn both and you will not know afterwards which knob caused the change. The common recommendation: leave top-p at 1 and steer with temperature only, or the other way round.

3. Context length costs twice. More tokens in the prompt mean more compute per step (quadratic) and more memory for the KV cache (linear). A tidy prompt is not just didactically better, it is also faster and cheaper.

Why can a model work in parallel during training but not while generating an answer?

Reflect

You are building a pipeline that extracts structured JSON from invoices. Which sampling setting fits?