From Prompt to Answer
Knowledge
In the beginner path you met analogies: the conference room for attention, the desk for the context window. Analogies are good entry points, but they hide what actually happens. Here is the real path.
Every token of your prompt passes through the same five stations. The model repeats this path for every single token it produces -- a sentence of 40 tokens means 40 complete passes.
iWhy GPT-2 as the example?
The numbers above come from GPT-2 small -- a model from 2019 with 124 million parameters. It is small enough to understand completely and architecturally identical to today's models: the models of 2026 have more blocks, more heads and wider vectors, but the same layout. Understand GPT-2 and you understand the rest.
Understanding
Query, key, value -- the three roles of every token
Attention sounds mysterious but it is a lookup operation. From every token vector the model derives three different versions of the same token through three learned matrices:
- Query: "What am I looking for right now?"
- Key: "What can I be found by?"
- Value: "This is what I pass on once I have been found."
The rest is simple: a token's query is matched against the keys of all tokens before it (a dot product). High agreement means a high score. The scores are divided by the square root of the dimension -- which keeps the numbers in a range where softmax does not collapse into extremes -- and then turned into weights that add up to 1. The result is a weighted sum of the values.
The crucial part: these weights are not learned. What is learned are only the three matrices that produce query, key and value. The weights themselves are recomputed on every pass, depending on what the text actually says.
Multi-head: several perspectives at once
A single set of query, key and value could only capture one kind of relationship. So the model does the same thing several times in parallel: GPT-2 small uses 12 heads of 64 dimensions each. Every head has its own matrices and therefore learns to attend to something different -- references, neighbourhood, sentence structure. At the end the results are stitched back together into 768 dimensions.
Switch between the heads below and compare how differently the same attention is distributed:
*What happens in real models
The division of labour between heads is not a theoretical construct -- in real models you can find heads that attend almost exclusively to the immediately preceding token, and others that reliably link pronouns to what they refer to. At the same time, by no means every head is this cleanly interpretable. The picture here is idealised so that the principle becomes visible.
Causal masking -- and why inference is slow
Before softmax runs, all scores pointing to tokens right of the current one are set to minus infinity. After the softmax they are exactly zero. That is the causal mask.
It has a practical consequence you notice every day:
- During training the model can process all positions of a text at once -- the mask makes sure no position peeks ahead. This is why training parallelises so well.
- During inference it cannot. Every new token depends on the previous one. That is why the answer arrives token by token, and why a long text takes linearly longer.
iKV cache
Without a trick, the model would have to recompute the keys and values of all previous tokens for every new token. Exactly that is cached -- the KV cache. It is the reason the first token of an answer takes noticeably longer than the ones that follow, and why large context costs memory above all. More on this in the "Context Strategies" section.
The compute cost of attention grows quadratically with sequence length: twice as many tokens means four times as many score computations. That is the real reason very large context windows are technically demanding.
The last row decides
After the final block there is still a matrix with one row per token. Only the last row matters for the prediction: it is projected to the size of the vocabulary -- 50,257 numbers, the logits. Softmax turns them into probabilities, and only then do the sampling parameters kick in:
- Temperature divides the logits before the softmax. Small values sharpen the distribution (the favourite almost always wins), large values flatten it.
- Top-k keeps only the k most likely candidates and renormalises.
- Top-p (nucleus sampling) keeps as many candidates as it takes for their probability to add up to p. The difference to top-k: the count adapts. For an unambiguous prediction one candidate remains, for an open slot dozens.
*For further exploration
The Transformer Explainer by the Polo Club (Georgia Tech) runs a real GPT-2 small directly in the browser and shows all intermediate results -- including the actual Q/K/V matrices. The visualisations here are deliberately simplified for teaching; if you want the real numbers, that is the place to go.
Apply
Three decisions follow from this process that you can make deliberately in practice:
1. Choose temperature by task, not by feel. Extraction, classification, code, structured JSON: temperature towards 0. As soon as you need an unambiguous, reproducible answer, any randomness is a risk.
2. Vary either temperature or top-p -- not both at once. Both parameters act on the same distribution. Turn both and you will not know afterwards which knob caused the change. The common recommendation: leave top-p at 1 and steer with temperature only, or the other way round.
3. Context length costs twice. More tokens in the prompt mean more compute per step (quadratic) and more memory for the KV cache (linear). A tidy prompt is not just didactically better, it is also faster and cheaper.
Why can a model work in parallel during training but not while generating an answer?
Reflect
You are building a pipeline that extracts structured JSON from invoices. Which sampling setting fits?