Zum Inhalt springen

Attention -- How LLMs Understand Context

Knowledge

In 2017, researchers at Google published a paper titled "Attention Is All You Need." With it, they introduced the Transformer model -- the architecture on which all major LLMs are based today. The key to it is the attention mechanism.

Earlier models read text the way we read a book: word by word, from left to right. The problem: with long sentences, they had already forgotten what was at the beginning. The attention mechanism solves this in an elegant way.

Imagine a conference room: every word in your text sits at a table. At each step, every word can "look at" all other words and decide which ones are most important for its context. In the sentence "The cat sat on the mat because she was tired," the word "she" can look back and recognize that it refers to "cat" and not to "mat."

The attention mechanism works like a conference — every word pays attention to the others

Try it yourself: click a word in the sentence -- you will immediately see what that word pays attention to.

Choose a sentence
Click a word -- you will see what it pays attention to.
Greyed-out words come later in the sentence. While predicting, a token may only ever look left -- to the right it is blind.
"she" looks almost entirely at "cat" -- not at "mat". This is exactly how a Transformer resolves references.
Attention distribution — «she»
cat62 %
mat14 %
she(itself)8 %
because6 %
the3 %
on3 %
Note: these values are set for teaching, not read out of a real model. They do follow the rules a real Transformer obeys: only look left, and every row adds up to 100%.

iWhy a word may only look left

While writing an answer, the model predicts one token after another. At the moment it works on a word, the rest of the sentence does not exist yet. That is why every token may only look at itself and at everything before it -- to the right it is blind. The technical term is causal masking.

Understanding

Context Window -- The Memory of the LLM

The context window is the area that an LLM can "see" at once. Think of a desk: everything that is on it, the LLM can read and consider. What is not on the desk does not exist for the LLM.

Context windows have grown dramatically in recent years:

  • Before (2022): about 8,000 tokens (roughly 6,000 words)
  • Today (2025-2026): 200,000 tokens as standard, 1 million tokens for top models (Claude Opus 5, GPT-5.6, Gemini 3.1 Pro)
  • Leaders: Up to 10 million tokens (Llama 4 Scout, open source) — on appropriate hardware

*Desk Analogy

An early LLM had a small side table -- room for a few pages. Today's LLMs have a huge conference table that can hold entire books. The bigger the desk, the more context the LLM can take into account.

Temperature -- Controlling Creativity

Besides the context window, there is another important parameter: temperature. It controls how "creative" or "random" the responses are.

  • Temperature 0: The model always chooses the most probable next word. Result: focused, predictable, consistent. Ideal for facts and code.
  • Temperature 1+: The model also chooses less probable words. Result: more creative, surprising, but also less accurate. Ideal for stories and brainstorming.

Here is what actually happens: the model computes a probability for every possible next token. Temperature does not change the order of the candidates -- it changes how clearly the favourite leads. Pull the slider and watch the bars.

Choose an example
The model has read:
The sky is ???
Sharpens or flattens the distribution before the dice are rolled.
Probability for the next token
blue71.1 %
today12.3 %
grey8.5 %
full3.5 %
clear2.1 %
overcast1.5 %
not0.9 %
green< 0.1 %
Balanced: the favourite usually wins, but there is room for variation.
Note: a real model scores all ~50,000 tokens of its vocabulary at every step. What you see here are the eight strongest candidates, normalised to 100%. The softmax behind it is real.

And this is how it plays out in the finished text:

Prompt to the model

What is the capital of Germany?

0.7

Balanced

Good balance of precision and creativity

Model response at temperature 0.7

“Berlin, the vibrant capital of Germany, blends history and modernity.”

All example responses

What best describes the context window of an LLM?

128K Tokens

GPT-4 Turbo / GPT-4o

128K

Books

~1.4 books

Pages

~191 Pages

Code Files

~512

~250 Tokens each

Chat Messages

~3,200

~40 Tokens each

Context Window Comparison

Note: Values are approximations. 1 Token equals roughly 3/4 of an English word.

Apply

If you notice that an LLM "forgets" early details in a long conversation, that is because of the context window. Tip: summarize important information at the beginning of your message so it stays in the context window.

Reflect

You want an LLM to come up with a creative advertising slogan. Which temperature do you choose?