Attention -- How LLMs Understand Context
Knowledge
In 2017, researchers at Google published a paper titled "Attention Is All You Need." With it, they introduced the Transformer model -- the architecture on which all major LLMs are based today. The key to it is the attention mechanism.
Earlier models read text the way we read a book: word by word, from left to right. The problem: with long sentences, they had already forgotten what was at the beginning. The attention mechanism solves this in an elegant way.
Imagine a conference room: every word in your text sits at a table. At each step, every word can "look at" all other words and decide which ones are most important for its context. In the sentence "The cat sat on the mat because she was tired," the word "she" can look back and recognize that it refers to "cat" and not to "mat."

Try it yourself: click a word in the sentence -- you will immediately see what that word pays attention to.
iWhy a word may only look left
While writing an answer, the model predicts one token after another. At the moment it works on a word, the rest of the sentence does not exist yet. That is why every token may only look at itself and at everything before it -- to the right it is blind. The technical term is causal masking.
Understanding
Context Window -- The Memory of the LLM
The context window is the area that an LLM can "see" at once. Think of a desk: everything that is on it, the LLM can read and consider. What is not on the desk does not exist for the LLM.
Context windows have grown dramatically in recent years:
- Before (2022): about 8,000 tokens (roughly 6,000 words)
- Today (2025-2026): 200,000 tokens as standard, 1 million tokens for top models (Claude Opus 5, GPT-5.6, Gemini 3.1 Pro)
- Leaders: Up to 10 million tokens (Llama 4 Scout, open source) — on appropriate hardware
*Desk Analogy
An early LLM had a small side table -- room for a few pages. Today's LLMs have a huge conference table that can hold entire books. The bigger the desk, the more context the LLM can take into account.
Temperature -- Controlling Creativity
Besides the context window, there is another important parameter: temperature. It controls how "creative" or "random" the responses are.
- Temperature 0: The model always chooses the most probable next word. Result: focused, predictable, consistent. Ideal for facts and code.
- Temperature 1+: The model also chooses less probable words. Result: more creative, surprising, but also less accurate. Ideal for stories and brainstorming.
Here is what actually happens: the model computes a probability for every possible next token. Temperature does not change the order of the candidates -- it changes how clearly the favourite leads. Pull the slider and watch the bars.
And this is how it plays out in the finished text:
Prompt to the model
What is the capital of Germany?
Balanced
Good balance of precision and creativity
Model response at temperature 0.7
“Berlin, the vibrant capital of Germany, blends history and modernity.”
All example responses
What best describes the context window of an LLM?
128K Tokens
GPT-4 Turbo / GPT-4o
Books
~1.4 books
Pages
~191 Pages
Code Files
~512
~250 Tokens each
Chat Messages
~3,200
~40 Tokens each
Context Window Comparison
Note: Values are approximations. 1 Token equals roughly 3/4 of an English word.
Apply
If you notice that an LLM "forgets" early details in a long conversation, that is because of the context window. Tip: summarize important information at the beginning of your message so it stays in the context window.
Reflect
You want an LLM to come up with a creative advertising slogan. Which temperature do you choose?