Context Strategies
Knowledge
Every LLM has a context window -- the maximum amount of text it can "see" at once. Everything you send to the model (system prompt, chat history, documents, your current question) plus the response must fit within this window.
Context window sizes have evolved rapidly:
- 2022: GPT-3.5 had 4,096 tokens (~3,000 words)
- 2023: GPT-4 launched with 8K tokens (March 2023), GPT-4 Turbo introduced 128K tokens (November 2023)
- 2024: Gemini 1.5 Pro reached 1M tokens (~750,000 words)
- 2025/26: Claude Opus 5 offers 1M tokens, Gemini 3.1 also 1M
128K Tokens
GPT-4 Turbo / GPT-4o
Books
~1.4 books
Pages
~191 Pages
Code Files
~512
~250 Tokens each
Chat Messages
~3,200
~40 Tokens each
Context Window Comparison
Note: Values are approximations. 1 Token equals roughly 3/4 of an English word.
But a large context window alone doesn't solve all problems. More context means higher costs, longer response times, and the risk that important information gets "lost" in the middle (the so-called Lost in the Middle problem).
Understand
Strategy 1: Sliding Window
With the sliding window approach, the model only keeps the last N messages in context. Older messages are discarded. This is the simplest strategy and works well for short conversations.
Advantage: Cheap, fast, predictable. Disadvantage: The model "forgets" early parts of the conversation.
Strategy 2: Summarization
A smarter approach: instead of discarding old messages, an LLM summarizes them. The summary is kept as context while the original messages are removed.
Advantage: Important information is preserved. Disadvantage: Summaries can lose details; additional API calls required.
Strategy 3: RAG (Retrieval-Augmented Generation)
Instead of cramming everything into the context, RAG retrieves only the relevant information from an external database. The search query determines which documents are loaded into the context window.
Advantage: Scales to millions of documents; context stays focused. Disadvantage: Requires infrastructure (vector database, embedding pipeline).
*Rule of Thumb
Small contexts (under 10K tokens): Sliding window is enough. Medium contexts (10K-100K): Summarization helps. Large knowledge bases (100K+ documents): RAG is the right approach.
Apply
Imagine you're building a chatbot for customer support at an online store. The chatbot should:
- Have access to 500 FAQ articles
- Take the chat history with the customer into account
- Respond quickly and cost-effectively
The optimal solution combines strategies: RAG for the FAQ articles (load only relevant ones), sliding window for the current chat history, and optionally summarization for longer support sessions.
Reflect
A user has been having a technical conversation with an AI assistant for 2 hours. Suddenly the model seems to have forgotten earlier instructions. What is the most likely cause?