Zum Inhalt springen

Chunking Strategies 2026

Knowledge

Chunking determines what your RAG system finds. No matter how good your embedding model is, no matter how fast your vector database -- if the chunks are poorly cut, the system delivers poor results. Chunking is the silent quality lever of every RAG pipeline.

The three key strategies in 2026:

  • Fixed-Size Chunking: Cut documents into sections of fixed token length
  • Semantic Chunking: Cut at content boundaries based on embedding similarity
  • Parent-Child Chunking: Hierarchical structure with coarse parent chunks and fine child chunks

Understand

Fixed-Size Chunking -- The Classic (and Its Limits)

Fixed-size chunking splits documents into sections of equal length -- for example, 512 tokens per chunk. Simple to implement, deterministic, easy to understand.

The problem: Text doesn't care about token counts. A chunk can end in the middle of a paragraph, split a thought, or cut off context from the previous section.

Overlap as a band-aid: To rescue split contexts, you add overlaps -- typically 10--20% of the chunk size. For 512 tokens, that's 50--100 tokens of overlap.

Chunk SizeOverlapAdvantageDisadvantage
256 tokens25High granularity, precise matchesMuch context is lost
512 tokens50Good compromiseDefault setting, not optimized
1024 tokens100More context per chunkDilutes relevance for short questions

iFixed-Size Isn't Dead

For homogeneous documents with uniform information density (e.g., log files, tabular data), fixed-size chunking still works well. It's outdated as a universal solution, not as a tool.

Semantic Chunking -- Cut Where the Content Demands It

Semantic chunking analyzes the text and identifies content boundaries. The algorithm:

  1. Split the text into sentences
  2. Compute the embedding for each sentence
  3. Compare consecutive sentence embeddings (cosine similarity)
  4. When similarity drops below a threshold, set a chunk break
  5. Group sentences between breaks into chunks

The result: chunks that form thematic units. A paragraph about return policies stays together instead of being cut off after 512 tokens mid-sentence.

Trade-offs:

  • Non-deterministic chunk sizes (some chunks are 200 tokens, others 800)
  • Dependent on embedding model quality
  • Threshold must be tuned per document type
  • Higher computational cost during indexing

A company indexes contracts with clearly structured sections (paragraphs, clauses, sub-items). Which chunking strategy maximizes retrieval quality?

Parent-Child Chunking -- The Best of Both Worlds

Parent-child chunking solves a fundamental dilemma: small chunks are precise for search but provide too little context for generation. Large chunks provide context but dilute relevance in search.

The solution: Two levels

  • Child chunks (small, 100--300 tokens): Used for vector search. High precision matching.
  • Parent chunks (large, 500--1500 tokens): Passed to the LLM as context. Contain the full surrounding context.

The flow:

  1. Query embedding is searched against child chunks
  2. The most relevant child chunks are identified
  3. Instead of the child chunks, the associated parent chunks are passed to the LLM
  4. The LLM has enough context to generate a well-grounded answer

Sample text

Machine learning is a subfield of artificial intelligence. Algorithms learn from data without being explicitly programmed. Typical applications include image recognition, natural language processing, and recommendation systems. Neural networks are inspired by the human brain. They consist of layers of artificial neurons. By training on large amounts of data, they can recognize complex patterns. Deep learning uses particularly deep networks with many layers. Transformer models have revolutionized language processing. The attention mechanism allows capturing relationships between all words simultaneously. GPT, BERT, and T5 are well-known transformer architectures. They are used for translation, summarization, and text generation.

20 words per chunk, 5-word overlap. Sometimes cuts mid-sentence.

#120 wordsSentence cutMachine learning is a subfield of artificial intelligence. Algorithms learn from data without being explicitly programmed. Typical applications include image
#220 wordsprogrammed. Typical applications include image recognition, natural language processing, and recommendation systems. Neural networks are inspired by the human brain.
#320 wordsSentence cutinspired by the human brain. They consist of layers of artificial neurons. By training on large amounts of data, they
#420 wordsSentence cutlarge amounts of data, they can recognize complex patterns. Deep learning uses particularly deep networks with many layers. Transformer models
#520 wordsSentence cutwith many layers. Transformer models have revolutionized language processing. The attention mechanism allows capturing relationships between all words simultaneously. GPT,
#620 wordsSentence cutbetween all words simultaneously. GPT, BERT, and T5 are well-known transformer architectures. They are used for translation, summarization, and text
#76 wordsfor translation, summarization, and text generation.

Chunks

7

Avg. words

18

Overlap

5 words

Compare the three strategies and observe how chunk size and quality differ.

Why does parent-child chunking use different chunks for retrieval and generation?

Apply

Choosing a Chunking Strategy

Document TypeRecommended StrategyRationale
Flowing text (articles, reports)Semantic ChunkingContent boundaries form natural units
Structured documents (contracts, manuals)Structure-based + Parent-ChildLeverage existing structure, map hierarchy
Code documentationFunction/class-basedLogical code units as chunks
Tabular dataRow or section-basedPreserve rows as atomic units
Chat historiesMessage-based with context windowDialog turns as natural units

Overlap Strategies in Detail

Overlap isn't a universal solution, but often useful:

  • Token overlap: Fixed number of tokens overlap. Simple but blind.
  • Sentence overlap: Repeat complete sentences from the edge of the previous chunk. Preserves grammatical completeness.
  • Sliding window: Chunks slide across the text with a fixed step. Maximum overlap, but also maximum redundancy in the DB.

Reflect

Think of a scenario from your work: which documents would you make accessible via RAG? Which chunking strategy fits best -- and why? Remember that the "right" strategy always depends on the document type and the nature of the questions.