Chunking Strategies 2026
Knowledge
Chunking determines what your RAG system finds. No matter how good your embedding model is, no matter how fast your vector database -- if the chunks are poorly cut, the system delivers poor results. Chunking is the silent quality lever of every RAG pipeline.
The three key strategies in 2026:
- Fixed-Size Chunking: Cut documents into sections of fixed token length
- Semantic Chunking: Cut at content boundaries based on embedding similarity
- Parent-Child Chunking: Hierarchical structure with coarse parent chunks and fine child chunks
Understand
Fixed-Size Chunking -- The Classic (and Its Limits)
Fixed-size chunking splits documents into sections of equal length -- for example, 512 tokens per chunk. Simple to implement, deterministic, easy to understand.
The problem: Text doesn't care about token counts. A chunk can end in the middle of a paragraph, split a thought, or cut off context from the previous section.
Overlap as a band-aid: To rescue split contexts, you add overlaps -- typically 10--20% of the chunk size. For 512 tokens, that's 50--100 tokens of overlap.
| Chunk Size | Overlap | Advantage | Disadvantage |
|---|---|---|---|
| 256 tokens | 25 | High granularity, precise matches | Much context is lost |
| 512 tokens | 50 | Good compromise | Default setting, not optimized |
| 1024 tokens | 100 | More context per chunk | Dilutes relevance for short questions |
iFixed-Size Isn't Dead
For homogeneous documents with uniform information density (e.g., log files, tabular data), fixed-size chunking still works well. It's outdated as a universal solution, not as a tool.
Semantic Chunking -- Cut Where the Content Demands It
Semantic chunking analyzes the text and identifies content boundaries. The algorithm:
- Split the text into sentences
- Compute the embedding for each sentence
- Compare consecutive sentence embeddings (cosine similarity)
- When similarity drops below a threshold, set a chunk break
- Group sentences between breaks into chunks
The result: chunks that form thematic units. A paragraph about return policies stays together instead of being cut off after 512 tokens mid-sentence.
Trade-offs:
- Non-deterministic chunk sizes (some chunks are 200 tokens, others 800)
- Dependent on embedding model quality
- Threshold must be tuned per document type
- Higher computational cost during indexing
A company indexes contracts with clearly structured sections (paragraphs, clauses, sub-items). Which chunking strategy maximizes retrieval quality?
Parent-Child Chunking -- The Best of Both Worlds
Parent-child chunking solves a fundamental dilemma: small chunks are precise for search but provide too little context for generation. Large chunks provide context but dilute relevance in search.
The solution: Two levels
- Child chunks (small, 100--300 tokens): Used for vector search. High precision matching.
- Parent chunks (large, 500--1500 tokens): Passed to the LLM as context. Contain the full surrounding context.
The flow:
- Query embedding is searched against child chunks
- The most relevant child chunks are identified
- Instead of the child chunks, the associated parent chunks are passed to the LLM
- The LLM has enough context to generate a well-grounded answer
Sample text
20 words per chunk, 5-word overlap. Sometimes cuts mid-sentence.
Chunks
7
Avg. words
18
Overlap
5 words
Compare the three strategies and observe how chunk size and quality differ.
Why does parent-child chunking use different chunks for retrieval and generation?
Apply
Choosing a Chunking Strategy
| Document Type | Recommended Strategy | Rationale |
|---|---|---|
| Flowing text (articles, reports) | Semantic Chunking | Content boundaries form natural units |
| Structured documents (contracts, manuals) | Structure-based + Parent-Child | Leverage existing structure, map hierarchy |
| Code documentation | Function/class-based | Logical code units as chunks |
| Tabular data | Row or section-based | Preserve rows as atomic units |
| Chat histories | Message-based with context window | Dialog turns as natural units |
Overlap Strategies in Detail
Overlap isn't a universal solution, but often useful:
- Token overlap: Fixed number of tokens overlap. Simple but blind.
- Sentence overlap: Repeat complete sentences from the edge of the previous chunk. Preserves grammatical completeness.
- Sliding window: Chunks slide across the text with a fixed step. Maximum overlap, but also maximum redundancy in the DB.
Reflect
Think of a scenario from your work: which documents would you make accessible via RAG? Which chunking strategy fits best -- and why? Remember that the "right" strategy always depends on the document type and the nature of the questions.