Advanced Chunking Techniques: What The Benchmarks Actually Show
Artificial Intelligence Pratik BhavsarKey takeaways
- Chunking impacts retrieval quality, vector database costs, and generation latency simultaneously; size should be tuned based on the specific document structure and the length of a complete answer within your corpus.
- Advanced techniques like context-enriched or late chunking should be reserved for specific, identified failures, as simpler strategies (like recursive splitting) often yield the majority of performance gains.
- Quantifiable performance monitoring using evaluators like Chunk Relevance and Context Precision is necessary to distinguish between retrieval bottlenecks and generation errors in production.
Your retriever returns ten chunks, your retrieval-augmented generation (RAG) app still hallucinates, and doubling chunk size doubles your time to first token. Chunk size is the one parameter that touches retrieval quality, vector storage bills, and generation latency at once.
What is chunking?
Chunking breaks text into smaller units. Each chunk is vectorized and stored, and that unit decides what your retriever can return, what your vector database costs, and how much context your large language model (LLM) reads.
In this screenshot, ChunkViz renders how different splitters cut the same paragraph:
The role of chunking
Chunk size is not one decision with one right answer. Chunk size is a single number pulling against four constraints at once, and they do not agree: retrieval quality wants chunks sized to your answers, storage cost wants them large, latency wants them small and few, and hallucination risk wants fewer of them in the prompt.
Knowing which constraint binds hardest in your system is most of the work.
Retrieval quality
There is no universally correct chunk size. There is a size that matches how long a complete answer runs in your corpus. A 2025 token-size sweep makes the point by contradiction: 64-token chunks won on SQuAD, where an answer is a phrase, while NarrativeQA kept improving all the way to 1,024 tokens because its answers only make sense with surrounding narrative. Same benchmark, same models, opposite conclusions. Before you tune anything, pull ten real answers from your own corpus and note how much surrounding text a person needs to justify each one. That length is your starting chunk size.
Vector database cost and query latency
Halving your chunk size doubles your vector count, and cost scales with that count rather than with the size of your corpus. Qdrant's sizing guidance puts full-precision RAM at num_vectors × dimensions × 4 bytes × 1.5, and Pinecone serverless charges 1 read unit per GB of namespace scanned. The same corpus at 256 tokens per chunk therefore costs roughly four times as much to query as it does at 1,024. Do that multiplication before a chunking sweep rather than after: a change you make for retrieval reasons lands on the infrastructure bill.
LLM latency and cost
Chunk size and top-k multiply, and most teams tune them in separate meetings. NVIDIA's enterprise RAG sizing guide measured time to first token climbing from about 2.0 seconds at 256-token chunks to 7.1 at 1,024 — and, separately, raising top-k from four to 10 pushed it from 3.6 to 7.9. Move both and you are not adding the two penalties, you are multiplying them. If you have a latency budget, spend it on one lever and hold the other fixed.
LLM hallucinations
The instinct when an answer comes back wrong is to retrieve more. The evidence says that usually makes it worse. The 2025 RAGUARD benchmark found that going from one retrieved document to five raised the recall of misleading content from 21.3% to 44.8%, and a 2025 study measured performance dropping by up to 20% from added documents even with total context length and the position of the relevant passage held constant. Extra context is not free insurance. Fewer, better chunks beat more chunks.
Factors influencing chunking
Four variables should drive the decision, and each one resolves to a rule you can apply today rather than a number to memorize.
Text structure
Split on the boundaries the document already has. Across 1,080 configurations spanning 36 strategies, six domains and five embedding models, paragraph-group chunking roughly doubled ranking quality against fixed character chunking. So, the question is not which algorithm to run, it is whether your documents have real structure to respect.
- Headings, paragraphs, and sections are boundaries an author already decided were meaningful, and honoring them beats any character count you could tune.
- Transcripts, scanned text and chat logs have none, which is the honest case for fixed-size splitting.
Embedding model limits
Treat the stated maximum as a ceiling to stay well under, not a target. All-MiniLM-L6-v2 truncates at 256 word pieces, OpenAI's text-embedding-3 models accept 8,192 tokens, and Cohere Embed v4 accepts 128,000 — but Jina AI's long-context study found retrieval quality degrading steadily with length well before any of those limits. A model that accepts 8,192 tokens does not embed 8,192 tokens well. Find the length where your model's retrieval still holds up on your data, then confirm it fits the limit. Not the other way round.
LLM context length
A large window is not a substitute for good retrieval. Windows have grown to 400,000 tokens for GPT-5, yet in the No Literal Matching benchmark, 11 of 13 models claiming 128,000+ context reached only half their short-context performance at 32,000 tokens, and a separate study found RAG with 600-token chunks beating a full 128,000-token context outright. Filling the window is the expensive way to get a worse answer. Read the window as the ceiling on chunk size multiplied by top-k — it bounds your choices, it does not make them for you.
Query type
The shape of the question sets the size of the chunk. NVIDIA's chunking study across five datasets found factoid queries favoring 256–512-token chunks while analytical queries favored 1,024-token or page-level chunks. Most production systems get both, and a compromise size in the middle tends to serve neither well. If you can route by query type, do that. If you cannot, size for whichever type carries your traffic and accept that the other will underperform — deliberately, rather than by accident.
How to determine your own requirements first
Before testing a single splitter, answer five questions about your corpus. The answers rule strategies out rather than ranking them, which narrows the field faster than any benchmark table.
| Requirement | Your answer |
| Document structure | Consistent headings / Mixed formats / Unstructured |
| Dominant query type | Factual / Analytical / Multi-hop |
| Required metadata | Source, temporal, entities, other |
| Embedding model limit | 512 / 1,024 / 8,192 tokens |
| Average paragraph length | X tokens |
Consistent headings plus analytical queries points at structure-aware chunking. Unstructured documents plus factual queries points at smaller fixed-size chunks with semantic boundaries. Either way you now have two candidates to test instead of eight.
Character and recursive character splitting
Managed platforms disagree on defaults: Bedrock Knowledge Bases use ~300 tokens with 20% overlap, Azure AI Search recommends 512 with 25%, and Vertex AI RAG Engine defaults to 1,024 with 256 overlap. All LangChain splitters share the TextSplitter base class, now in the langchain-text-splitters package (v1.1.2).
CharacterTextSplitter splits on one separator and merges the results, ignoring where sentences end.
RecursiveCharacterTextSplitter tries separators in order — paragraph, line, space, character — and recurses into any split still larger than chunk_size. LangChain's docs call it the recommended splitter for generic text.
Chroma's token-level chunking evaluation reported that “the heuristic RecursiveCharacterTextSplitter with chunk size 200 and no overlap performs well,” with 88.1 recall at n=5. If retrieval quality is poor and you are using character splitting, look at where the splits land before changing anything else.
Sentence and semantic splitting
Sentence splitting uses spaCy's sents to find boundaries, then groups whole sentences by stride and overlap — useful where sentence integrity beats uniform size, like FAQ pairs and legal clauses. spaCy 3.8's en_core_web_sm ships a senter component the spaCy docs describe as roughly 10x faster than the parser.
Semantic splitting goes further, embedding each sentence and cutting where cosine similarity to the previous one drops below a threshold (credit: agamm/semantic-split).
Qu et al. (2025) tested semantic chunkers against fixed-size across 10 datasets. Fixed-size won on three of five evidence-retrieval datasets, and swapping the embedding model produced a 7.44% average F1 gain — larger than any chunking-strategy difference: “The computational costs associated with semantic chunking are not justified by consistent performance gains.” An August 2026 evaluation reached the same verdict for semantic, contextual, and late chunking.
Price it before you commit. A 10,000-document corpus averaging 100 sentences per document is a million embedding calls just to chunk — around $100 of preprocessing before you have indexed anything.
Context-enriched, late and agentic chunking
Three strategies give a chunk awareness of its surroundings, and they pay for it in different currencies.
- Context-enriched chunking prepends a short LLM-written statement to each chunk before embedding. “Revenue increased 15% quarter-over-quarter” becomes “From Q3 2024 Earnings Report, Financial Performance section: Revenue increased 15% quarter-over-quarter,” so a query about Q3 2024 matches a chunk that never mentions it. Anthropic's Contextual Retrieval cut top-20 retrieval failures by 35% with embeddings alone, 49% with BM25, and 67% with reranking added, at $1.02 per million document tokens with prompt caching. The cost is one LLM call per chunk.
- Late chunking reverses the order: run a long-context embedding model over the whole document, then apply chunk boundaries at the final pooling step. Each chunk embedding carries its position in the document, with no extra LLM call. Across three models and four BEIR datasets it improved nDCG@10 by 3.63% relative over naive chunking. The cost is that the whole document must fit the model's context, so it works best under about 8K tokens.
- Agentic chunking hands the boundary decision to an LLM that reads the text and reasons about where the breaks belong. Chonkie's SlumberChunker implements this. The cost is roughly $1,000+ per 100,000 documents zlus non-deterministic boundaries between runs, which rules it out for most production corpora. Reserve it for a small corpus of genuinely high-value documents.
LLM-based chunking
When coreferences keep atomic facts from retrieving cleanly, Dense X Retrieval introduces “propositions”: atomic, self-contained factoids with coreferences resolved. Proposition-level indexing improved Recall@20 by +10.1 over passages for unsupervised dense retrievers. The authors' propositionizer model is still on Hugging Face. Two caveats:
- A 2025 short paper found the technique “does not generalize well to domains with lower fact density, such as novels or scientific papers.”
- The 2026 evaluation calls propositions a specialized choice rather than a default.
Multi-vector indexing searches on a vector derived from something other than the raw text, then returns the original document: embed child chunks and return the parent, embed a generated summary, or embed hypothetical questions the document would answer.
Document-specific splitting
For mixed-format “Chat with PDF,” parsing counts as chunking. Unstructured partitions document, spreadsheet, presentation, image, JSON and audio formats across four partitioning strategies. Note that infer_table_structure now defaults to False, so turn it on for table extraction.
This returns typed elements (Title, NarrativeText, Table, Image, FigureCaption) with tables rendered as HTML that preserves rows and columns. Open a representative PDF from your corpus first: if it contains tables, charts or code blocks, character and sentence splitting will destroy that structure.
Choosing a chunking method
Match the row to your corpus, then read the constraint column — that is usually what decides it.
| Chunking method | Use when | Key consideration |
| Character splitter | Prototyping, or unstructured text | Ignores meaning; splits break mid-sentence |
| Recursive character splitter | Size control with some semantic preservation | Still breaks on dense paragraphs |
| Sentence splitter | Legal docs, FAQs, complete sentences matter | Variable sizes may exceed token limits |
| Semantic splitting | High-value docs with clear topic transitions | Expensive; embeds every sentence |
| Document-specific splitting | PDFs with tables, code, mixed content | Complex setup; slower than text splitters |
| Context-enriched chunking | Finance, law, medicine; precision is critical | One LLM call per chunk |
| Late chunking | Documents under 8K tokens, heavy cross-references | Whole document must fit the model context |
| Agentic chunking | Small corpus of extremely valuable documents | Most expensive; $1,000+ per 100K documents |
If two rows fit, test both. The benchmarks disagree with each other often enough that corpus-specific evidence beats any general ranking.
How to measure chunking effectiveness
Chunking is the one RAG parameter you cannot tune by reading the output. A response can look right while half the retrieved context was noise, and it can look wrong because one chunk was cut in the wrong place. Splunk Agent Observability scores retrieval separately from generation on every trace, which is what makes a chunking change measurable rather than a matter of taste.
Three evaluators carry most of the signal:
- Chunk Relevance is binary and per chunk: each retrieved chunk comes back Relevant or Not Relevant. If a query returns nothing Relevant, the chunk boundaries are what failed, not the prompt.
- Context Precision aggregates that into one number for the whole retrieved context, weighted by rank — consistently low Context Precision means your chunks are longer than they need to be.
- Completeness sits on the generation side and catches the opposite failure: a change that raises Chunk Relevance but drops Completeness has usually made chunks too small to carry a whole thought.
from splunk_ao.agent_streams import enable_evaluators
enable_evaluators(
project_name="chunking-sweep",
agent_stream_name="production",
metrics=[
SplunkAOEvaluators.context_precision,
SplunkAOEvaluators.context_relevance,
SplunkAOEvaluators.completeness,
],
)
Two notes. enable_evaluators replaces the whole set each time it runs, so list every evaluator you want active. And Chunk Relevance is computed by prompting an LLM, so budget it on the sweep or a held-out sample rather than every production query. These run on Luna evaluation models, which is what makes it affordable to score every configuration in a sweep instead of sampling.
Measure first, then add complexity
Start with a tuned recursive or token-based splitter and a strong embedding model. That combination captures most of the available gain, and the 2026 benchmarks are consistent on the point: swapping the embedding model moved F1 more than any chunking-strategy difference. Add semantic, context-enriched, late or agentic chunking when Chunk Relevance and Context Precision expose a failure you can name.
Splunk closes that measurement loop. Agent-reliability capabilities in Splunk Agent Observability, putting chunk-level evaluations beside production traces and metrics. Read The Agentic Shift to see how Splunk extends observability to every agent and retriever you ship.
FAQs
Related Articles

Superapps: What They Are and How They Work

The Role of Behavioral Analytics in Cybersecurity
