Advanced Chunking Techniques: What The Benchmarks Actually Show

Artificial Intelligence Pratik Bhavsar

Key takeaways

  • Chunking impacts retrieval quality, vector database costs, and generation latency simultaneously; size should be tuned based on the specific document structure and the length of a complete answer within your corpus.
  • Advanced techniques like context-enriched or late chunking should be reserved for specific, identified failures, as simpler strategies (like recursive splitting) often yield the majority of performance gains.
  • Quantifiable performance monitoring using evaluators like Chunk Relevance and Context Precision is necessary to distinguish between retrieval bottlenecks and generation errors in production.

Your retriever returns ten chunks, your retrieval-augmented generation (RAG) app still hallucinates, and doubling chunk size doubles your time to first token. Chunk size is the one parameter that touches retrieval quality, vector storage bills, and generation latency at once.

What is chunking?

Chunking breaks text into smaller units. Each chunk is vectorized and stored, and that unit decides what your retriever can return, what your vector database costs, and how much context your large language model (LLM) reads.

In this screenshot, ChunkViz renders how different splitters cut the same paragraph:

This image shows the document chunking process: one source document becomes many indexed units, and only a few come back per query.

The role of chunking

Chunk size is not one decision with one right answer. Chunk size is a single number pulling against four constraints at once, and they do not agree: retrieval quality wants chunks sized to your answers, storage cost wants them large, latency wants them small and few, and hallucination risk wants fewer of them in the prompt.

Knowing which constraint binds hardest in your system is most of the work.

Retrieval quality

There is no universally correct chunk size. There is a size that matches how long a complete answer runs in your corpus. A 2025 token-size sweep makes the point by contradiction: 64-token chunks won on SQuAD, where an answer is a phrase, while NarrativeQA kept improving all the way to 1,024 tokens because its answers only make sense with surrounding narrative. Same benchmark, same models, opposite conclusions. Before you tune anything, pull ten real answers from your own corpus and note how much surrounding text a person needs to justify each one. That length is your starting chunk size.

Vector database cost and query latency

Halving your chunk size doubles your vector count, and cost scales with that count rather than with the size of your corpus. Qdrant's sizing guidance puts full-precision RAM at num_vectors × dimensions × 4 bytes × 1.5, and Pinecone serverless charges 1 read unit per GB of namespace scanned. The same corpus at 256 tokens per chunk therefore costs roughly four times as much to query as it does at 1,024. Do that multiplication before a chunking sweep rather than after: a change you make for retrieval reasons lands on the infrastructure bill.

LLM latency and cost

Chunk size and top-k multiply, and most teams tune them in separate meetings. NVIDIA's enterprise RAG sizing guide measured time to first token climbing from about 2.0 seconds at 256-token chunks to 7.1 at 1,024 — and, separately, raising top-k from four to 10 pushed it from 3.6 to 7.9. Move both and you are not adding the two penalties, you are multiplying them. If you have a latency budget, spend it on one lever and hold the other fixed.

LLM hallucinations

The instinct when an answer comes back wrong is to retrieve more. The evidence says that usually makes it worse. The 2025 RAGUARD benchmark found that going from one retrieved document to five raised the recall of misleading content from 21.3% to 44.8%, and a 2025 study measured performance dropping by up to 20% from added documents even with total context length and the position of the relevant passage held constant. Extra context is not free insurance. Fewer, better chunks beat more chunks.

Factors influencing chunking

Four variables should drive the decision, and each one resolves to a rule you can apply today rather than a number to memorize.

Text structure

Split on the boundaries the document already has. Across 1,080 configurations spanning 36 strategies, six domains and five embedding models, paragraph-group chunking roughly doubled ranking quality against fixed character chunking. So, the question is not which algorithm to run, it is whether your documents have real structure to respect.

Embedding model limits

Treat the stated maximum as a ceiling to stay well under, not a target. All-MiniLM-L6-v2 truncates at 256 word pieces, OpenAI's text-embedding-3 models accept 8,192 tokens, and Cohere Embed v4 accepts 128,000 — but Jina AI's long-context study found retrieval quality degrading steadily with length well before any of those limits. A model that accepts 8,192 tokens does not embed 8,192 tokens well. Find the length where your model's retrieval still holds up on your data, then confirm it fits the limit. Not the other way round.

LLM context length

A large window is not a substitute for good retrieval. Windows have grown to 400,000 tokens for GPT-5, yet in the No Literal Matching benchmark, 11 of 13 models claiming 128,000+ context reached only half their short-context performance at 32,000 tokens, and a separate study found RAG with 600-token chunks beating a full 128,000-token context outright. Filling the window is the expensive way to get a worse answer. Read the window as the ceiling on chunk size multiplied by top-k — it bounds your choices, it does not make them for you.

Query type

The shape of the question sets the size of the chunk. NVIDIA's chunking study across five datasets found factoid queries favoring 256–512-token chunks while analytical queries favored 1,024-token or page-level chunks. Most production systems get both, and a compromise size in the middle tends to serve neither well. If you can route by query type, do that. If you cannot, size for whichever type carries your traffic and accept that the other will underperform — deliberately, rather than by accident.

How to determine your own requirements first

Before testing a single splitter, answer five questions about your corpus. The answers rule strategies out rather than ranking them, which narrows the field faster than any benchmark table.

Requirement Your answer
Document structure Consistent headings / Mixed formats / Unstructured
Dominant query type Factual / Analytical / Multi-hop
Required metadata Source, temporal, entities, other
Embedding model limit 512 / 1,024 / 8,192 tokens
Average paragraph length X tokens

Consistent headings plus analytical queries points at structure-aware chunking. Unstructured documents plus factual queries points at smaller fixed-size chunks with semantic boundaries. Either way you now have two candidates to test instead of eight.

Character and recursive character splitting

Managed platforms disagree on defaults: Bedrock Knowledge Bases use ~300 tokens with 20% overlap, Azure AI Search recommends 512 with 25%, and Vertex AI RAG Engine defaults to 1,024 with 256 overlap. All LangChain splitters share the TextSplitter base class, now in the langchain-text-splitters package (v1.1.2).

CharacterTextSplitter splits on one separator and merges the results, ignoring where sentences end.

Character splitting: a fixed character count, applied without regard for where the sentences are.

RecursiveCharacterTextSplitter tries separators in order — paragraph, line, space, character — and recurses into any split still larger than chunk_size. LangChain's docs call it the recommended splitter for generic text.

Recursive splitting falls back only as far as it has to, so no chunk straddles a paragraph break.

Chroma's token-level chunking evaluation reported that “the heuristic RecursiveCharacterTextSplitter with chunk size 200 and no overlap performs well,” with 88.1 recall at n=5. If retrieval quality is poor and you are using character splitting, look at where the splits land before changing anything else.

Sentence and semantic splitting

Sentence splitting uses spaCy's sents to find boundaries, then groups whole sentences by stride and overlap — useful where sentence integrity beats uniform size, like FAQ pairs and legal clauses. spaCy 3.8's en_core_web_sm ships a senter component the spaCy docs describe as roughly 10x faster than the parser.

Sentence splitting: identify boundaries first, then group whole sentences by stride and overlap.

Semantic splitting goes further, embedding each sentence and cutting where cosine similarity to the previous one drops below a threshold (credit: agamm/semantic-split).

Semantic splitting scores similarity between consecutive sentences and cuts where it drops below the threshold.

Qu et al. (2025) tested semantic chunkers against fixed-size across 10 datasets. Fixed-size won on three of five evidence-retrieval datasets, and swapping the embedding model produced a 7.44% average F1 gain — larger than any chunking-strategy difference: “The computational costs associated with semantic chunking are not justified by consistent performance gains.” An August 2026 evaluation reached the same verdict for semantic, contextual, and late chunking.

Price it before you commit. A 10,000-document corpus averaging 100 sentences per document is a million embedding calls just to chunk — around $100 of preprocessing before you have indexed anything.

Context-enriched, late and agentic chunking

Three strategies give a chunk awareness of its surroundings, and they pay for it in different currencies.

LLM-based chunking

When coreferences keep atomic facts from retrieving cleanly, Dense X Retrieval introduces “propositions”: atomic, self-contained factoids with coreferences resolved. Proposition-level indexing improved Recall@20 by +10.1 over passages for unsupervised dense retrievers. The authors' propositionizer model is still on Hugging Face. Two caveats:

Multi-vector indexing searches on a vector derived from something other than the raw text, then returns the original document: embed child chunks and return the parent, embed a generated summary, or embed hypothetical questions the document would answer.

Single-vector versus multi-vector embeddings: one vector per document, or one per token with late interaction at query time.

Document-specific splitting

For mixed-format “Chat with PDF,” parsing counts as chunking. Unstructured partitions document, spreadsheet, presentation, image, JSON and audio formats across four partitioning strategies. Note that infer_table_structure now defaults to False, so turn it on for table extraction.

This returns typed elements (Title, NarrativeText, Table, Image, FigureCaption) with tables rendered as HTML that preserves rows and columns. Open a representative PDF from your corpus first: if it contains tables, charts or code blocks, character and sentence splitting will destroy that structure.

Choosing a chunking method

Match the row to your corpus, then read the constraint column — that is usually what decides it.

Chunking method Use when Key consideration
Character splitter Prototyping, or unstructured text Ignores meaning; splits break mid-sentence
Recursive character splitter Size control with some semantic preservation Still breaks on dense paragraphs
Sentence splitter Legal docs, FAQs, complete sentences matter Variable sizes may exceed token limits
Semantic splitting High-value docs with clear topic transitions Expensive; embeds every sentence
Document-specific splitting PDFs with tables, code, mixed content Complex setup; slower than text splitters
Context-enriched chunking Finance, law, medicine; precision is critical One LLM call per chunk
Late chunking Documents under 8K tokens, heavy cross-references Whole document must fit the model context
Agentic chunking Small corpus of extremely valuable documents Most expensive; $1,000+ per 100K documents

If two rows fit, test both. The benchmarks disagree with each other often enough that corpus-specific evidence beats any general ranking.

How to measure chunking effectiveness

Chunking is the one RAG parameter you cannot tune by reading the output. A response can look right while half the retrieved context was noise, and it can look wrong because one chunk was cut in the wrong place. Splunk Agent Observability scores retrieval separately from generation on every trace, which is what makes a chunking change measurable rather than a matter of taste.

Three evaluators carry most of the signal:

title
label
type
python
snippet
from splunk_ao import SplunkAOEvaluators
from splunk_ao.agent_streams import enable_evaluators

enable_evaluators(
project_name="chunking-sweep",
agent_stream_name="production",
metrics=[
SplunkAOEvaluators.context_precision,
SplunkAOEvaluators.context_relevance,
SplunkAOEvaluators.completeness,
],
)
showcopybutton
true

Two notes. enable_evaluators replaces the whole set each time it runs, so list every evaluator you want active. And Chunk Relevance is computed by prompting an LLM, so budget it on the sweep or a held-out sample rather than every production query. These run on Luna evaluation models, which is what makes it affordable to score every configuration in a sweep instead of sampling.

Measure first, then add complexity

Start with a tuned recursive or token-based splitter and a strong embedding model. That combination captures most of the available gain, and the 2026 benchmarks are consistent on the point: swapping the embedding model moved F1 more than any chunking-strategy difference. Add semantic, context-enriched, late or agentic chunking when Chunk Relevance and Context Precision expose a failure you can name.

Splunk closes that measurement loop. Agent-reliability capabilities in Splunk Agent Observability, putting chunk-level evaluations beside production traces and metrics. Read The Agentic Shift to see how Splunk extends observability to every agent and retriever you ship.

FAQs

How should teams determine the optimal chunk size for their RAG system?
Optimal chunk size is determined by analyzing the corpus to match the amount of text required to contain a complete, standalone answer, while also factoring in infrastructure storage costs and system latency limits.
Why is paragraph-aware splitting often superior to fixed-size chunking?
Paragraph-aware splitting respects natural document boundaries that authors have already established as meaningful, leading to higher retrieval ranking quality compared to arbitrary character-based splits that can fragment thoughts.
What are the primary operational risks of using semantic or agentic chunking?
Semantic and agentic chunking methods are computationally expensive, often costing hundreds of dollars per corpus in extra embedding calls while failing to provide consistent performance improvements over simpler, structure-aware recursive splitting strategies.
How do latency budgets dictate the selection of chunking and retrieval parameters?
Latency budgets establish an upper bound on system performance that forces teams to balance chunk size against the number of retrieved documents, preventing the compounding delays caused by increasing both parameters simultaneously.
What evaluators identify if chunking is the root cause of a RAG failure?
Evaluators such as Chunk Relevance and Context Precision isolate the retrieval phase, allowing teams to verify if failures result from poor boundary definitions or if the underlying issues lie elsewhere in the generation logic.

Related Articles

Superapps: What They Are and How They Work
Learn
4 Minute Read

Superapps: What They Are and How They Work

Approving and adopting new enterprise software is hard — superapps simplify the process by putting everything in one box, an ecosystem of tools in one interface
The Role of Behavioral Analytics in Cybersecurity
Learn
7 Minute Read

The Role of Behavioral Analytics in Cybersecurity

Analyzing behaviors has a lot of use cases. In this article, we are hyper-focused on using BA for the cybersecurity of your enterprise. Learn all about BA here.
Blacklist & Whitelist: Terms To Avoid
Learn
5 Minute Read

Blacklist & Whitelist: Terms To Avoid

In this article, we will dive into why “blacklist” and “whitelist” are not inclusive terms and explore potential alternatives that can promote a more inclusive language.