AI Summary: Semantic chunking is an embedding-driven document segmentation technique that places chunk boundaries at natural topic transitions. By calculating cosine distance between rolling sentence windows, it produces coherent retrieval units that outperform naive token sliding windows while preserving code blocks and conceptual integrity.
The Limitation of Fixed Character & Token Splitting
In first-generation RAG implementations, text splitting relied on fixed token sizes (e.g. 512 tokens with 50-token overlap). On structured technical documentation, fixed splitting introduces severe retrieval defects:
- Topic Smearing: A chunk ends midway through a conceptual explanation, and the second half is appended to an unrelated subsequent topic. The resulting embedding vector sits in an intermediate subspace that matches neither query well.
- Micro-Fragments: Splitting on hard token counts can leave a 3-line code example isolated in a single chunk with zero explanatory context.
- Redundant Embeddings: Overlapping windows generate duplicate vector entries, inflating vector database storage costs and diluting Top-K retrieval results.
Semantic chunking solves this by computing mathematical inflection points in semantic continuity.
The Semantic Distance Algorithm (Kamradt Methodology)
The canonical semantic chunking algorithm evaluates rolling windows of sentences to detect topic shifts:
Document Text ──► Sentence Tokenization [s1, s2, s3, s4, ...]
│
▼
Rolling Window Merging (size=3)
g1 = [s1, s2, s3], g2 = [s2, s3, s4], ...
│
▼
Generate Dense Vector Embeddings
v1 = Embed(g1), v2 = Embed(g2), ...
│
▼
Compute Pairwise Cosine Distance
d_i = 1.0 - CosineSimilarity(v_i, v_{i+1})
│
▼
Calculate Breakpoint Threshold
(e.g., 95th Percentile of Distances)
│
▼
Split Chunks at Points Where d_i > Threshold
# Semantic Chunking Implementation (Python / NumPy)
import numpy as np
def compute_chunk_breakpoints(distances: list[float], percentile: float = 92.0) -> list[int]:
"""Identify split indices where semantic distance spikes above percentile threshold."""
threshold = np.percentile(distances, percentile)
return [i + 1 for i, d in enumerate(distances) if d > threshold]
Comparative Chunking Strategies for Technical Context
| Strategy | Computational Cost | Semantic Coherence | Code & Table Safety | Ideal Workload |
|---|---|---|---|---|
| Fixed Token Window | Zero extra compute (fast) | Poor (arbitrary cuts) | Extremely dangerous (breaks fences) | Plain literary prose |
| Recursive Character / AST | Low (regex / Markdown parser) | High (respects headers/paragraphs) | High (preserves fences & tables) | Standard technical Markdown docs |
| Pure Semantic Chunking | High (requires embedding all sentences) | Very High (groups related thoughts) | Medium (must guard code fences) | Dense, unformatted technical papers |
| Hybrid AST-Semantic Splitter | Moderate (embeds sub-sections only) | Maximum (optimal precision) | Maximum (deterministic code handling) | Enterprise documentation suites |
The Hybrid Recommendation: AST Structure First, Semantic Fallback
For technical documentation indexed in llms.txt and llms-full.txt, pure semantic chunking should never split code blocks or tables. The recommended production pattern is Hybrid AST-Semantic Chunking:
- Pass 1: Markdown AST Splitting: Split the document by Markdown H2 and H3 boundaries using
MarkdownHeaderTextSplitter. - Pass 2: Size & Semantics Verification:
- If a section is under 800 tokens, keep it as a single chunk.
- If an H2 section contains 3,000 tokens of dense prose, run the semantic distance algorithm only within that section's body to locate natural sub-topic breakpoints without crossing into adjacent sections.
Related guidance
To understand how clean Markdown enables reliable chunking, read our guide on Markdown for RAG, review context allocations in Token Budgeting, and evaluate syntax choices in JSON vs Markdown for AI.
References
- Greg Kamradt: 5 Levels of Text Splitting: Benchmark implementations of semantic distance chunking.
- Pinecone: Chunking Strategies for LLM Applications: Comprehensive engineering guide on embedding spaces and boundary decisions.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.