AI Summary:
llms-full.txtis an aggregated context bundle hosted at/llms-full.txt. It concatenates your project's essential documentation, API specifications, and type definitions into a single, clean Markdown stream. This allows frontier reasoning models to ingest complete technical documentation in a single prompt without performing multi-hop HTTP requests.
The Paradigm Shift: From Search to Massive Context Ingestion
In early RAG (Retrieval-Augmented Generation) architectures, models were constrained by 4k to 32k context windows. Engineers were forced to rely on vector databases, semantic search, and cosine similarity chunking. However, this approach often broke down for technical codebases:
- Vector search retrieves disconnected paragraphs, stripping method signatures of their enclosing class context.
- Cross-module dependencies and inheritance chains fail to resolve when chunks are evaluated in isolation.
- Agents make dozens of sequential tool calls to explore linked pages, racking up latency and network overhead.
With the advent of frontier models offering 500k to 2M token context windows (such as Claude Opus 5, Sonnet 5, GPT-6 Astra, and Gemini 3.0 Ultra), developers can pass entire documentation suites directly into context. The llms-full.txt specification creates a standard contract for this bundle.
Document Assembly Protocol & Delimiter Standard
A well-architected llms-full.txt document must preserve the source provenance of every concatenated document. If an agent cannot trace an API definition back to its canonical live URL, it cannot verify freshness or reference the source when explaining code to an engineer.
Every included page must be separated by an explicit horizontal rule, an H2 document header, a source metadata block, and the sanitized markdown content:
# Acme Platform — Complete Technical Reference
> Consolidated documentation and API specifications. Generated 2026-09-10T12:00:00Z.
> Total Tokens: ~42,500 (BPE o200k / Claude Tokenizer)
---
## Getting Started: Core Concepts & Setup
- Source URL: https://acme.dev/docs/getting-started
- Last Updated: 2026-09-08
- Checksum: sha256:4a8b1...
Acme is a distributed key-value store with ACID transactions over Raft consensus...
---
## API Reference: Session Management
- Source URL: https://acme.dev/docs/api/session
- Last Updated: 2026-09-05
- Checksum: sha256:7f9e2...
### `POST /v1/sessions`
Creates an authenticated caller session with signed JWT claims...
The Provenance Header Rule
Never omit the - Source URL: and - Last Updated: fields. Autonomous agents use these lines to detect whether the captured content conflicts with a newer version retrieved via real-time web search or local git diffs.
Token Economics: Maximizing Prompt Caching
Serving a 100,000-token llms-full.txt file might seem cost-prohibitive at first glance. However, modern frontier LLM providers (Anthropic, Google, and OpenAI) feature Prompt Caching:
Anthropic Claude API Cost Model (Example):
- Uncached Prompt Tokens: $3.00 / million tokens
- Cached Write Tokens: $3.75 / million tokens (once per 5-min TTL)
- Cache Read Hit Tokens: $0.30 / million tokens (90% discount!)
When an agent like Claude Code or Cursor reads a static llms-full.txt file at the start of a conversation, that 50k token block gets written to the provider's KV cache. Every subsequent question, refactoring command, or test generation run hits the cache, achieving:
- Sub-second Time-to-First-Token (TTFT): Bypasses model prefill latency.
- 90% Cost Reduction: Makes deep context ingestion cheaper than issuing multiple fragmented API requests.
Critical Caching Rule
Keep the top portion of llms-full.txt deterministic and unchanging. Dynamic timestamps or volatile counters placed at the very top of the file invalidate prefix hashes and destroy prompt cache hit rates!
Comparison Matrix: Bundle vs Incremental Index
| Dimension | llms-full.txt (Full Bundle) | llms.txt (Curated Map) | Vector RAG Embeddings |
|---|---|---|---|
| Ingestion Pattern | Zero-shot single context load | Multi-hop sequential HTTP fetch | Embedding query + Top-K chunk retrieval |
| Token Footprint | 40,000 – 250,000+ tokens | 1,000 – 3,500 tokens | 1,500 – 8,000 tokens per query |
| Architectural Coherence | Complete cross-module visibility | High (agent chooses relevant URLs) | Fragmented (loses holistic AST view) |
| Freshness Maintenance | Automated CI nightly build | Low maintenance / static links | Complex vector index sync pipelines |
| Cache Efficiency | Maximum (leverages KV prompt cache) | Low per-hop caching | Variable |
Production Pipeline: Automating Bundle Generation
Do not maintain llms-full.txt by hand. Implement a headless CI build step (such as via Turndown or Cheerio) that executes during your documentation deployment:
// scripts/build-llms-full.ts
import fs from 'node:fs'
import path from 'node:path'
async function buildFullContext() {
const manifest = JSON.parse(fs.readFileSync('docs-manifest.json', 'utf8'))
const chunks: string[] = [
'# Acme Platform — Complete Technical Reference\n',
`> Automated build: ${new Date().toISOString()}\n\n`,
]
for (const page of manifest.pages) {
const rawMarkdown = fs.readFileSync(path.join('content', page.path), 'utf8')
chunks.push(`---\n## ${page.title}\n- Source URL: ${page.canonicalUrl}\n\n${rawMarkdown}\n\n`)
}
fs.writeFileSync('public/llms-full.txt', chunks.join(''), 'utf8')
console.log('Successfully generated public/llms-full.txt')
}
buildFullContext()
Related guidance
For developers building high-efficiency agent loops, review the concise llms.txt map guide for fast navigation, master Claude Code context optimization for token limits, and establish strict token budgeting limits.
References
- The /llms.txt file, v2: the proposal's guidance on concise maps, source links, Markdown versions, and context limits.
- Anthropic Prompt Caching Architecture Guide: Detailed technical breakdown of KV cache lifecycles, prefix matching, and tokenomics.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.