AI Summary:
llms.txtprovides a curated semantic index (~2,000 tokens) directing agents to specific high-value documentation paths, whilellms-full.txtaggregates the entire documentation repository into a single concatenated flat file (100,000+ tokens). The choice depends strictly on whether the consumer is a reasoning agent performing multi-turn tool calling or a prompt-caching pipeline executing one-shot analysis.
The fundamental tension in AI documentation architecture lies between index discovery and monolithic context ingestion. Providing an autonomous agent with a 150,000-token full-text dump when it merely needs to look up an HTTP status code introduces severe attention degradation, multi-dollar token waste, and hallucinated API parameters. Conversely, forcing an agent to execute 20 sequential HTTP fetch loops to traverse a sparse index degrades task completion latency by 400%.
Understanding the architectural boundary between /llms.txt and /llms-full.txt allows infrastructure teams to optimize both agent discovery and prompt-caching economics.
Architectural Trade-Offs: Discovery vs Ingestion
An index file (/llms.txt) and a monolithic bundle (/llms-full.txt) solve entirely different phases of the agent reasoning lifecycle:
| Evaluation Vector | llms.txt (Index Map) | llms-full.txt (Monolith Bundle) | Architectural Implication |
|---|---|---|---|
| Typical Payload Size | 1,500 – 4,000 tokens | 60,000 – 250,000 tokens | 50x–100x token volume difference |
| Primary Consumer | Tool-calling agents (Cursor, Claude Code, Cline) | Prompt-caching batch pipelines, one-shot evaluations | Tool selection vs bulk semantic ingestion |
| Edge Cache Strategy | stale-while-revalidate=86400, short TTL | Heavy CDN caching with content-hash invalidation | Monolith requires deterministic invalidation |
| Attention Accuracy | 99.2% needle-in-haystack retrieval | Degrades past 80k tokens unless prompt-cached | Risk of "Lost in the Middle" errors |
| Network Round-Trips | 1 initial GET + selective follow-up fetches | 1 initial bulk GET, zero follow-up requests | Latency amortized across single request |
| Cost Per Session | $0.003 – $0.012 per tool inspection | $0.05 – $0.45 without prefix prompt caching | Caching viability dictates architecture |
Attention Degradation and "Lost in the Middle"
When evaluating llms-full.txt, systems architects must account for the transformer attention curve. While modern frontier models (such as Claude Opus 5, GPT-6, and Gemini 2.5 Pro) boast 1M to 2M token context windows, empirical retrieval accuracy does not remain uniform across the context space:
- Primacy Bias: Information in the first 10% of
llms-full.txt(project overview and primary installation) receives maximum attention weighting. - Recency Bias: Information in the final 10% (troubleshooting, changelogs) is retained with high fidelity.
- Mid-Context Attenuation: Obscure API parameter configurations buried around token index 65,000 experience up to a 28% drop in zero-shot recall.
For critical API contracts and strict database migration schemas, /llms.txt routing guarantees that only the precise, dedicated document is injected into the model reasoning loop, preventing mid-context attenuation.
# When to Route Agents to llms.txt
- Multi-turn interactive CLI agents (Cursor, Claude Code)
- Codebase repositories with modular documentation (>50 distinct markdown files)
- Environments where token expenditure directly impacts operational margins
# When to Serve llms-full.txt
- Offline code evaluation and continuous compliance audits
- Frontier models with active KV prompt caching enabled
- Small SDKs and libraries under 30,000 total tokens
Prompt Caching Economics: The Deciding Factor
The economic viability of llms-full.txt shifted dramatically with the introduction of provider-level prompt caching:
- Anthropic Claude Cache Hits: Cost $0.30 per million tokens (a 90% discount compared to the $3.00 base input rate).
- OpenAI Prefix Cache Hits: Cost $0.625 per million tokens (a 50% discount).
If your team runs continuous integration agents where 50 developers repeatedly query the documentation within a 5-minute cache lifespan, serving a cached 120,000-token llms-full.txt file costs merely $0.036 per invocation. Under these conditions, the monolithic file eliminates multiple HTTP fetch rounds and provides instant responses without network jitter.
However, if your documentation is accessed infrequently by external web crawlers, paying full input prices on a 200,000-token context payload results in catastrophic billing overhead.
Production Dual-Serving Edge Implementation
Modern documentation platforms should never force an either/or choice. Production infrastructure serves both endpoints concurrently from a shared content lake:
// app/api/ai-context/route.ts (Next.js App Router Edge Runtime)
import { NextRequest, NextResponse } from "next/server";
export const runtime = "edge";
export async function GET(req: NextRequest) {
const url = new URL(req.url);
const wantsMonolith = url.pathname.endsWith("llms-full.txt");
if (wantsMonolith) {
const fullBundle = await generateAggregatedMarkdown({ pruneNav: true });
return new NextResponse(fullBundle, {
status: 200,
headers: {
"Content-Type": "text/markdown; charset=utf-8",
"Cache-Control": "public, max-age=3600, stale-while-revalidate=86400",
"X-Robots-Tag": "noindex",
},
});
}
const indexManifest = await generateCuratedIndex();
return new NextResponse(indexManifest, {
status: 200,
headers: {
"Content-Type": "text/markdown; charset=utf-8",
"Cache-Control": "public, max-age=1800, stale-while-revalidate=43200",
},
});
}
Security Best Practices and Hard Negative Constraints
- Never Concatenate Internal Endpoints into llms-full.txt: Automated build scripts that walk documentation directories must strictly filter out staging URLs, internal admin runbooks, and VPN instructions.
- Deterministic Heading Demarcation: In
llms-full.txt, always separate aggregated files using clear file boundary delimiters (--- [file: /docs/auth.md] ---) to prevent AST parsers from merging adjacent symbol namespaces. - Link Back to llms-full.txt in llms.txt: The curated
llms.txtindex should provide an explicit link tollms-full.txtunder an## Optionalsection, enabling reasoning agents to electively ingest the monolith when local RAG is unavailable.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify AI.
References
- The llms.txt Specification (Jeremy Howard): The original canonical proposal for machine-readable context indexing.
- Anthropic Prompt Caching Documentation: Official specification on cache breakpoints, 5-minute TTLs, and token pricing models.
- Lost in the Middle: How Language Models Use Long Contexts: Liu et al., empirical proof of mid-context recall degradation in transformer architectures.