AI Summary: Serving a pre-rendered static
llms.txtfile from a CDN provides zero-latency responses and infinite scaling for public open-source libraries. Conversely, deploying a dynamic Edge Worker allows enterprise platforms to tailor token volume to the requesting model, inject tenant-specific API keys, and enforce RBAC authentication boundaries.
When publishing machine-readable documentation, engineering teams encounter an immediate infrastructure bifurcation: Should llms.txt be a plain file sitting in public/llms.txt generated at build time, or should it be a dynamic HTTP route executed on Cloudflare Workers, Fastly Compute, or Next.js Edge?
While static deployment provides unmatched reliability, enterprise platforms requiring multi-tenant isolation, real-time feature gating, and token-aware context slicing require dynamic edge generation.
Architectural Trade-Off Matrix
Evaluating static build-time generation against dynamic edge compute:
| Operational Metric | Static Build (public/llms.txt) | Dynamic Edge Worker (app/llms.txt/route.ts) | Architectural Decision Driver |
|---|---|---|---|
| Response Latency | ~10ms – 25ms (Global CDN Edge) | ~45ms – 120ms (Edge Compute execution) | Target latency tolerance |
| Hosting Cost | Effectively $0 (served as static asset) | Micro-costs per request ($0.15–$0.50/M reqs) | Traffic volume and scale |
| Tenant Isolation | None (identical public document for all callers) | High (injects tenant-specific scopes and endpoints) | Multi-tenant SaaS requirements |
| Model Token Tailoring | Fixed token count (e.g. always 8,000 tokens) | Dynamic slicing (e.g. 2k for mini models, 50k for Opus) | Consumer diversity optimization |
| Deployment Dependency | Requires full CI build to update docs | Instant updates via KV / Redis sync or webhooks | Documentation update velocity |
| Authentication Enforcement | None (publicly reachable) | Bearer token / API Key validation at the edge | Intellectual property and API security |
The Mechanics of Dynamic Context Slicing
The primary limitation of a static llms.txt is that it cannot adapt to the consumer. If an autonomous agent operating on a small reasoning model (such as Claude Haiku or GPT-4o-mini with a strict token budget) fetches a 12,000-token static manifest, it expends 80% of its working memory on the index alone.
A dynamic edge worker inspects request headers and query parameters to dynamically slice the documentation AST:
// app/llms.txt/route.ts (Next.js App Router Edge Runtime)
import { NextRequest, NextResponse } from "next/server";
import { getDocIndexFromKV } from "@/lib/kv";
export const runtime = "edge";
export async function GET(req: NextRequest) {
const modelHeader = req.headers.get("X-LLM-Model") || "default";
const authHeader = req.headers.get("Authorization");
// Determine token budget based on caller model
const maxTokens = modelHeader.includes("mini") || modelHeader.includes("haiku")
? 2000
: 15000;
// Verify enterprise tenant permissions if bearer token is present
const isEnterprise = authHeader ? await verifyTenantAuth(authHeader) : false;
// Retrieve cached documentation manifest from KV
const rawManifest = await getDocIndexFromKV();
// Slice AST to fit budget and permission boundaries
const tailoredManifest = sliceManifestToBudget(rawManifest, {
maxTokens,
includeInternalAPIs: isEnterprise,
});
return new NextResponse(tailoredManifest, {
status: 200,
headers: {
"Content-Type": "text/markdown; charset=utf-8",
"Cache-Control": "private, no-cache, no-store", // Dynamic responses avoid shared CDN caching
"Vary": "X-LLM-Model, Authorization",
},
});
}
When to Use Static vs Dynamic Architecture
# Use Static Build-Time Deployment When:
1. The documentation is 100% open source and public.
2. The codebase changes via Git commits and deploys via GitHub Actions.
3. Maximum traffic resilience (millions of scraping requests during major releases) is required.
4. Infrastructure budget is zero.
# Use Dynamic Edge-Generated Architecture When:
1. Documentation includes customer-specific API keys and custom sandbox URLs.
2. Enterprise tiering hides advanced enterprise SDK documentation behind auth.
3. Real-time feature flags toggle documentation visibility based on deployed feature rollout percentages.
4. Content is pulled from a headless CMS (Contentful, Sanity) where Git rebuilds are impractical.
Hybrid Caching Strategy: The Best of Both Worlds
To achieve dynamic tailoring without paying computation overhead on every request, high-scale architectures employ Edge KV with Cache Tags:
flowchart LR
Caller[AI Agent / Spider] --> Edge[Edge Worker Router]
Edge --> CheckCache{In Edge KV Cache?}
CheckCache -->|Hit| FastResponse[Return Cached Markdown (~15ms)]
CheckCache -->|Miss| Rebuild[Slice AST & Store in KV (~80ms)]
Rebuild --> FastResponse
Security Best Practices and Hard Negative Constraints
- Never Cache Authenticated Responses in Public CDN Edges: If your dynamic
llms.txtinjects customer-specific endpoints or tokens, theCache-Controlheader must strictly specifyprivateorno-storeto prevent cache leakage across tenants. - Deterministic Fallback to Static Root: Ensure that if the Edge KV or database fails, the worker gracefully falls back to a pre-bundled static public
llms.txtfile rather than returning an HTTP 500 error to the agent. - Rate Limiting by Requester IP and API Key: Autonomous agents can loop infinitely during debugging. Enforce rate limits (e.g. 100 requests per minute per IP) to protect edge compute budgets.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify AI.
References
- Cloudflare Workers Documentation: KV Storage: Low-latency global key-value data storage at the edge.
- RFC 9111: HTTP Caching: Authoritative specification for Cache-Control and Vary headers.
- Next.js Edge Runtime Documentation: Standards for running lightweight serverless code at CDN edges.