AI Summary: Automating
llms.txtandllms-full.txtgeneration in CI/CD pipelines ensures that AI coding agents always ingest fresh, valid technical documentation. A robust automation architecture extracts canonical markdown sources, validates all outbound hyperlinks against 404/429 errors, enforces strict token budget gates, and triggers edge CDN cache invalidation upon merge to main.
The Stale Documentation Liability
Maintaining llms.txt and llms-full.txt manually introduces severe operational debt:
- Documentation pages get renamed or deleted, leaving dead links inside
llms.txt. - Developers update API signatures in code but forget to update the context bundle, causing agents to generate deprecated function calls.
- Unmonitored documentation additions cause
llms-full.txtto silently balloon from 40k tokens to 250k tokens, blowing context budgets and degrading model reasoning.
To guarantee documentation integrity, treat llms.txt as a Compiled Build Artifact.
The 4-Stage Automation Pipeline
[Git Commit / PR Trigger]
│
▼
[Stage 1: Extraction & Generation]
• Traverse approved markdown content directories
• Extract H1 titles, descriptions, and canonical URLs
• Compile public/llms.txt and public/llms-full.txt
│
▼
[Stage 2: Validation & Link Checking]
• Verify every link returns HTTP 200 OK (no 404s or circular redirects)
• Validate Markdown syntax against the llms.txt specification
│
▼
[Stage 3: Token Budget Gate]
• Calculate exact BPE token count using native tokenizers
• Fail CI if total bundle exceeds target budget (e.g. > 65,000 tokens)
│
▼
[Stage 4: Edge Cache Invalidation (Post-Deploy)]
• Purge Cloudflare / Fastly CDN cache for /llms.txt and /llms-full.txt
Production GitHub Actions Workflow
# .github/workflows/llms-txt-automation.yml
name: Compile & Validate llms.txt
on:
push:
branches: [main]
paths: ['content/**', 'src/docs/**', 'scripts/build-llms-txt.ts']
pull_request:
paths: ['content/**', 'src/docs/**', 'scripts/build-llms-txt.ts']
jobs:
build-and-verify:
runs-on: ubuntu-latest
steps:
- name: Checkout Repository
uses: actions/checkout@v4
- name: Setup Node.js Runtime
uses: actions/setup-node@v4
with:
node-version: 22
cache: 'pnpm'
- name: Install Dependencies
run: corepack enable && pnpm install --frozen-lockfile
- name: Compile llms.txt Manifests
run: pnpm tsx scripts/build-llms-txt.ts
- name: Validate Outbound Links (Dead Link Detector)
run: pnpm tsx scripts/validate-links.ts
- name: Enforce Token Budget Threshold
run: pnpm tsx scripts/verify-token-budget.ts --max-tokens=65000
- name: Assert Zero Git Working Directory Drift
if: github.event_name == 'pull_request'
run: |
git diff --exit-code public/llms.txt public/llms-full.txt || (
echo "ERROR: Committed llms.txt is out of sync with docs. Run 'pnpm tsx scripts/build-llms-txt.ts' and commit the diff."
exit 1
)
- name: Purge Cloudflare Edge Cache
if: github.ref == 'refs/heads/main' && github.event_name == 'push'
env:
CLOUDFLARE_ZONE_ID: ${{ secrets.CLOUDFLARE_ZONE_ID }}
CLOUDFLARE_API_TOKEN: ${{ secrets.CLOUDFLARE_API_TOKEN }}
run: |
curl -X POST "https://api.cloudflare.com/client/v4/zones/${CLOUDFLARE_ZONE_ID}/purge_cache" \
-H "Authorization: Bearer ${CLOUDFLARE_API_TOKEN}" \
-H "Content-Type: application/json" \
--data '{"files":["https://acme.dev/llms.txt", "https://acme.dev/llms-full.txt"]}'
Dead Link Verification Script (Node.js / TypeScript)
// scripts/validate-links.ts
import fs from 'node:fs'
async function validateLinks() {
const content = fs.readFileSync('public/llms.txt', 'utf8')
const urlMatches = Array.from(content.matchAll(/\[.*?\]\((https?:\/\/.*?)\)/g)).map(m => m[1])
console.log(`Found ${urlMatches.length} links to validate in llms.txt...`)
let hasFailures = false
for (const url of urlMatches) {
try {
const res = await fetch(url, { method: 'HEAD', headers: { 'User-Agent': 'llmstxt-validator/1.0' } })
if (!res.ok) {
console.error(`FAIL: ${url} returned HTTP ${res.status}`)
hasFailures = true
} else {
console.log(`OK (HTTP ${res.status}): ${url}`)
}
} catch (err) {
console.error(`ERROR fetching ${url}:`, err)
hasFailures = true
}
}
if (hasFailures) {
process.exit(1)
}
}
validateLinks()
Automation Failure Mode Checklist
| Risk Area | Mitigation Pattern |
|---|---|
| Silent Token Bloat | Strict CI gate fails pull requests if token count exceeds designated quota |
| Broken Anchor Links | Automated HEAD check validates all target URLs return HTTP 200 OK |
| Stale Edge CDN Cache | Automated Cloudflare API purge triggered on merge to main |
| Dirty Repo Drift | CI asserts git diff --exit-code public/ to catch uncommitted generated files |
Related guidance
To design clean, token-efficient manifests, review llms.txt Core Architecture, understand context sizing in llms-full.txt, and establish strict boundaries with Prohibitions in AGENTS.md.
References
- GitHub Actions Documentation: Official specifications for workflow triggers, secrets, and runners.
- Cloudflare API: Purge Cache by URL: Technical guide on edge cache invalidation.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.