AI Summary: The
ai.txtstandard (formalized under the IETF Internet-Draft track) defines a standardized machine-readable policy format hosted at/.well-known/ai.txtand/.well-known/ai.json. It provides granular policy toggles for AI training, scraping, indexing, attribution, and commercial licensing, addressing capabilities that the legacy 1994 Robots Exclusion Protocol cannot represent.
Why Robots.txt Cannot Express Modern AI Governance
The Robots Exclusion Protocol (robots.txt, standardized in RFC 9309) was architected in 1994 to solve a single engineering challenge: preventing search engine spiders from overwhelming web servers with concurrent socket requests.
In modern AI architectures, however, data consumption takes three distinct forms:
- Foundation Model Pre-training: Massively scraping billions of tokens to train base model weights (e.g. Common Crawl, GPTBot, ClaudeBot).
- Retrieval-Augmented Generation (RAG) & Answer Engines: Reading live documentation to answer an immediate query with direct citation (e.g. Perplexity, ChatGPT Search).
- User-Delegated Autonomous Agents: An authenticated engineer using Cursor or Claude Code to inspect an API schema to write code.
If a company blocks User-agent: * in robots.txt, it inadvertently cuts off developer IDEs and breaks search engine visibility. The ai.txt standard decouples crawling mechanics from intellectual property rights and usage consent.
The Specification: Dual Text & JSON Formats
The IETF draft specifies two canonical formats:
- Plaintext representation:
/.well-known/ai.txt(or/ai.txtfallback) - Structured JSON representation:
/.well-known/ai.json
Production Plaintext Specification (ai.txt)
# /.well-known/ai.txt
Spec-Version: 1.0.0
Site-Name: Acme Technologies Corp
Site-URL: https://acme.dev
Policy-URL: https://acme.dev/legal/ai-terms
Contact: ai-compliance@acme.dev
# Default Organization Policy
Training: deny
Scraping: allow
Indexing: allow
Caching: allow
Attribution: required
License-URL: https://acme.dev/licenses/commercial-eval
# Path-specific Overrides
Scope: /docs/*
Training: deny
Inference: allow
Attribution: required
Scope: /benchmarks/*
Training: deny
Inference: deny
Structured JSON Representation (ai.json)
For programmatic validation in automated crawler pipelines:
{
"$schema": "https://standards.ietf.org/ai-txt/v1.json",
"specVersion": "1.0.0",
"organization": "Acme Technologies Corp",
"policyUrl": "https://acme.dev/legal/ai-terms",
"contact": "ai-compliance@acme.dev",
"defaults": {
"training": false,
"inference": true,
"indexing": true,
"caching": true,
"attributionRequired": true
},
"scopedRules": [
{
"pathPattern": "/docs/**",
"training": false,
"inference": true,
"attributionUrl": "https://acme.dev/docs"
},
{
"pathPattern": "/internal/**",
"training": false,
"inference": false
}
]
}
Matrix of Directives & Behavioral Contracts
| Directive | Permitted Values | Architectural Meaning |
|---|---|---|
Training | allow | deny | Foundation model weight training and fine-tuning permission. |
Inference | allow | deny | Real-time RAG context retrieval and synthetic answer generation. |
Indexing | allow | deny | Vector indexing or lexical search inverted-index construction. |
Caching | allow | deny | Storing prompt cache KV tokens on provider edge infrastructure. |
Attribution | required | optional | Requirement to display brand name and canonical hyperlink in generative output. |
License-URL | Valid HTTPS URL | Machine-actionable link to commercial data licensing terms. |
Legal Standing: EU Copyright Directive Art. 4 & US Fair Use
Publishing an ai.txt file is not merely technical documentation—it constitutes an explicit legal act in major jurisdictions:
- European Union: Article 4 of the Digital Single Market (DSM) Directive 2019/790 permits text and data mining (TDM) unless the rightsholder has expressly reserved their rights in an appropriate manner, such as machine-readable means. An
ai.txtdeclaringTraining: denyfulfills the requirement for machine-readable reservation. - United States: In US copyright litigation regarding fair use, documented machine-readable opt-outs demonstrate that commercial scraping occurred in clear opposition to published publisher terms.
Edge Deployment Architecture
Serve /.well-known/ai.txt and /.well-known/ai.json with immutable cache headers and permissive CORS at your CDN boundary:
// Cloudflare Worker / Next.js Edge Handler
export const runtime = 'edge'
export async function GET(req: Request) {
const isJson = req.url.endsWith('.json')
const contentType = isJson ? 'application/json' : 'text/plain; charset=utf-8'
const body = isJson ? JSON.stringify(AI_POLICY_JSON, null, 2) : AI_POLICY_TEXT
return new Response(body, {
status: 200,
headers: {
'Content-Type': contentType,
'Cache-Control': 'public, max-age=86400, stale-while-revalidate=604800',
'Access-Control-Allow-Origin': '*',
'X-Robots-Tag': 'noindex', // Prevent the policy file itself from ranking in search
},
})
}
Related guidance
To understand how to declare agent-specific navigation maps, read our ai.txt overview, examine What is llms.txt?, and ensure your internal engineering repositories follow AGENTS.md best practices.
References
- IETF Internet-Draft: The ai.txt Specification: Formal technical draft specifying directive parsing, grammar, and well-known URI schemas.
- RFC 9309: Robots Exclusion Protocol: The IETF standard for robots.txt parsing and crawler rules.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.