ClaudeBot: robots token, verification and how to allow or disallow it
ClaudeBot is an AI crawler operated by Anthropic. Its robots.txt token is ClaudeBot, and the vendor documentation we fetched on 2026-09-06 is the source for every statement on this page.
| Robots token | ClaudeBot |
|---|---|
| Vendor | Anthropic |
| Purpose | training corpus |
| Documented verification | none published |
| Vendor documentation last verified | 2026-09-06 |
| Share of read domains disallowing it for / | 17.2% of 128 |
| Last verified |
Identity
| Robots token | ClaudeBot |
|---|---|
| Matched as | claudebot |
| Vendor | Anthropic |
| Purpose | training corpus |
| Documentation | https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler |
| Documented verification | none published |
Vendor text: 'ClaudeBot helps enhance the utility and safety of our generative AI models by collecting web content that could potentially contribute to their training.' No published IP ranges and no reverse DNS convention were found on 2026-09-06, so a hit carrying this token is recorded as unverifiable, never as verified (same rule as app/botverify.py).
Allow or disallow it
User-agent: ClaudeBot Allow: /
User-agent: ClaudeBot Disallow: /
A robots.txt rule is a request that a well behaved crawler honours. It is not an access control, and this page does not tell anyone what to choose.
Questions
Does ClaudeBot obey robots.txt?
The vendor documentation at https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler states that ClaudeBot respects robots.txt. We fetched that page and it answered our checker on 2026-09-06. This records what the vendor documents, not what any individual request did.
How do I allow ClaudeBot in robots.txt?
Add this group to the robots.txt at the root of the host: User-agent: ClaudeBot Allow: /. The token is matched case insensitively as a substring of the user agent by RFC 9309, and the longest matching rule wins.
How do I disallow ClaudeBot in robots.txt?
Add this group to the robots.txt at the root of the host: User-agent: ClaudeBot Disallow: /. A robots.txt rule is a request that a well behaved crawler honours; it is not an access control.
How do I verify a request really came from ClaudeBot?
We have not found a documented verification method for ClaudeBot: the vendor publishes neither an IP range file nor a reverse DNS convention that we could fetch. Without one, a user agent string is not evidence of origin.
What we measure
Of the 128 seeded domains whose robots.txt we have read, 22 disallow ClaudeBot for / and 106 allow it, which is 17.2% disallowed, week 2026-W37. Counts only: no domain is named.
The weekly AI crawler access index
Sources
| Source | Type | HTTP | Verified | Note |
|---|---|---|---|---|
| https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler | vendor | 200 | 2026-09-06 | URL taken from the support.claude.com sitemap, not guessed; 12 occurrences of ClaudeBot and the exact snippet 'User-agent: ClaudeBot Disallow: /' in the body |
| https://darkvisitors.com/agents/claudebot | aggregator | 200 | 2026-09-06 | third-party aggregator, redirects to knownagents.com; corroboration only; volatile source: identical byte count but a different content hash on every fetch (verified twice on 2026-09-06), so the refresher compares byte counts for it, not hashes |
Free data
Check a domain against every token
AEO Watch is an independent, factual monitor. It is not affiliated with, endorsed by or speaking for any crawler vendor. Every statement is an observation with the date it was made and the raw evidence behind it: a robots.txt line, an HTTP status code, a header. There are no scores, no grades and no verdicts here, and nothing on this page is advice. A site opts out at any time and the opt out is honoured automatically and permanently.