Which AI crawlers are there and how do I allow or block each one? 21 documented tokens, each dated on its own row
21 AI crawler tokens from 11 vendors are documented here, each with the purpose its vendor's own documentation states, the documented way to verify a request came from it, the robots.txt line that names it, and the share of the domains we have read a robots.txt for that disallow it for /, counted against the domains measured for that token. Each row carries the date its own documentation was last read; the newest is 2026-09-17.
| Documented tokens | 21 |
|---|---|
| Vendors | 11 |
| Newest documentation check | 2026-09-17 |
| Tokens with a weekly index row | 21 |
| Domains read this week, all tokens | 12585 |
| Week | 2026-W41 |
| Last verified |
The documented tokens
Ordered by the share of read domains that disallow the token for /, the same order the weekly index uses. Each token links to its own guide, which cites the vendor documentation and the date we read it.
| Token | Vendor | Purpose | Documented verification | Documentation read | Share of read domains disallowing it for / |
|---|---|---|---|---|---|
| CCBot | Common Crawl Foundation | training corpus, public web corpus | none published | 13.7% of 12585 | |
| Bytespider | ByteDance | training corpus | none published | 13.2% of 12585 | |
| GPTBot | OpenAI | training corpus | published IP ranges | 13.2% of 12585 | |
| ClaudeBot | Anthropic | training corpus | published IP ranges | 12.1% of 12585 | |
| meta-externalagent | Meta | training corpus, search index | none published | 10.7% of 12585 | |
| Google-Extended | training corpus, grounding | published IP ranges | 10.5% of 12585 | ||
| Amazonbot | Amazon | search index, user-initiated fetch | reverse DNS | 10.5% of 12585 | |
| Applebot-Extended | Apple | training corpus | reverse DNS | 10.4% of 12585 | |
| anthropic-ai | Anthropic | undocumented | none published | 9.1% of 12585 | |
| PerplexityBot | Perplexity | search index | published IP ranges | 8.1% of 12585 | |
| ChatGPT-User | OpenAI | user-initiated fetch | published IP ranges | 7.6% of 12585 | |
| DuckAssistBot | DuckDuckGo | user-initiated fetch, answer generation | none published | 5.9% of 12585 | |
| OAI-SearchBot | OpenAI | search index | published IP ranges | 5.9% of 12585 | |
| Perplexity-User | Perplexity | user-initiated fetch | published IP ranges | 5.8% of 12585 | |
| Claude-User | Anthropic | user-initiated fetch | published IP ranges | 5.6% of 12585 | |
| Claude-SearchBot | Anthropic | search index | published IP ranges | 5.5% of 12585 | |
| Amzn-SearchBot | Amazon | search index | published IP ranges | 4.3% of 12373 | |
| Amzn-User | Amazon | user-initiated fetch | published IP ranges | 4.2% of 12373 | |
| Applebot | Apple | search index, user-initiated fetch | reverse DNS | 4.0% of 12585 | |
| Bingbot | Microsoft | search index | published IP ranges | 2.5% of 12585 | |
| Googlebot | search index | published IP ranges and reverse DNS | 2.1% of 12585 |
One robots.txt block for every token
# Every AI crawler token in the AEO Watch registry: 21 tokens, newest documentation check 2026-09-17; each token's own date is on https://bikoosh.com/aeo/crawlers. # A robots.txt rule is a request, not an access control. # Delete the User-agent line of any crawler you want to keep reading the site. User-agent: CCBot User-agent: Bytespider User-agent: GPTBot User-agent: ClaudeBot User-agent: meta-externalagent User-agent: Google-Extended User-agent: Amazonbot User-agent: Applebot-Extended User-agent: anthropic-ai User-agent: PerplexityBot User-agent: ChatGPT-User User-agent: DuckAssistBot User-agent: OAI-SearchBot User-agent: Perplexity-User User-agent: Claude-User User-agent: Claude-SearchBot User-agent: Amzn-SearchBot User-agent: Amzn-User User-agent: Applebot User-agent: Bingbot User-agent: Googlebot Disallow: /
A robots.txt rule is a request. Whether a crawler honours it is what that vendor's own documentation says, and every token above links to a guide page that cites the document and the date we read it. A rule is not an access control: it does not stop a request from arriving, and this page does not tell anyone what to choose. The block names every token in the table above, the ones whose documented purpose is a search index included, because the table is the whole registry. The allow and the disallow snippet for a single token sit on its guide page.
Allow some tokens and disallow others: the robots.txt builder
Quick answers
How many sites block CCBot?
13.7% of the 12585 domains whose robots.txt we have read disallow CCBot for /, week 2026-W41: 1721 of 12585. Counts only, no domain is named.
How many sites block Bytespider?
13.2% of the 12585 domains whose robots.txt we have read disallow Bytespider for /, week 2026-W41: 1667 of 12585. Counts only, no domain is named.
How many sites block GPTBot?
13.2% of the 12585 domains whose robots.txt we have read disallow GPTBot for /, week 2026-W41: 1663 of 12585. Counts only, no domain is named.
How many sites block ClaudeBot?
12.1% of the 12585 domains whose robots.txt we have read disallow ClaudeBot for /, week 2026-W41: 1519 of 12585. Counts only, no domain is named.
How many sites block meta-externalagent?
10.7% of the 12585 domains whose robots.txt we have read disallow meta-externalagent for /, week 2026-W41: 1343 of 12585. Counts only, no domain is named.
How many sites block Google-Extended?
10.5% of the 12585 domains whose robots.txt we have read disallow Google-Extended for /, week 2026-W41: 1325 of 12585. Counts only, no domain is named.
How many sites block Amazonbot?
10.5% of the 12585 domains whose robots.txt we have read disallow Amazonbot for /, week 2026-W41: 1323 of 12585. Counts only, no domain is named.
How many sites block Applebot-Extended?
10.4% of the 12585 domains whose robots.txt we have read disallow Applebot-Extended for /, week 2026-W41: 1305 of 12585. Counts only, no domain is named.
How many sites block anthropic-ai?
9.1% of the 12585 domains whose robots.txt we have read disallow anthropic-ai for /, week 2026-W41: 1150 of 12585. Counts only, no domain is named.
How many sites block PerplexityBot?
8.1% of the 12585 domains whose robots.txt we have read disallow PerplexityBot for /, week 2026-W41: 1016 of 12585. Counts only, no domain is named.
How many sites block ChatGPT-User?
7.6% of the 12585 domains whose robots.txt we have read disallow ChatGPT-User for /, week 2026-W41: 957 of 12585. Counts only, no domain is named.
How many sites block DuckAssistBot?
5.9% of the 12585 domains whose robots.txt we have read disallow DuckAssistBot for /, week 2026-W41: 744 of 12585. Counts only, no domain is named.
How many sites block OAI-SearchBot?
5.9% of the 12585 domains whose robots.txt we have read disallow OAI-SearchBot for /, week 2026-W41: 744 of 12585. Counts only, no domain is named.
How many sites block Perplexity-User?
5.8% of the 12585 domains whose robots.txt we have read disallow Perplexity-User for /, week 2026-W41: 730 of 12585. Counts only, no domain is named.
How many sites block Claude-User?
5.6% of the 12585 domains whose robots.txt we have read disallow Claude-User for /, week 2026-W41: 702 of 12585. Counts only, no domain is named.
How many sites block Claude-SearchBot?
5.5% of the 12585 domains whose robots.txt we have read disallow Claude-SearchBot for /, week 2026-W41: 692 of 12585. Counts only, no domain is named.
How many sites block Amzn-SearchBot?
4.3% of the 12373 domains whose robots.txt we have read disallow Amzn-SearchBot for /, week 2026-W41: 534 of 12373. Counts only, no domain is named.
How many sites block Amzn-User?
4.2% of the 12373 domains whose robots.txt we have read disallow Amzn-User for /, week 2026-W41: 518 of 12373. Counts only, no domain is named.
How many sites block Applebot?
4.0% of the 12585 domains whose robots.txt we have read disallow Applebot for /, week 2026-W41: 505 of 12585. Counts only, no domain is named.
How many sites block Bingbot?
2.5% of the 12585 domains whose robots.txt we have read disallow Bingbot for /, week 2026-W41: 313 of 12585. Counts only, no domain is named.
How many sites block Googlebot?
2.1% of the 12585 domains whose robots.txt we have read disallow Googlebot for /, week 2026-W41: 259 of 12585. Counts only, no domain is named.
Where these numbers come from
- The weekly AI crawler access index, week 2026-W41: the share column above, with the chart and the history.
- One guide page per token, linked from the table: the vendor documentation we fetched, its HTTP status and the date we last read it.
- The AEO Watch dataset
- Tranco top 1,000 AI crawler policy (CSV, CC BY 4.0)
- The AEO Watch API
- The checker as a command line tool (one file, no install)
Check your own site
The same checks, run on your domain now: what robots.txt tells each AI crawler, whether the edge answers them, and whether the text is readable without JavaScript. Free, no signup.
Limits: 8 checks and 5 different domains per address per hour; a repeat of the same domain inside 7 days is answered from what we already hold.
AEO Watch is an independent, factual monitor. It is not affiliated with, endorsed by or speaking for any crawler vendor. Every statement is an observation with the date it was made and the raw evidence behind it: a robots.txt line, an HTTP status code, a header. There are no scores, no grades and no verdicts here, and nothing on this page is advice. A site opts out at any time and the opt out is honoured automatically and permanently.