# Which AI crawlers are there and how do I allow or block each one? 21 documented tokens, each dated on its own row

## In short

21 AI crawler tokens from 11 vendors are documented here, each with the purpose its vendor's own documentation states, the documented way to verify a request came from it, the robots.txt line that names it, and the share of the domains we have read a robots.txt for that disallow it for /, counted against the domains measured for that token. Each row carries the date its own documentation was last read; the newest is 2026-09-17.

| Fact | Value |
| --- | --- |
| Documented tokens | 21 |
| Vendors | 11 |
| Newest documentation check | 2026-09-17 |
| Tokens with a weekly index row | 21 |
| Domains read this week, all tokens | 12585 |
| Week | 2026-W41 |
| Last verified | 2026-09-17 |

## The documented tokens

Ordered by the share of read domains that disallow the token for /, the same order the weekly index uses. Each token links to its own guide, which cites the vendor documentation and the date we read it.

| Token | Vendor | Purpose | Documented verification | Share of read domains disallowing it for / |
|---|---|---|---|---|
| CCBot | Common Crawl Foundation | training corpus, public web corpus | none published | 13.7% of 12585 |
| Bytespider | ByteDance | training corpus | none published | 13.2% of 12585 |
| GPTBot | OpenAI | training corpus | published IP ranges | 13.2% of 12585 |
| ClaudeBot | Anthropic | training corpus | published IP ranges | 12.1% of 12585 |
| meta-externalagent | Meta | training corpus, search index | none published | 10.7% of 12585 |
| Google-Extended | Google | training corpus, grounding | published IP ranges | 10.5% of 12585 |
| Amazonbot | Amazon | search index, user-initiated fetch | reverse DNS | 10.5% of 12585 |
| Applebot-Extended | Apple | training corpus | reverse DNS | 10.4% of 12585 |
| anthropic-ai | Anthropic | undocumented | none published | 9.1% of 12585 |
| PerplexityBot | Perplexity | search index | published IP ranges | 8.1% of 12585 |
| ChatGPT-User | OpenAI | user-initiated fetch | published IP ranges | 7.6% of 12585 |
| DuckAssistBot | DuckDuckGo | user-initiated fetch, answer generation | none published | 5.9% of 12585 |
| OAI-SearchBot | OpenAI | search index | published IP ranges | 5.9% of 12585 |
| Perplexity-User | Perplexity | user-initiated fetch | published IP ranges | 5.8% of 12585 |
| Claude-User | Anthropic | user-initiated fetch | published IP ranges | 5.6% of 12585 |
| Claude-SearchBot | Anthropic | search index | published IP ranges | 5.5% of 12585 |
| Amzn-SearchBot | Amazon | search index | published IP ranges | 4.3% of 12373 |
| Amzn-User | Amazon | user-initiated fetch | published IP ranges | 4.2% of 12373 |
| Applebot | Apple | search index, user-initiated fetch | reverse DNS | 4.0% of 12585 |
| Bingbot | Microsoft | search index | published IP ranges | 2.5% of 12585 |
| Googlebot | Google | search index | published IP ranges and reverse DNS | 2.1% of 12585 |

## One robots.txt block for every token

```
# Every AI crawler token in the AEO Watch registry: 21 tokens, newest documentation check 2026-09-17; each token's own date is on https://bikoosh.com/aeo/crawlers.
# A robots.txt rule is a request, not an access control.
# Delete the User-agent line of any crawler you want to keep reading the site.
User-agent: CCBot
User-agent: Bytespider
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: meta-externalagent
User-agent: Google-Extended
User-agent: Amazonbot
User-agent: Applebot-Extended
User-agent: anthropic-ai
User-agent: PerplexityBot
User-agent: ChatGPT-User
User-agent: DuckAssistBot
User-agent: OAI-SearchBot
User-agent: Perplexity-User
User-agent: Claude-User
User-agent: Claude-SearchBot
User-agent: Amzn-SearchBot
User-agent: Amzn-User
User-agent: Applebot
User-agent: Bingbot
User-agent: Googlebot
Disallow: /
```

A robots.txt rule is a request. Whether a crawler honours it is what that vendor's own documentation says, and every token above links to a guide page that cites the document and the date we read it. A rule is not an access control: it does not stop a request from arriving, and this page does not tell anyone what to choose. The block names every token in the table above, the ones whose documented purpose is a search index included, because the table is the whole registry. The allow and the disallow snippet for a single token sit on its guide page.

Allow some tokens and disallow others: the robots.txt builder: https://bikoosh.com/aeo/robots-builder

## Quick answers

### How many sites block CCBot?

13.7% of the 12585 domains whose robots.txt we have read disallow CCBot for /, week 2026-W41: 1721 of 12585. Counts only, no domain is named.

### How many sites block Bytespider?

13.2% of the 12585 domains whose robots.txt we have read disallow Bytespider for /, week 2026-W41: 1667 of 12585. Counts only, no domain is named.

### How many sites block GPTBot?

13.2% of the 12585 domains whose robots.txt we have read disallow GPTBot for /, week 2026-W41: 1663 of 12585. Counts only, no domain is named.

### How many sites block ClaudeBot?

12.1% of the 12585 domains whose robots.txt we have read disallow ClaudeBot for /, week 2026-W41: 1519 of 12585. Counts only, no domain is named.

### How many sites block meta-externalagent?

10.7% of the 12585 domains whose robots.txt we have read disallow meta-externalagent for /, week 2026-W41: 1343 of 12585. Counts only, no domain is named.

### How many sites block Google-Extended?

10.5% of the 12585 domains whose robots.txt we have read disallow Google-Extended for /, week 2026-W41: 1325 of 12585. Counts only, no domain is named.

### How many sites block Amazonbot?

10.5% of the 12585 domains whose robots.txt we have read disallow Amazonbot for /, week 2026-W41: 1323 of 12585. Counts only, no domain is named.

### How many sites block Applebot-Extended?

10.4% of the 12585 domains whose robots.txt we have read disallow Applebot-Extended for /, week 2026-W41: 1305 of 12585. Counts only, no domain is named.

### How many sites block anthropic-ai?

9.1% of the 12585 domains whose robots.txt we have read disallow anthropic-ai for /, week 2026-W41: 1150 of 12585. Counts only, no domain is named.

### How many sites block PerplexityBot?

8.1% of the 12585 domains whose robots.txt we have read disallow PerplexityBot for /, week 2026-W41: 1016 of 12585. Counts only, no domain is named.

### How many sites block ChatGPT-User?

7.6% of the 12585 domains whose robots.txt we have read disallow ChatGPT-User for /, week 2026-W41: 957 of 12585. Counts only, no domain is named.

### How many sites block DuckAssistBot?

5.9% of the 12585 domains whose robots.txt we have read disallow DuckAssistBot for /, week 2026-W41: 744 of 12585. Counts only, no domain is named.

### How many sites block OAI-SearchBot?

5.9% of the 12585 domains whose robots.txt we have read disallow OAI-SearchBot for /, week 2026-W41: 744 of 12585. Counts only, no domain is named.

### How many sites block Perplexity-User?

5.8% of the 12585 domains whose robots.txt we have read disallow Perplexity-User for /, week 2026-W41: 730 of 12585. Counts only, no domain is named.

### How many sites block Claude-User?

5.6% of the 12585 domains whose robots.txt we have read disallow Claude-User for /, week 2026-W41: 702 of 12585. Counts only, no domain is named.

### How many sites block Claude-SearchBot?

5.5% of the 12585 domains whose robots.txt we have read disallow Claude-SearchBot for /, week 2026-W41: 692 of 12585. Counts only, no domain is named.

### How many sites block Amzn-SearchBot?

4.3% of the 12373 domains whose robots.txt we have read disallow Amzn-SearchBot for /, week 2026-W41: 534 of 12373. Counts only, no domain is named.

### How many sites block Amzn-User?

4.2% of the 12373 domains whose robots.txt we have read disallow Amzn-User for /, week 2026-W41: 518 of 12373. Counts only, no domain is named.

### How many sites block Applebot?

4.0% of the 12585 domains whose robots.txt we have read disallow Applebot for /, week 2026-W41: 505 of 12585. Counts only, no domain is named.

### How many sites block Bingbot?

2.5% of the 12585 domains whose robots.txt we have read disallow Bingbot for /, week 2026-W41: 313 of 12585. Counts only, no domain is named.

### How many sites block Googlebot?

2.1% of the 12585 domains whose robots.txt we have read disallow Googlebot for /, week 2026-W41: 259 of 12585. Counts only, no domain is named.

- Every guide page: https://bikoosh.com/aeo/crawler/<token>
- The weekly access index: https://bikoosh.com/aeo/index

AEO Watch is an independent, factual monitor. It is not affiliated with, endorsed by or speaking for any crawler vendor. Every statement is an observation with the date it was made and the raw evidence behind it: a robots.txt line, an HTTP status code, a header. There are no scores, no grades and no verdicts here, and nothing on this page is advice. A site opts out at any time and the opt out is honoured automatically and permanently.
