Who actually blocks
the AI crawlers?
We resolved 28 AI user-agents against the live robots.txt of 379 well-known websites. Every verdict below is measured, dated, and traceable to the rule that produced it.
See the AI Visibility Index — openness ranked by sector and by provider →
Which crawlers get blocked most
Blocking is not evenly distributed. Training crawlers are declined far more often than the search and on-demand fetchers that produce citations — which is the distinction most publishers are actually trying to draw.
| User-agent | Operator | Purpose | Sites blocking | Share |
|---|---|---|---|---|
CCBot | Common Crawl | training dataset | 90 | 24% |
Bytespider | ByteDance | model training | 87 | 23% |
ClaudeBot | Anthropic | model training | 81 | 21% |
Applebot-Extended | Apple | Apple Intelligence control | 74 | 20% |
GPTBot | OpenAI | model training | 72 | 19% |
Diffbot | Diffbot | structured-data extraction | 71 | 19% |
Amazonbot | Amazon | model training | 67 | 18% |
Meta-ExternalAgent | Meta | model training | 66 | 17% |
cohere-ai | Cohere | model training | 65 | 17% |
Google-Extended | Gemini training control | 62 | 16% | |
PerplexityBot | Perplexity | search indexing | 58 | 15% |
Omgilibot | Webz.io | training dataset | 58 | 15% |
Timpibot | Timpi | training dataset | 55 | 15% |
YouBot | You.com | search indexing | 54 | 14% |
ImagesiftBot | ImageSift | training dataset | 48 | 13% |
Perplexity-User | Perplexity | on-demand fetch | 47 | 12% |
PetalBot | Huawei | search indexing | 45 | 12% |
ChatGPT-User | OpenAI | on-demand fetch | 43 | 11% |
Claude-User | Anthropic | on-demand fetch | 42 | 11% |
Meta-ExternalFetcher | Meta | on-demand fetch | 42 | 11% |
Claude-SearchBot | Anthropic | search indexing | 41 | 11% |
MistralAI-User | Mistral AI | on-demand fetch | 41 | 11% |
DuckAssistBot | DuckDuckGo | on-demand fetch | 40 | 11% |
Google-CloudVertexBot | Vertex AI on-demand fetch | 39 | 10% | |
AI2Bot | Allen Institute for AI | model training | 39 | 10% |
OAI-SearchBot | OpenAI | search indexing | 28 | 7% |
GoogleOther | research and development fetch | 27 | 7% | |
Bingbot | Microsoft | Copilot search index | 2 | 1% |
How this was measured
For each domain we requested /robots.txt once, over HTTPS, on the apex and then www. The file was parsed into groups and each of the 28 user-agents resolved against it: a direct group wins, otherwise the wildcard group applies, otherwise the crawler is recorded as not listed. Domains that returned nothing on either host are excluded rather than guessed at — 379 of the sites attempted produced a readable policy.
What this does not measure: whether any crawler honoured the policy, whether the content is useful once fetched, or anything about pages behind authentication. robots.txt is a voluntary protocol and a stated preference.
Browse by category
| Category | Domains | Block ≥1 crawler | Median access |
|---|---|---|---|
| News publishers | 44 | 40 | 43% |
| Technology media | 22 | 17 | 79% |
| Developer references | 34 | 3 | 100% |
| SaaS platforms | 53 | 6 | 100% |
| AI companies | 27 | 4 | 100% |
| SEO and marketing tools | 25 | 2 | 100% |
| Hosting and infrastructure | 22 | 0 | 100% |
| Reference and research | 25 | 8 | 100% |
| E-commerce and retail | 19 | 5 | 100% |
| Finance and business media | 29 | 11 | 100% |
| Government and institutions | 18 | 1 | 100% |
| Health publishers | 16 | 7 | 100% |
| Travel | 19 | 7 | 100% |
| Social and media platforms | 26 | 15 | 93% |
Every domain measured
Sorted by the share of tracked AI crawlers permitted, lowest first.
Where does your site sit?
Run the same 28-crawler resolution against your own domain, free and without an account, with the AI crawler access checker. If the answer surprises you, the policy builder composes a corrected robots.txt block, and the ChatGPT citation guide covers the training-versus-search distinction that catches most publishers out.