IndexHalo
MEASURED STUDY · 12 AUGUST 2026

Who actually blocks
the AI crawlers?

We resolved fifteen AI user-agents against the live robots.txt of 87 well-known websites. Every verdict below is measured, dated, and traceable to the rule that produced it.

31of 87 sites block at least one AI crawler · 54 permit all fifteen · median access 100%

Which crawlers get blocked most

Blocking is not evenly distributed. Training crawlers are declined far more often than the search and on-demand fetchers that produce citations — which is the distinction most publishers are actually trying to draw.

User-agentOperatorPurposeSites blockingShare
ClaudeBotAnthropicmodel training2832%
CCBotCommon Crawltraining dataset2731%
BytespiderByteDancemodel training2529%
Applebot-ExtendedAppleApple Intelligence control2428%
PerplexityBotPerplexitysearch indexing2326%
Meta-ExternalAgentMetamodel training2225%
GPTBotOpenAImodel training2124%
Google-ExtendedGoogleGemini training control1922%
AmazonbotAmazonmodel training1922%
Claude-UserAnthropicon-demand fetch1821%
Claude-SearchBotAnthropicsearch indexing1720%
Perplexity-UserPerplexityon-demand fetch1618%
ChatGPT-UserOpenAIon-demand fetch1517%
OAI-SearchBotOpenAIsearch indexing1011%
BingbotMicrosoftCopilot search index22%

How this was measured

For each domain we requested /robots.txt once, over HTTPS, on the apex and then www. The file was parsed into groups and each of the fifteen user-agents resolved against it: a direct group wins, otherwise the wildcard group applies, otherwise the crawler is recorded as not listed. Domains that returned nothing on either host are excluded rather than guessed at — 87 of the sites attempted produced a readable policy.

What this does not measure: whether any crawler honoured the policy, whether the content is useful once fetched, or anything about pages behind authentication. robots.txt is a voluntary protocol and a stated preference.

Every domain measured

Sorted by the share of tracked AI crawlers permitted, lowest first.

DomainAccessAllowedBlockedNot listed
reddit.com0%0150
bloomberg.com13%2130
cnn.com13%2130
nytimes.com13%2130
quora.com13%2130
telegraph.co.uk13%2130
amazon.com20%3120
bbc.co.uk20%3120
economist.com20%3120
reuters.com20%3120
sciencedirect.com20%3120
arstechnica.com27%4110
investopedia.com27%4110
theverge.com27%4110
wired.com27%4110
wsj.com27%4110
nature.com33%5100
figma.com40%690
theguardian.com40%690
apnews.com47%780
forbes.com47%780
ft.com47%780
tripadvisor.com47%780
healthline.com53%870
canva.com60%960
medium.com60%960
aljazeera.com67%1050
webmd.com73%1140
jstor.org93%1410
notion.so93%1410
who.int93%0114
adobe.com100%1500
ahrefs.com100%1500
airbnb.com100%1500
anthropic.com100%1500
apple.com100%1500
arxiv.org100%1500
asana.com100%1500
atlassian.com100%1500
booking.com100%1500
cloudflare.com100%1500
cohere.com100%1500
developer.mozilla.org100%1500
digitalocean.com100%1500
drift.com100%0015
europa.eu100%1500
expedia.com100%1500
freshworks.com100%1500
github.com100%1500
gitlab.com100%1500
google.com100%1500
gov.uk100%1500
hubspot.com100%1500
huggingface.co100%1500
independent.co.uk100%1500
intercom.com100%1500
klaviyo.com100%1500
linkedin.com100%0015
mailchimp.com100%1500
microsoft.com100%1500
mistral.ai100%1500
monday.com100%1500
moz.com100%1500
netlify.com100%1500
nhs.uk100%1500
oecd.org100%1500
openai.com100%1500
perplexity.ai100%1500
pubmed.ncbi.nlm.nih.gov100%1500
railway.com100%1500
salesforce.com100%1500
screamingfrog.co.uk100%1500
semrush.com100%1500
shopify.com100%1500
similarweb.com100%1500
slack.com100%1500
squarespace.com100%1500
stripe.com100%1500
substack.com100%1500
vercel.com100%1500
w3schools.com100%1500
webflow.com100%1500
wikipedia.org100%1500
wix.com100%1500
worldbank.org100%1500
zendesk.com100%1500
zoom.us100%1500

Where does your site sit?

Run the same fifteen-crawler resolution against your own domain, free and without an account, with the AI crawler access checker. If the answer surprises you, the policy builder composes a corrected robots.txt block, and the ChatGPT citation guide covers the training-versus-search distinction that catches most publishers out.

ACCESS IS ONLY THE FIRST LAYER

Reachable is not the same as citable.

The free report measures whether the content itself can be extracted, attributed and quoted once a crawler arrives.

Run a free report →