IndexHalo
MEASURED STUDY · 21 SEPTEMBER 2026

Who actually blocks
the AI crawlers?

We resolved 28 AI user-agents against the live robots.txt of 379 well-known websites. Every verdict below is measured, dated, and traceable to the rule that produced it.

See the AI Visibility Index — openness ranked by sector and by provider →

126of 379 sites block at least one AI crawler · 240 permit all 28 · median access 100%

Which crawlers get blocked most

Blocking is not evenly distributed. Training crawlers are declined far more often than the search and on-demand fetchers that produce citations — which is the distinction most publishers are actually trying to draw.

User-agentOperatorPurposeSites blockingShare
CCBotCommon Crawltraining dataset9024%
BytespiderByteDancemodel training8723%
ClaudeBotAnthropicmodel training8121%
Applebot-ExtendedAppleApple Intelligence control7420%
GPTBotOpenAImodel training7219%
DiffbotDiffbotstructured-data extraction7119%
AmazonbotAmazonmodel training6718%
Meta-ExternalAgentMetamodel training6617%
cohere-aiCoheremodel training6517%
Google-ExtendedGoogleGemini training control6216%
PerplexityBotPerplexitysearch indexing5815%
OmgilibotWebz.iotraining dataset5815%
TimpibotTimpitraining dataset5515%
YouBotYou.comsearch indexing5414%
ImagesiftBotImageSifttraining dataset4813%
Perplexity-UserPerplexityon-demand fetch4712%
PetalBotHuaweisearch indexing4512%
ChatGPT-UserOpenAIon-demand fetch4311%
Claude-UserAnthropicon-demand fetch4211%
Meta-ExternalFetcherMetaon-demand fetch4211%
Claude-SearchBotAnthropicsearch indexing4111%
MistralAI-UserMistral AIon-demand fetch4111%
DuckAssistBotDuckDuckGoon-demand fetch4011%
Google-CloudVertexBotGoogleVertex AI on-demand fetch3910%
AI2BotAllen Institute for AImodel training3910%
OAI-SearchBotOpenAIsearch indexing287%
GoogleOtherGoogleresearch and development fetch277%
BingbotMicrosoftCopilot search index21%

How this was measured

For each domain we requested /robots.txt once, over HTTPS, on the apex and then www. The file was parsed into groups and each of the 28 user-agents resolved against it: a direct group wins, otherwise the wildcard group applies, otherwise the crawler is recorded as not listed. Domains that returned nothing on either host are excluded rather than guessed at — 379 of the sites attempted produced a readable policy.

What this does not measure: whether any crawler honoured the policy, whether the content is useful once fetched, or anything about pages behind authentication. robots.txt is a voluntary protocol and a stated preference.

Browse by category

Who started blocking which crawler this week, measured from live robots.txt. Unsubscribe in one click.

Every domain measured

Sorted by the share of tracked AI crawlers permitted, lowest first.

DomainAccessAllowedBlockedNot listed
reddit.com0%0280
usatoday.com4%1270
globeandmail.com7%2260
quora.com7%2260
imdb.com11%3250
instacart.com11%3250
pinterest.com11%3250
telegraph.co.uk11%3250
amazon.com14%4240
cnn.com14%4240
reuters.com14%4240
lifehacker.com18%5230
mashable.com18%5230
sciencedirect.com18%5230
thetimes.co.uk18%5230
theverge.com18%5230
vox.com18%5230
arstechnica.com21%6220
nytimes.com21%6220
theregister.com21%6220
wsj.com21%6220
asahi.com25%7210
barrons.com25%7210
cntraveler.com25%7210
flickr.com25%7210
hackernoon.com25%7210
investopedia.com25%7210
marketwatch.com25%7210
newyorker.com25%7210
travelandleisure.com25%7210
verywellhealth.com25%7210
wired.com25%7210
bloomberg.com29%8200
buzzfeednews.com29%8200
dw.com29%8200
huffpost.com29%8200
nbcnews.com29%8200
smh.com.au29%8200
spiegel.de29%8200
theage.com.au29%8200
academia.edu32%9190
lemonde.fr32%9190
netflix.com32%9190
bbc.co.uk36%10180
theatlantic.com36%10180
pcmag.com39%11170
france24.com43%12160
nature.com43%12160
cnbc.com46%13150
fastcompany.com50%14140
inc.com50%14140
corriere.it54%15130
economist.com54%15130
ft.com54%15130
slate.com54%15130
theguardian.com54%15130
behance.net57%16120
elpais.com57%16120
forbes.com57%16120
tripadvisor.com57%16120
zeit.de57%16120
drugs.com61%17110
figma.com61%17110
newsweek.com61%17110
washingtonpost.com61%17110
apnews.com64%18100
canva.com64%18100
disneyplus.com64%18100
hulu.com64%18100
healthline.com68%1990
howtogeek.com68%1990
makeuseof.com68%1990
medicalnewstoday.com68%1990
xda-developers.com68%1990
ebay.com71%2080
vimeo.com71%2080
medium.com75%2170
pexels.com75%2170
abcnews.go.com79%2260
aljazeera.com79%2260
androidcentral.com79%2260
latimes.com79%2260
techradar.com79%2260
tomshardware.com79%2260
axios.com82%2350
businessinsider.com82%2350
sciencemag.org82%2350
venturebeat.com82%2350
metacritic.com86%2440
strava.com86%2440
bleepingcomputer.com89%2530
dictionary.com89%2530
fortune.com89%2530
geeksforgeeks.org89%2530
ssrn.com89%2530
tutorialspoint.com89%2530
webmd.com89%2530
alibaba.com93%2620
descript.com93%2620
digitaltrends.com93%2620
goodreads.com93%2620
irishtimes.com93%2620
ai.meta.com96%2710
airbnb.com96%2710
amplitude.com96%2710
calendly.com96%2710
cbsnews.com96%2710
chewy.com96%2710
fool.com96%2710
giphy.com96%2710
github.com96%2710
jasper.ai96%2710
jstor.org96%2710
lonelyplanet.com96%2710
loom.com96%2710
myfitnesspal.com96%2710
nerdwallet.com96%2710
notion.so96%2710
omio.com96%2710
pnas.org96%2710
searchenginejournal.com96%2710
skyscanner.net96%2710
straitstimes.com96%2710
together.ai96%2710
who.int96%0127
yoast.com96%2710
500px.com100%0028
acm.org100%2800
adobe.com100%2800
afar.com100%2800
ahrefs.com100%2800
airtable.com100%2800
aliexpress.com100%2800
anandtech.com100%0028
angular.io100%2800
anthropic.com100%2800
archive.org100%2800
arxiv.org100%2800
asana.com100%2800
atlassian.com100%2800
australia.gov.au100%0028
aws.amazon.com100%2800
azure.microsoft.com100%2800
backlinko.com100%2800
bankofamerica.com100%2800
bankrate.com100%2800
barclays.co.uk100%2800
basecamp.com100%2800
belastingdienst.nl100%2800
berkeley.edu100%2800
bestbuy.com100%2800
binance.com100%2800
booking.com100%2800
botify.com100%2800
box.com100%2800
brex.com100%2800
brightedge.com100%2800
britannica.com100%2800
britishairways.com100%2800
bund.de100%2800
calm.com100%2800
cam.ac.uk100%2800
canada.ca100%2800
cdc.gov100%0028
character.ai100%2800
chase.com100%2800
clearscope.io100%0028
clevelandclinic.org100%2800
clickup.com100%2800
close.com100%2800
cloud.google.com100%0028
cloudflare.com100%2800
cmu.edu100%2800
codecademy.com100%2800
codeium.com100%2800
cohere.com100%2800
coinbase.com100%2800
conductor.com100%2800
contentkingapp.com100%0028
copy.ai100%2800
costco.com100%2800
crates.io100%2800
creditkarma.com100%2800
css-tricks.com100%2800
cursor.com100%2800
datadog.com100%2800
deel.com100%2800
deepmind.google100%2800
deliveroo.co.uk100%2800
delta.com100%2800
dev.to100%2800
developer.mozilla.org100%2800
deviantart.com100%2800
digitalocean.com100%2800
dnsimple.com100%2800
docker.com100%2800
docusign.com100%2800
dribbble.com100%2800
drift.com100%0028
dropbox.com100%2800
elevenlabs.io100%2800
engadget.com100%2800
etsy.com100%2800
europa.eu100%2800
expedia.com100%2800
expensify.com100%2800
fastly.com100%2800
fda.gov100%2800
fireflies.ai100%2800
fitbit.com100%2800
fly.io100%2800
frase.io100%2800
freecodecamp.org100%2800
freshworks.com100%2800
gandi.net100%2800
gitlab.com100%2800
golang.org100%2800
goodrx.com100%2800
gov.uk100%2800
groq.com100%2800
gutenberg.org100%2800
hashnode.com100%2800
hbomax.com100%2800
headspace.com100%2800
hetzner.com100%2800
homedepot.com100%2800
hopkinsmedicine.org100%2800
hostelworld.com100%2800
hover.com100%2800
hsbc.com100%2800
hubspot.com100%2800
huggingface.co100%2800
ikea.com100%2800
imf.org100%2800
independent.co.uk100%2800
intercom.com100%2800
irs.gov100%2800
justeattakeaway.com100%2800
kayak.com100%2800
kinsta.com100%2800
klaviyo.com100%2800
kraken.com100%2800
kubernetes.io100%2800
leetcode.com100%2800
linear.app100%2800
linkedin.com100%0028
linode.com100%2800
lufthansa.com100%2800
lumar.io100%2800
macrumors.com100%2800
mailchimp.com100%2800
majestic.com100%2800
marketmuse.com100%2800
merriam-webster.com100%2800
miro.com100%2800
mistral.ai100%2800
mit.edu100%2800
mixpanel.com100%2800
monday.com100%2800
mongodb.com100%2800
monzo.com100%2800
morningstar.com100%2800
moz.com100%2800
mysql.com100%2800
n26.com100%2800
namecheap.com100%2800
neilpatel.com100%2800
netlify.app100%2800
netlify.com100%2800
newrelic.com100%2800
nhs.uk100%2800
nih.gov100%2800
nike.com100%2800
nodejs.org100%2800
oecd.org100%2800
oncrawl.com100%2800
openai.com100%2800
otter.ai100%2800
ouraring.com100%2800
ovhcloud.com100%2800
packagist.org100%2800
pagerduty.com100%2800
pantheon.io100%2800
paypal.com100%2800
perplexity.ai100%2800
php.net100%2800
pipedrive.com100%2800
plos.org100%2800
porkbun.com100%2800
postgresql.org100%2800
pubmed.ncbi.nlm.nih.gov100%2800
pypi.org100%2800
python.org100%2800
quickbooks.intuit.com100%2800
railway.com100%2800
ramp.com100%2800
reactjs.org100%2800
redis.io100%2800
render.com100%2800
replicate.com100%2800
researchgate.net100%2800
revolut.com100%2800
rippling.com100%2800
rottentomatoes.com100%2800
roughguides.com100%2800
rubygems.org100%2800
runwayml.com100%2800
ryanair.com100%0028
sage.com100%2800
salesforce.com100%2800
santander.com100%2800
scmp.com100%2800
screamingfrog.co.uk100%2800
sec.gov100%2800
segment.com100%2800
semrush.com100%2800
sendgrid.com100%2800
seranking.com100%2800
seroundtable.com100%2800
serpstat.com100%2800
shopify.com100%2800
similarweb.com100%2800
sistrix.com100%0028
sitepoint.com100%2800
skatteverket.se100%2800
sky.com100%2800
slack.com100%2800
slashdot.org100%2800
smashingmagazine.com100%2800
soundcloud.com100%2800
sourcegraph.com100%2800
spotify.com100%2800
springer.com100%2800
spyfu.com100%2800
sqlite.org100%2800
squarespace.com100%2800
stability.ai100%2800
stanford.edu100%2800
starlingbank.com100%2800
stripe.com100%2800
surferseo.com100%2800
surveymonkey.com100%2800
svelte.dev100%2800
synthesia.io100%2800
tabnine.com100%2800
target.com100%2800
teladoc.com100%0028
thenextweb.com100%2800
time.com100%2800
trainline.com100%2800
trello.com100%2800
twilio.com100%2800
twitch.tv100%2800
typeform.com100%2800
ulta.com100%2800
un.org100%2800
unsplash.com100%2800
usa.gov100%2800
vercel.com100%2800
vrbo.com100%2800
vuejs.org100%2800
w3schools.com100%2800
walmart.com100%2800
wayfair.com100%2800
wellsfargo.com100%2800
wikipedia.org100%2800
wiktionary.org100%2800
wise.com100%2800
wix.com100%2800
worldbank.org100%2800
wpengine.com100%2800
writesonic.com100%2800
x.ai100%2800
xero.com100%2800
youtube.com100%2800
zalando.com100%0028
zendesk.com100%2800
zoho.com100%2800
zoom.us100%2800

Where does your site sit?

Run the same 28-crawler resolution against your own domain, free and without an account, with the AI crawler access checker. If the answer surprises you, the policy builder composes a corrected robots.txt block, and the ChatGPT citation guide covers the training-versus-search distinction that catches most publishers out.

ACCESS IS ONLY THE FIRST LAYER

Reachable is not the same as citable.

The free report measures whether the content itself can be extracted, attributed and quoted once a crawler arrives.

Run a free report →