IndexHalo
MEASURED STUDY · 13 AUGUST 2026

AI crawler access:
news publishers

The sites with the most at stake: their archives train models and their reporting feeds live answers. 44 domains measured; 39 block at least one of the fifteen tracked AI crawlers.

39of 44 news publishers block at least one AI crawler · median access 33%

Every domain in this category

Sorted by the share of tracked AI crawlers permitted, lowest first. Each report shows all fifteen verdicts and the robots.txt rule behind them.

DomainAccessAllowedBlockedNot listed
asahi.com7%1140
buzzfeednews.com7%1140
huffpost.com7%1140
nbcnews.com7%1140
usatoday.com7%1140
bloomberg.com13%2130
cnn.com13%2130
globeandmail.com13%2130
nytimes.com13%2130
telegraph.co.uk13%2130
bbc.co.uk20%3120
economist.com20%3120
reuters.com20%3120
dw.com27%4110
newyorker.com27%4110
smh.com.au27%4110
spiegel.de27%4110
theage.com.au27%4110
theatlantic.com27%4110
thetimes.co.uk27%4110
vox.com27%4110
wsj.com27%4110
france24.com33%5100
theguardian.com40%690
zeit.de40%690
apnews.com47%780
corriere.it47%780
elpais.com47%780
ft.com47%780
lemonde.fr47%780
newsweek.com47%780
slate.com47%780
washingtonpost.com53%870
abcnews.go.com67%1050
aljazeera.com67%1050
axios.com80%1230
latimes.com80%1230
irishtimes.com87%1320
cbsnews.com93%1410
independent.co.uk100%1500
scmp.com100%1500
sky.com100%1500
straitstimes.com100%1500
time.com100%1500

Method and caveats are described on the full study page: one robots.txt fetch per domain, resolved against fifteen published AI user-agents, reported as the voluntary policy it is. Check any site yourself with the free crawler access checker.

Who started blocking which crawler this week, measured from live robots.txt. Unsubscribe in one click.