IndexHalo
MEASURED STUDY · 21 SEPTEMBER 2026

AI crawler access:
news publishers

The sites with the most at stake: their archives train models and their reporting feeds live answers. 44 domains measured; 40 block at least one of the 28 tracked AI crawlers.

40of 44 news publishers block at least one AI crawler · median access 43%

Every domain in this category

Sorted by the share of tracked AI crawlers permitted, lowest first. Each report shows all 28 verdicts and the robots.txt rule behind them.

DomainAccessAllowedBlockedNot listed
usatoday.com4%1270
globeandmail.com7%2260
telegraph.co.uk11%3250
cnn.com14%4240
reuters.com14%4240
thetimes.co.uk18%5230
vox.com18%5230
nytimes.com21%6220
wsj.com21%6220
asahi.com25%7210
newyorker.com25%7210
bloomberg.com29%8200
buzzfeednews.com29%8200
dw.com29%8200
huffpost.com29%8200
nbcnews.com29%8200
smh.com.au29%8200
spiegel.de29%8200
theage.com.au29%8200
lemonde.fr32%9190
bbc.co.uk36%10180
theatlantic.com36%10180
france24.com43%12160
corriere.it54%15130
economist.com54%15130
ft.com54%15130
slate.com54%15130
theguardian.com54%15130
elpais.com57%16120
zeit.de57%16120
newsweek.com61%17110
washingtonpost.com61%17110
apnews.com64%18100
abcnews.go.com79%2260
aljazeera.com79%2260
latimes.com79%2260
axios.com82%2350
irishtimes.com93%2620
cbsnews.com96%2710
straitstimes.com96%2710
independent.co.uk100%2800
scmp.com100%2800
sky.com100%2800
time.com100%2800

Method and caveats are described on the full study page: one robots.txt fetch per domain, resolved against 28 published AI user-agents, reported as the voluntary policy it is. Check any site yourself with the free crawler access checker.

Who started blocking which crawler this week, measured from live robots.txt. Unsubscribe in one click.