AI crawler access:
news publishers
The sites with the most at stake: their archives train models and their reporting feeds live answers. 44 domains measured; 39 block at least one of the fifteen tracked AI crawlers.
39of 44 news publishers block at least one AI crawler · median access 33%
Every domain in this category
Sorted by the share of tracked AI crawlers permitted, lowest first. Each report shows all fifteen verdicts and the robots.txt rule behind them.
| Domain | Access | Allowed | Blocked | Not listed |
|---|---|---|---|---|
| asahi.com | 7% | 1 | 14 | 0 |
| buzzfeednews.com | 7% | 1 | 14 | 0 |
| huffpost.com | 7% | 1 | 14 | 0 |
| nbcnews.com | 7% | 1 | 14 | 0 |
| usatoday.com | 7% | 1 | 14 | 0 |
| bloomberg.com | 13% | 2 | 13 | 0 |
| cnn.com | 13% | 2 | 13 | 0 |
| globeandmail.com | 13% | 2 | 13 | 0 |
| nytimes.com | 13% | 2 | 13 | 0 |
| telegraph.co.uk | 13% | 2 | 13 | 0 |
| bbc.co.uk | 20% | 3 | 12 | 0 |
| economist.com | 20% | 3 | 12 | 0 |
| reuters.com | 20% | 3 | 12 | 0 |
| dw.com | 27% | 4 | 11 | 0 |
| newyorker.com | 27% | 4 | 11 | 0 |
| smh.com.au | 27% | 4 | 11 | 0 |
| spiegel.de | 27% | 4 | 11 | 0 |
| theage.com.au | 27% | 4 | 11 | 0 |
| theatlantic.com | 27% | 4 | 11 | 0 |
| thetimes.co.uk | 27% | 4 | 11 | 0 |
| vox.com | 27% | 4 | 11 | 0 |
| wsj.com | 27% | 4 | 11 | 0 |
| france24.com | 33% | 5 | 10 | 0 |
| theguardian.com | 40% | 6 | 9 | 0 |
| zeit.de | 40% | 6 | 9 | 0 |
| apnews.com | 47% | 7 | 8 | 0 |
| corriere.it | 47% | 7 | 8 | 0 |
| elpais.com | 47% | 7 | 8 | 0 |
| ft.com | 47% | 7 | 8 | 0 |
| lemonde.fr | 47% | 7 | 8 | 0 |
| newsweek.com | 47% | 7 | 8 | 0 |
| slate.com | 47% | 7 | 8 | 0 |
| washingtonpost.com | 53% | 8 | 7 | 0 |
| abcnews.go.com | 67% | 10 | 5 | 0 |
| aljazeera.com | 67% | 10 | 5 | 0 |
| axios.com | 80% | 12 | 3 | 0 |
| latimes.com | 80% | 12 | 3 | 0 |
| irishtimes.com | 87% | 13 | 2 | 0 |
| cbsnews.com | 93% | 14 | 1 | 0 |
| independent.co.uk | 100% | 15 | 0 | 0 |
| scmp.com | 100% | 15 | 0 | 0 |
| sky.com | 100% | 15 | 0 | 0 |
| straitstimes.com | 100% | 15 | 0 | 0 |
| time.com | 100% | 15 | 0 | 0 |
Method and caveats are described on the full study page: one robots.txt fetch per domain, resolved against fifteen published AI user-agents, reported as the voluntary policy it is. Check any site yourself with the free crawler access checker.