AI crawler access:
news publishers
The sites with the most at stake: their archives train models and their reporting feeds live answers. 44 domains measured; 40 block at least one of the 28 tracked AI crawlers.
40of 44 news publishers block at least one AI crawler · median access 43%
Every domain in this category
Sorted by the share of tracked AI crawlers permitted, lowest first. Each report shows all 28 verdicts and the robots.txt rule behind them.
| Domain | Access | Allowed | Blocked | Not listed |
|---|---|---|---|---|
| usatoday.com | 4% | 1 | 27 | 0 |
| globeandmail.com | 7% | 2 | 26 | 0 |
| telegraph.co.uk | 11% | 3 | 25 | 0 |
| cnn.com | 14% | 4 | 24 | 0 |
| reuters.com | 14% | 4 | 24 | 0 |
| thetimes.co.uk | 18% | 5 | 23 | 0 |
| vox.com | 18% | 5 | 23 | 0 |
| nytimes.com | 21% | 6 | 22 | 0 |
| wsj.com | 21% | 6 | 22 | 0 |
| asahi.com | 25% | 7 | 21 | 0 |
| newyorker.com | 25% | 7 | 21 | 0 |
| bloomberg.com | 29% | 8 | 20 | 0 |
| buzzfeednews.com | 29% | 8 | 20 | 0 |
| dw.com | 29% | 8 | 20 | 0 |
| huffpost.com | 29% | 8 | 20 | 0 |
| nbcnews.com | 29% | 8 | 20 | 0 |
| smh.com.au | 29% | 8 | 20 | 0 |
| spiegel.de | 29% | 8 | 20 | 0 |
| theage.com.au | 29% | 8 | 20 | 0 |
| lemonde.fr | 32% | 9 | 19 | 0 |
| bbc.co.uk | 36% | 10 | 18 | 0 |
| theatlantic.com | 36% | 10 | 18 | 0 |
| france24.com | 43% | 12 | 16 | 0 |
| corriere.it | 54% | 15 | 13 | 0 |
| economist.com | 54% | 15 | 13 | 0 |
| ft.com | 54% | 15 | 13 | 0 |
| slate.com | 54% | 15 | 13 | 0 |
| theguardian.com | 54% | 15 | 13 | 0 |
| elpais.com | 57% | 16 | 12 | 0 |
| zeit.de | 57% | 16 | 12 | 0 |
| newsweek.com | 61% | 17 | 11 | 0 |
| washingtonpost.com | 61% | 17 | 11 | 0 |
| apnews.com | 64% | 18 | 10 | 0 |
| abcnews.go.com | 79% | 22 | 6 | 0 |
| aljazeera.com | 79% | 22 | 6 | 0 |
| latimes.com | 79% | 22 | 6 | 0 |
| axios.com | 82% | 23 | 5 | 0 |
| irishtimes.com | 93% | 26 | 2 | 0 |
| cbsnews.com | 96% | 27 | 1 | 0 |
| straitstimes.com | 96% | 27 | 1 | 0 |
| independent.co.uk | 100% | 28 | 0 | 0 |
| scmp.com | 100% | 28 | 0 | 0 |
| sky.com | 100% | 28 | 0 | 0 |
| time.com | 100% | 28 | 0 | 0 |
Method and caveats are described on the full study page: one robots.txt fetch per domain, resolved against 28 published AI user-agents, reported as the voluntary policy it is. Check any site yourself with the free crawler access checker.