AI crawler access:
technology media
Outlets covering the AI industry, and deciding whether its crawlers may read them. 22 domains measured; 17 block at least one of the 28 tracked AI crawlers.
17of 22 technology media block at least one AI crawler · median access 79%
Every domain in this category
Sorted by the share of tracked AI crawlers permitted, lowest first. Each report shows all 28 verdicts and the robots.txt rule behind them.
| Domain | Access | Allowed | Blocked | Not listed |
|---|---|---|---|---|
| lifehacker.com | 18% | 5 | 23 | 0 |
| mashable.com | 18% | 5 | 23 | 0 |
| theverge.com | 18% | 5 | 23 | 0 |
| arstechnica.com | 21% | 6 | 22 | 0 |
| theregister.com | 21% | 6 | 22 | 0 |
| hackernoon.com | 25% | 7 | 21 | 0 |
| wired.com | 25% | 7 | 21 | 0 |
| pcmag.com | 39% | 11 | 17 | 0 |
| howtogeek.com | 68% | 19 | 9 | 0 |
| makeuseof.com | 68% | 19 | 9 | 0 |
| xda-developers.com | 68% | 19 | 9 | 0 |
| androidcentral.com | 79% | 22 | 6 | 0 |
| techradar.com | 79% | 22 | 6 | 0 |
| tomshardware.com | 79% | 22 | 6 | 0 |
| venturebeat.com | 82% | 23 | 5 | 0 |
| bleepingcomputer.com | 89% | 25 | 3 | 0 |
| digitaltrends.com | 93% | 26 | 2 | 0 |
| anandtech.com | 100% | 0 | 0 | 28 |
| engadget.com | 100% | 28 | 0 | 0 |
| macrumors.com | 100% | 28 | 0 | 0 |
| slashdot.org | 100% | 28 | 0 | 0 |
| thenextweb.com | 100% | 28 | 0 | 0 |
Method and caveats are described on the full study page: one robots.txt fetch per domain, resolved against 28 published AI user-agents, reported as the voluntary policy it is. Check any site yourself with the free crawler access checker.