AI crawler access:
reference and research
Encyclopaedias, journals and universities — the citation bedrock of the web. 25 domains measured; 8 block at least one of the fifteen tracked AI crawlers.
8of 25 reference and research block at least one AI crawler · median access 100%
Every domain in this category
Sorted by the share of tracked AI crawlers permitted, lowest first. Each report shows all fifteen verdicts and the robots.txt rule behind them.
| Domain | Access | Allowed | Blocked | Not listed |
|---|---|---|---|---|
| sciencedirect.com | 20% | 3 | 12 | 0 |
| nature.com | 33% | 5 | 10 | 0 |
| academia.edu | 40% | 6 | 9 | 0 |
| sciencemag.org | 67% | 10 | 5 | 0 |
| dictionary.com | 80% | 12 | 3 | 0 |
| ssrn.com | 80% | 12 | 3 | 0 |
| jstor.org | 93% | 14 | 1 | 0 |
| pnas.org | 93% | 14 | 1 | 0 |
| acm.org | 100% | 15 | 0 | 0 |
| archive.org | 100% | 15 | 0 | 0 |
| arxiv.org | 100% | 15 | 0 | 0 |
| berkeley.edu | 100% | 15 | 0 | 0 |
| britannica.com | 100% | 15 | 0 | 0 |
| cam.ac.uk | 100% | 15 | 0 | 0 |
| cmu.edu | 100% | 15 | 0 | 0 |
| gutenberg.org | 100% | 15 | 0 | 0 |
| merriam-webster.com | 100% | 15 | 0 | 0 |
| mit.edu | 100% | 15 | 0 | 0 |
| plos.org | 100% | 15 | 0 | 0 |
| pubmed.ncbi.nlm.nih.gov | 100% | 15 | 0 | 0 |
| researchgate.net | 100% | 15 | 0 | 0 |
| springer.com | 100% | 15 | 0 | 0 |
| stanford.edu | 100% | 15 | 0 | 0 |
| wikipedia.org | 100% | 15 | 0 | 0 |
| wiktionary.org | 100% | 15 | 0 | 0 |
Method and caveats are described on the full study page: one robots.txt fetch per domain, resolved against fifteen published AI user-agents, reported as the voluntary policy it is. Check any site yourself with the free crawler access checker.