IndexHalo
MEASURED STUDY · 13 AUGUST 2026

AI crawler access:
reference and research

Encyclopaedias, journals and universities — the citation bedrock of the web. 25 domains measured; 8 block at least one of the fifteen tracked AI crawlers.

8of 25 reference and research block at least one AI crawler · median access 100%

Every domain in this category

Sorted by the share of tracked AI crawlers permitted, lowest first. Each report shows all fifteen verdicts and the robots.txt rule behind them.

DomainAccessAllowedBlockedNot listed
sciencedirect.com20%3120
nature.com33%5100
academia.edu40%690
sciencemag.org67%1050
dictionary.com80%1230
ssrn.com80%1230
jstor.org93%1410
pnas.org93%1410
acm.org100%1500
archive.org100%1500
arxiv.org100%1500
berkeley.edu100%1500
britannica.com100%1500
cam.ac.uk100%1500
cmu.edu100%1500
gutenberg.org100%1500
merriam-webster.com100%1500
mit.edu100%1500
plos.org100%1500
pubmed.ncbi.nlm.nih.gov100%1500
researchgate.net100%1500
springer.com100%1500
stanford.edu100%1500
wikipedia.org100%1500
wiktionary.org100%1500

Method and caveats are described on the full study page: one robots.txt fetch per domain, resolved against fifteen published AI user-agents, reported as the voluntary policy it is. Check any site yourself with the free crawler access checker.

Who started blocking which crawler this week, measured from live robots.txt. Unsubscribe in one click.