IndexHalo
MEASURED STUDY · 21 SEPTEMBER 2026

AI crawler access:
reference and research

Encyclopaedias, journals and universities — the citation bedrock of the web. 25 domains measured; 8 block at least one of the 28 tracked AI crawlers.

8of 25 reference and research block at least one AI crawler · median access 100%

Every domain in this category

Sorted by the share of tracked AI crawlers permitted, lowest first. Each report shows all 28 verdicts and the robots.txt rule behind them.

DomainAccessAllowedBlockedNot listed
sciencedirect.com18%5230
academia.edu32%9190
nature.com43%12160
sciencemag.org82%2350
dictionary.com89%2530
ssrn.com89%2530
jstor.org96%2710
pnas.org96%2710
acm.org100%2800
archive.org100%2800
arxiv.org100%2800
berkeley.edu100%2800
britannica.com100%2800
cam.ac.uk100%2800
cmu.edu100%2800
gutenberg.org100%2800
merriam-webster.com100%2800
mit.edu100%2800
plos.org100%2800
pubmed.ncbi.nlm.nih.gov100%2800
researchgate.net100%2800
springer.com100%2800
stanford.edu100%2800
wikipedia.org100%2800
wiktionary.org100%2800

Method and caveats are described on the full study page: one robots.txt fetch per domain, resolved against 28 published AI user-agents, reported as the voluntary policy it is. Check any site yourself with the free crawler access checker.

Who started blocking which crawler this week, measured from live robots.txt. Unsubscribe in one click.