IndexHalo
MEASURED STUDY · 13 AUGUST 2026

AI crawler access:
developer references

Documentation and Q&A sites — the sources answer engines lean on hardest for technical queries. 34 domains measured; 2 block at least one of the fifteen tracked AI crawlers.

2of 34 developer references block at least one AI crawler · median access 100%

Every domain in this category

Sorted by the share of tracked AI crawlers permitted, lowest first. Each report shows all fifteen verdicts and the robots.txt rule behind them.

DomainAccessAllowedBlockedNot listed
tutorialspoint.com80%1230
geeksforgeeks.org87%1320
angular.io100%1500
codecademy.com100%1500
crates.io100%1500
css-tricks.com100%1500
dev.to100%1500
developer.mozilla.org100%1500
digitalocean.com100%1500
docker.com100%1500
freecodecamp.org100%1500
github.com100%1500
gitlab.com100%1500
golang.org100%1500
hashnode.com100%1500
kubernetes.io100%1500
leetcode.com100%1500
mongodb.com100%1500
mysql.com100%1500
nodejs.org100%1500
packagist.org100%1500
php.net100%1500
postgresql.org100%1500
pypi.org100%1500
python.org100%1500
reactjs.org100%1500
redis.io100%1500
rubygems.org100%1500
sitepoint.com100%1500
smashingmagazine.com100%1500
sqlite.org100%1500
svelte.dev100%1500
vuejs.org100%1500
w3schools.com100%1500

Method and caveats are described on the full study page: one robots.txt fetch per domain, resolved against fifteen published AI user-agents, reported as the voluntary policy it is. Check any site yourself with the free crawler access checker.

Who started blocking which crawler this week, measured from live robots.txt. Unsubscribe in one click.