IndexHalo
MEASURED STUDY · 13 AUGUST 2026

AI crawler access:
ai companies

The AI industry's own websites, resolved against the industry's own crawlers. 28 domains measured; 2 block at least one of the fifteen tracked AI crawlers.

2of 28 ai companies block at least one AI crawler · median access 100%

Every domain in this category

Sorted by the share of tracked AI crawlers permitted, lowest first. Each report shows all fifteen verdicts and the robots.txt rule behind them.

DomainAccessAllowedBlockedNot listed
descript.com87%1320
together.ai93%1410
ai.meta.com100%1500
anthropic.com100%1500
character.ai100%1500
codeium.com100%1500
cohere.com100%1500
copy.ai100%1500
cursor.com100%1500
deepmind.google100%1500
elevenlabs.io100%1500
fireflies.ai100%1500
gamma.app100%1500
groq.com100%1500
huggingface.co100%1500
jasper.ai100%1500
mistral.ai100%1500
openai.com100%1500
otter.ai100%1500
perplexity.ai100%1500
replicate.com100%1500
runwayml.com100%1500
sourcegraph.com100%1500
stability.ai100%1500
synthesia.io100%1500
tabnine.com100%1500
writesonic.com100%1500
x.ai100%1500

Method and caveats are described on the full study page: one robots.txt fetch per domain, resolved against fifteen published AI user-agents, reported as the voluntary policy it is. Check any site yourself with the free crawler access checker.

Who started blocking which crawler this week, measured from live robots.txt. Unsubscribe in one click.