IndexHalo
MEASURED STUDY · 21 SEPTEMBER 2026

AI crawler access:
ai companies

The AI industry's own websites, resolved against the industry's own crawlers. 27 domains measured; 4 block at least one of the 28 tracked AI crawlers.

4of 27 ai companies block at least one AI crawler · median access 100%

Every domain in this category

Sorted by the share of tracked AI crawlers permitted, lowest first. Each report shows all 28 verdicts and the robots.txt rule behind them.

DomainAccessAllowedBlockedNot listed
descript.com93%2620
ai.meta.com96%2710
jasper.ai96%2710
together.ai96%2710
anthropic.com100%2800
character.ai100%2800
codeium.com100%2800
cohere.com100%2800
copy.ai100%2800
cursor.com100%2800
deepmind.google100%2800
elevenlabs.io100%2800
fireflies.ai100%2800
groq.com100%2800
huggingface.co100%2800
mistral.ai100%2800
openai.com100%2800
otter.ai100%2800
perplexity.ai100%2800
replicate.com100%2800
runwayml.com100%2800
sourcegraph.com100%2800
stability.ai100%2800
synthesia.io100%2800
tabnine.com100%2800
writesonic.com100%2800
x.ai100%2800

Method and caveats are described on the full study page: one robots.txt fetch per domain, resolved against 28 published AI user-agents, reported as the voluntary policy it is. Check any site yourself with the free crawler access checker.

Who started blocking which crawler this week, measured from live robots.txt. Unsubscribe in one click.