Is the web letting
the answer engines in?
Across 380 public sites, the median site allows 100% of known AI crawlers, and 68% are fully open. This index reads live robots policies — the same measurement we publish per domain — and aggregates it by sector and by AI provider. Every figure is measured, carries its sample size, and is dated from the crawl itself.
Openness by sector
Median share of AI crawlers allowed, and the share of sites in the sector that are fully open. Ranked most-open first.
| Sector | Median access | Fully open | |
|---|---|---|---|
| Hosting and infrastructure | 100% | 100% · n=22 | |
| Developer references | 100% | 94% · n=34 | |
| Government and institutions | 100% | 94% · n=18 | |
| SEO and marketing tools | 100% | 92% · n=26 | |
| SaaS platforms | 100% | 91% · n=54 | |
| AI companies | 100% | 86% · n=28 | |
| E-commerce and retail | 100% | 71% · n=17 | |
| Reference and research | 100% | 67% · n=24 | |
| Finance and business media | 100% | 62% · n=29 | |
| Travel | 100% | 61% · n=18 | |
| Health publishers | 100% | 56% · n=16 | |
| Social and media platforms | 96% | 48% · n=27 | |
| Technology media | 68% | 22% · n=23 | |
| News publishers | 40% | 11% · n=44 |
Full figures as data
[
{
"sector": "Hosting and infrastructure",
"domains": 22,
"medianAccessPct": 100,
"meanAccessPct": 100,
"fullyOpenPct": 100
},
{
"sector": "Developer references",
"domains": 34,
"medianAccessPct": 100,
"meanAccessPct": 99,
"fullyOpenPct": 94
},
{
"sector": "Government and institutions",
"domains": 18,
"medianAccessPct": 100,
"meanAccessPct": 100,
"fullyOpenPct": 94
},
{
"sector": "SEO and marketing tools",
"domains": 26,
"medianAccessPct": 100,
"meanAccessPct": 100,
"fullyOpenPct": 92
},
{
"sector": "SaaS platforms",
"domains": 54,
"medianAccessPct": 100,
"meanAccessPct": 98,
"fullyOpenPct": 91
},
{
"sector": "AI companies",
"domains": 28,
"medianAccessPct": 100,
"meanAccessPct": 99,
"fullyOpenPct": 86
},
{
"sector": "E-commerce and retail",
"domains": 17,
"medianAccessPct": 100,
"meanAccessPct": 87,
"fullyOpenPct": 71
},
{
"sector": "Reference and research",
"domains": 24,
"medianAccessPct": 100,
"meanAccessPct": 89,
"fullyOpenPct": 67
},
{
"sector": "Finance and business media",
"domains": 29,
"medianAccessPct": 100,
"meanAccessPct": 87,
"fullyOpenPct": 62
},
{
"sector": "Travel",
"domains": 18,
"medianAccessPct": 100,
"meanAccessPct": 89,
"fullyOpenPct": 61
},
{
"sector": "Health publishers",
"domains": 16,
"medianAccessPct": 100,
"meanAccessPct": 87,
"fullyOpenPct": 56
},
{
"sector": "Social and media platforms",
"domains": 27,
"medianAccessPct": 96,
"meanAccessPct": 74,
"fullyOpenPct": 48
},
{
"sector": "Technology media",
"domains": 23,
"medianAccessPct": 68,
"meanAccessPct": 62,
"fullyOpenPct": 22
},
{
"sector": "News publishers",
"domains": 44,
"medianAccessPct": 40,
"meanAccessPct": 49,
"fullyOpenPct": 11
}
]Which AI providers get blocked most
Share of the 380 measured sites that block at least one crawler from each provider, with the 95% confidence interval.
| Provider | Sites blocking | 95% CI |
|---|---|---|
| Common Crawl | 23% | 19–28% · n=380 |
| ByteDance | 22% | 18–27% · n=380 |
| Anthropic | 21% | 17–25% · n=380 |
| OpenAI | 20% | 16–24% · n=380 |
| Apple | 19% | 15–23% · n=380 |
| 18% | 14–22% · n=380 | |
| Meta | 18% | 15–22% · n=380 |
| Diffbot | 18% | 14–22% · n=380 |
| Amazon | 17% | 14–21% · n=380 |
| Cohere | 17% | 13–21% · n=380 |
| Perplexity | 16% | 13–20% · n=380 |
| Webz.io | 15% | 12–19% · n=380 |
| Timpi | 14% | 11–18% · n=380 |
| You.com | 13% | 10–16% · n=380 |
| ImageSift | 12% | 9–16% · n=380 |
| Huawei | 12% | 9–16% · n=380 |
| Allen Institute for AI | 11% | 8–14% · n=380 |
| Mistral AI | 10% | 8–14% · n=380 |
| DuckDuckGo | 10% | 8–14% · n=380 |
| Microsoft | 1% | 0–2% · n=380 |
How this is measured
Each domain's live robots policy is parsed and every known AI crawler is classified allowed or blocked (the same measurement published per-domain at /ai-crawler-report/). Access % is the share of AI crawlers a domain allows. Category and provider figures carry their domain count n and a Wilson 95% confidence interval, treating the measured set as a sample of the broader web. Nothing is modelled and no customer data is used.
Evidence class: measured. Source: 380 public domains, last crawled 2026-08-17. No customer data is used, and nothing on this page is modelled or estimated.