Skip to content

Latest commit

 

History

79 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

Agent Web Index — how much of the web can AI assistants actually read?

50,256 domains measured live. 26% of them cannot be read by at least one of ChatGPT, Claude, Perplexity or Gemini. Updated daily. Live index: https://shop.lumnika.com/ai-readiness/

Every row here is the result of real HTTP requests, not an estimate and not a re-publication of someone else's crawl: each domain's homepage is requested once as a browser and once as each of the published AI crawler user-agents, its robots.txt, llms.txt and sitemap are read, and the answers are compared.

The number that exists nowhere else

robots.txt says the crawler may come in — and the server refuses it anyway. You can only see this by making the request as that crawler, which is why no robots.txt study reports it.

Crawler Domains measured Served robots.txt says no Server says no anyway
claudebot 50,256 84% 2,049 7,178
gptbot 50,256 85% 2,600 6,842
oai-searchbot 50,256 90% 793 4,741
perplexitybot 50,256 90% 1,207 4,607
meta-externalagent 41,514 89% 1,161 4,188
amazonbot 41,514 85% 1,261 5,561
bytespider 24,805 86% 660 3,182
applebot 24,805 95% 100 1,112

Broken down by whoever answers in front of the site (read from the response headers of the same request: cf-ray, akamai-grn, x-amz-cf-id, x-fastly-request-id…). Unit: one domain × one crawler, counted only where that site's robots.txt allows that crawler — so every refusal below contradicts the site's own stated policy, and is almost always an edge default nobody chose.

Edge in front of the site Domains Requests robots.txt allows Refused anyway
Akamai 791 4,396 38%
Google 805 5,093 33%
Sucuri 78 506 22%
AWS CloudFront 3,200 17,721 16%
no known edge 12,821 75,070 11%
Azure Front Door 314 1,639 11%
Cloudflare 28,235 197,517 10%
DDoS-Guard 310 1,794 10%
Varnish 418 2,187 9%
Fastly 1,536 8,205 9%

no known edge is an upper bound, not a vendor: response headers are not kept, so a domain read before a signature was added to the table stays in that bucket until it is re-requested (a 250-domain sample on 19 Sep 2026: 14% already carry a signature the current table recognises). A weekly pass re-reads them, so the named vendors above are undercounts.

Files

File What it is
agent-web-index.csv one row per domain: score, grade, per-check breakdown, per-crawler verdict, edge vendor, date
aggregate.json today's aggregate, exactly as the live index publishes it
daily/<date>.json the immutable daily snapshot of the aggregate — this directory is the series

Method, and what it does not cover

  • One vantage point (Europe), one page per domain (the homepage), 12-second timeout per request.
  • Rank-bearing domains come from the Tranco research list (30-day average of five rankings); the rest come from public e-commerce platform seeds.
  • Google-Extended and Applebot-Extended never make requests — they are robots.txt opt-out tokens — so for those only robots.txt is reported and no "served" verdict exists.
  • Domains that answer with HTTP 429 are excluded from that pass rather than counted as blocking.
  • 16,011 measured hosts are infrastructure (CDNs, telemetry, resolvers) rather than sites and are counted apart; 22,689 failed to answer and are excluded from every percentage.
  • Blocking AI crawlers is a legitimate choice, not a failure. This dataset records what is true, not what should be.

Citing

Agent Web Index, 2026-10-06. 50,256 domains. https://shop.lumnika.com/ai-readiness/

Licence: CC BY 4.0 — use it, say where it came from.

Audit a single site live, for free, no signup: https://aivis.lumnika.com/en?src=agentindex#scan

About

How much of the web can AI assistants actually read? Daily measurement of every crawler, domain by domain.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors