The Crawl Price Index.

Coverage · edition of 2026-09-23

Anatomy of the frame

Every figure this index publishes is a share of the 30,712 domains that served a readable robots.txt in the 50,000-domain frame — not of 50,000, and not of the web. This page is where that number comes from, and what is inside it.

Frame
50,000
ranked domains
Readable robots.txt
30,712
61.42% — the denominator
No reading
19,288
38.58% — broken out below
Block ≥1 crawler
4,942
16.09% of the denominator
50,000 in the ranked frame30,712 served a readable robots.txt19,288 no reading10,347 answered8,941never gave a usable responserefused, no file, or not a robots.txt53.64% of the group
Widths are proportional to domain counts. The domains that answered a request and still returned no robots.txt outnumber those that never returned a usable response — and every branch below is measured by the scanner, not inferred.

Why 19,288 domains told us nothing

The scanner records a reason for every domain it cannot read. These 9 groups cover 75 distinct raw reasons and sum to 19,288 exactly, with nothing left over.

ReasonDomainsOf the 19,288
Name never resolved, or the connection was refused4,16721.6%
Refused an identified crawler (403 or 429)4,09221.22%
No robots.txt file (404 or 410)3,74819.43%
Answered 200, but the body was not a robots.txt2,50713%
Connected, then timed out1,8599.64%
TLS would not negotiate1,7499.07%
Connection dropped mid-transfer5152.67%
Other HTTP status3591.86%
Server error (5xx)2921.51%
10,347 of them answered the door and still said nothing — 53.64% of the group, more than everything that failed to connect (8,290), with 651 more that connected and answered a server error. A 404 is not the same as an unlisted crawler: 3,748 domains have no robots.txt at all and have therefore made no declaration about anything, while an unlisted crawler means a file exists and does not mention it.

Inside the 30,712

These branches are computed from the census itself. Each level sums to the one above it.

30,712 with a readable robots.txt25,770 block none of the 18 · 83.91%4,942 block ≥12,712search AND machine2,132machine only98search only
Grouping: Cloudflare’s — training and user-initiated crawlers together as machine consumption, against search. The home page’s asymmetry card uses the index’s own role tags (training against search, user-initiated on neither side), which is why it reads 2,121 / 99. Four in five domains that publish a robots.txt block none of the 18 tracked crawlers. Of those that do block, most draw no distinction between a crawler that sends readers back and one that does not.
BranchDomainsOf 30,712
Block none of the 1825,77083.91%
Name none of the 18 at all24,00578.16%
Named one and explicitly allowed it9393.06%
Path-scoped rule only, never the whole site8262.69%
Block at least one4,94216.09%
Search and machine-consumption alike2,7128.83%
Machine-consumption only, search left open2,1326.94%
Search only, machine-consumption left open980.32%

The 1,012 that block exactly one crawler

Blocking a single named crawler and nothing else cannot be a copied snippet or a hosting default. Someone chose that one.

CrawlerSole targetShare
GPTBot27927.57%
CCBot22322.04%
Bytespider20920.65%
Amazonbot949.29%
meta-externalagent777.61%
ClaudeBot363.56%
Google-Extended333.26%
Applebot-Extended252.47%
7 other crawlers363.56%

How many crawlers each blocker blocks

Crawlers blockedDomainsOf blockers
11,01220.48%
2–71,54131.18%
8 — the copied list and its variants2124.29%
9–171,94639.38%
All 182314.67%
Only 231 domains block every tracked crawler without exception — 0.75% of everything with a readable robots.txt. Blanket refusal, as a stated policy, is rare.

Wildcard policy coverage

Raw-body wildcard analysis is unavailable for this edition in this view. This is a coverage gap, not a finding of zero wildcard restrictions. The named-crawler figures above come from the parsed census and remain available.

At the level of individual states

30,712 domains × 18 crawlers is 552,816 individual declarations.

StateCellsOf 552,816Meaning
unlisted500,88590.61%The file exists and does not mention this crawler
blocked37,7686.83%Named, with Disallow: /
allowed7,5671.37%Named, and explicitly permitted
partial6,5961.19%Named, with a path-scoped disallow only
Nine declarations in ten are silence. Most sites have said nothing about most AI crawlers, and that is the most honest one-line summary of the declared-policy layer.

By rank band

CPI-50K v1 rank bandParsedBlock ≥1
1-1006627.27%
101-50027130.63%
501-100035224.15%
1001-50002,75220.68%
5001-100003,18417.81%
10001-250009,21015.9%
25001-5000014,87714.49%
Blocking declines with rank below the top thousand. This gradient is what reconciles our headline with indices that sample deeper: extend this frame to 100,000 along the observed curve and it predicts a materially lower rate.

How to read this page

The census figures above are computed from the edition of 2026-09-23, from the census file each time this page is built. Raw-body wildcard analysis is unavailable in this view for this edition. These figures describe declared robots.txt policy, not observed crawler access.

The per-domain rows behind these counts are the subscription. The shape is free, and free to cite with attribution.