Coverage · edition of 2026-09-23
Every figure this index publishes is a share of the 30,712 domains that served a readable robots.txt in the 50,000-domain frame — not of 50,000, and not of the web. This page is where that number comes from, and what is inside it.
The scanner records a reason for every domain it cannot read. These 9 groups cover 75 distinct raw reasons and sum to 19,288 exactly, with nothing left over.
| Reason | Domains | Of the 19,288 | |
|---|---|---|---|
| Name never resolved, or the connection was refused | 4,167 | 21.6% | |
| Refused an identified crawler (403 or 429) | 4,092 | 21.22% | |
| No robots.txt file (404 or 410) | 3,748 | 19.43% | |
| Answered 200, but the body was not a robots.txt | 2,507 | 13% | |
| Connected, then timed out | 1,859 | 9.64% | |
| TLS would not negotiate | 1,749 | 9.07% | |
| Connection dropped mid-transfer | 515 | 2.67% | |
| Other HTTP status | 359 | 1.86% | |
| Server error (5xx) | 292 | 1.51% |
These branches are computed from the census itself. Each level sums to the one above it.
| Branch | Domains | Of 30,712 |
|---|---|---|
| Block none of the 18 | 25,770 | 83.91% |
| Name none of the 18 at all | 24,005 | 78.16% |
| Named one and explicitly allowed it | 939 | 3.06% |
| Path-scoped rule only, never the whole site | 826 | 2.69% |
| Block at least one | 4,942 | 16.09% |
| Search and machine-consumption alike | 2,712 | 8.83% |
| Machine-consumption only, search left open | 2,132 | 6.94% |
| Search only, machine-consumption left open | 98 | 0.32% |
Blocking a single named crawler and nothing else cannot be a copied snippet or a hosting default. Someone chose that one.
| Crawler | Sole target | Share | |
|---|---|---|---|
| GPTBot | 279 | 27.57% | |
| CCBot | 223 | 22.04% | |
| Bytespider | 209 | 20.65% | |
| Amazonbot | 94 | 9.29% | |
| meta-externalagent | 77 | 7.61% | |
| ClaudeBot | 36 | 3.56% | |
| Google-Extended | 33 | 3.26% | |
| Applebot-Extended | 25 | 2.47% | |
| 7 other crawlers | 36 | 3.56% |
| Crawlers blocked | Domains | Of blockers |
|---|---|---|
| 1 | 1,012 | 20.48% |
| 2–7 | 1,541 | 31.18% |
| 8 — the copied list and its variants | 212 | 4.29% |
| 9–17 | 1,946 | 39.38% |
| All 18 | 231 | 4.67% |
Raw-body wildcard analysis is unavailable for this edition in this view. This is a coverage gap, not a finding of zero wildcard restrictions. The named-crawler figures above come from the parsed census and remain available.
30,712 domains × 18 crawlers is 552,816 individual declarations.
| State | Cells | Of 552,816 | Meaning |
|---|---|---|---|
| unlisted | 500,885 | 90.61% | The file exists and does not mention this crawler |
| blocked | 37,768 | 6.83% | Named, with Disallow: / |
| allowed | 7,567 | 1.37% | Named, and explicitly permitted |
| partial | 6,596 | 1.19% | Named, with a path-scoped disallow only |
| CPI-50K v1 rank band | Parsed | Block ≥1 | |
|---|---|---|---|
| 1-100 | 66 | 27.27% | |
| 101-500 | 271 | 30.63% | |
| 501-1000 | 352 | 24.15% | |
| 1001-5000 | 2,752 | 20.68% | |
| 5001-10000 | 3,184 | 17.81% | |
| 10001-25000 | 9,210 | 15.9% | |
| 25001-50000 | 14,877 | 14.49% |
The census figures above are computed from the edition of 2026-09-23, from the census file each time this page is built. Raw-body wildcard analysis is unavailable in this view for this edition. These figures describe declared robots.txt policy, not observed crawler access.
The per-domain rows behind these counts are the subscription. The shape is free, and free to cite with attribution.