Independent · twice weekly · edition of —
Every edition we read what a ranked 50,000-domain frame tells 18 named AI crawlers in robots.txt, then knock on those doors presenting one of them and record the answer (thousands reached this edition). The counts for the current edition — refused, refused while serving our own identified crawler, and priced — are in the feed at /index.json and on Explore data. Every figure here states what it is a share of.
A crawler is a program that fetches web pages automatically. There have always been crawlers — that is how search engines find anything — but what they are for has split in two, and the split is the whole reason this site exists.
A training crawler copies your page into the material a model is built from. You are not paid, and no reader arrives. A search crawler indexes your page so a person can find it, and sends that person to you. Same request, same protocol, completely different bargain.
A site answers both of them in one plain text file at its root, called robots.txt. It can name a crawler and say no. Almost nothing else on the web is as consequential and as little recorded.
Nobody was writing the answers down. So twice a week we ask 50,000 sites what they have said, one at a time, and keep the record. That record is this index.
Every edition we read the robots.txt of a ranked 50,000-domain frame and record what each domain declares to 18 named AI crawlers: blocked, allowed, partial, no instruction, or no file at all. Rates are quoted against the domains that actually serve a readable file — never against “the web”.
robots.txt. Five states, never collapsed into “blocked or not”.Every rate on this site is quoted against the domains that actually served a readable file — never against “the web”. Full methodology →
The three figures are read from the current edition’s feed when the page loads; without scripts, the same figures with their denominators are at /index.json.
Declared policy first, then what the wire showed. These are the figures the rest of this page unpacks — and they are free to cite with attribution.
This is the most economically loaded thing in the dataset, and it does not show up in a headline block rate. A crawler that ingests your work to train a model and a crawler that might send you a visitor are different propositions — and domains that draw a line overwhelmingly draw it in one direction.
Declared robots.txt policy only. Role tags describe each crawler’s stated function; this is not evidence of intent, and not proof any crawler was denied.