The Crawl Price Index.

Independent · twice weekly · edition of

The web is refusing AI crawlers far faster than it is pricing them.

Every edition we read what a ranked 50,000-domain frame tells 18 named AI crawlers in robots.txt, then knock on those doors presenting one of them and record the answer (thousands reached this edition). The counts for the current edition — refused, refused while serving our own identified crawler, and priced — are in the feed at /index.json and on Explore data. Every figure here states what it is a share of.

Start here

Most of what reads your website now is software. It does not all want the same thing.

A crawler is a program that fetches web pages automatically. There have always been crawlers — that is how search engines find anything — but what they are for has split in two, and the split is the whole reason this site exists.

A training crawler copies your page into the material a model is built from. You are not paid, and no reader arrives. A search crawler indexes your page so a person can find it, and sends that person to you. Same request, same protocol, completely different bargain.

A site answers both of them in one plain text file at its root, called robots.txt. It can name a crawler and say no. Almost nothing else on the web is as consequential and as little recorded.

Nobody was writing the answers down. So twice a week we ask 50,000 sites what they have said, one at a time, and keep the record. That record is this index.

One page, one gate, three kinds of reader
A person robots.txtdoes not apply your page
A search crawlerindexes it, may send a reader robots.txtallow or refuse your page
A training crawlercopies it into a model robots.txtallow or refuse your page
robots.txt is a request, not a lock. It states a policy; it does not enforce one. Whether the door behaves the same way is a separate measurement, and we keep the two apart everywhere on this site.
What we measure

What we count, and what we count it against.

Every edition we read the robots.txt of a ranked 50,000-domain frame and record what each domain declares to 18 named AI crawlers: blocked, allowed, partial, no instruction, or no file at all. Rates are quoted against the domains that actually serve a readable file — never against “the web”.

domains serving a readable robots.txt
18
AI crawlers tracked by name
domains in the ranked frame

 

The declared layerEvery edition: 50,000 ranked domains × 18 named AI crawlers, read one domain at a time from robots.txt. Five states, never collapsed into “blocked or not”.
Is the domain even thereThe same frame is swept for whether it answers at all — alive, dead name, timeout, or serving a file then refusing an identified crawler at the door.
What the door actually doesThousands of those domains are knocked on twice in one sweep — once presenting a named AI crawler, once as our own self-identified crawler (the honest-identity control) — and the two answers are recorded side by side: served, refused, or the payment code used to refuse.
Who names a priceA public registry of endpoints advertising a price a machine can pay directly — captured alongside, at its honest size, because it is the same question asked from the other side.

Every rate on this site is quoted against the domains that actually served a readable file — never against “the web”. Full methodology →

This edition

Told no, refused at the door, or priced — three figures, in plain English.

The three figures are read from the current edition’s feed when the page loads; without scripts, the same figures with their denominators are at /index.json.

Declared policy first, then what the wire showed. These are the figures the rest of this page unpacks — and they are free to cite with attribution.

Training versus traffic

Sites treat the crawler that trains differently from the one that sends a reader.

This is the most economically loaded thing in the dataset, and it does not show up in a headline block rate. A crawler that ingests your work to train a model and a crawler that might send you a visitor are different propositions — and domains that draw a line overwhelmingly draw it in one direction.

Declared robots.txt policy only. Role tags describe each crawler’s stated function; this is not evidence of intent, and not proof any crawler was denied.

Where to go next

Three ways in, depending on what you came for.