How we measure
What we measure, how we measure it, what counts as an observation versus an inference, what we deliberately do not measure, and how to check our work.
Twice a week — Sunday and Wednesday — we read the robots.txt and homepage response headers of a ranked 50,000-domain frame using our own identified, cryptographically signed crawler. We record what each site tells AI crawlers: who is blocked, who is allowed, whether a price or payment wall is quoted, and what changed since the previous edition. We publish the aggregates free and the per-domain detail to subscribers. The census runs twice a week; the newsletter mails the Sunday edition, and its change window runs from one Sunday edition to the next.
robots.txt — never of the frame, never of “the web”.robots.txt with our signed, identified crawler.corrections array; past editions are never silently amended.| Signal | How | Type |
|---|---|---|
| Robots stance per crawler, per domain | GET /robots.txt with our identified user agent; parse User-agent / Disallow / Allow groups. Values: blocked, allowed, partial, unlisted, no_robots. | observed |
| Observed price | Recorded only when a site returns a price in a machine-readable payment response — HTTP 402 with a crawler-price or payment header, or an equivalent x402 offer. Where a price is printed we print the quote and the domain that gave it: this edition that is the identity-matrix panel’s one posted price, $0.50 at stackoverflow.com. The wide probe’s protocol-confirmed prices are a different instrument and are published as a count only — the home page’s “named a price” figure — so the quotes and domains behind that count are not on this site. Counted, not quoted; the two are never summed. | observed |
| Payment signals | Status codes and headers on identified-crawler requests: 402 responses, TollBit-style token walls, licensing redirects, payment:free declarations. | observed |
| Enforcement | Compares a site’s declaration against what an identified crawler actually receives. Wide probe: 7,491 domains reached, on a set built to over-represent blockers, so a rate over that set describes the set and nothing wider. | observed (wide probe) |
| Block rates | Share of parsed domains blocking each crawler. Denominator is domains successfully parsed, never the nominal top-N. | derived |
| Suffix groups | Segmented by country-code TLD suffix. Every group with at least 8 domains is listed with its n; a rate is published only where at least 100 domains are present, on the public site and in the Terminal alike. A ccTLD is a domain suffix, not a country — not operator location, ownership, audience or hosting. | derived |
| Trends & movers | Difference between editions in our archive — the previous one, or any baseline the reader chooses. | derived |
Our crawler sends a truthful user agent and signs every request cryptographically, so any site can verify that traffic claiming to be ours really is:
User-Agent: CrawlPriceIndexBot/1.0 (crawl-economy measurement;
robots.txt study; contact: hello@crawlpriceindex.com)
Signature-Agent: "https://crawlpriceindex.com"
Signature-Input: sig=("@authority" "signature-agent");created=…;keyid="…";tag="web-bot-auth"
Signatures follow RFC 9421 under the Web Bot Auth profile. Our public key directory is served at /.well-known/http-message-signatures-directory. We have applied to Cloudflare’s verified bots programme using signature verification.
The wide probe above covers the whole reached frame and presents one named crawler, which is why a wall that fires at ClaudeBot and not at GPTBot is invisible to it. The identity matrix makes the opposite trade: 331 domains rather than tens of thousands, asked under six identities apiece. That is where ClaudeBot is asked, and the two instruments should be read as the pair they are rather than as one claim.
A sixth identity starts on 17 September 2026, and why it exists — and why it waits — is worth stating. The fifth identity — ClaudeBot's user-agent string carrying crawler-max-price: 0.001 — differs from the plain ClaudeBot leg in two ways at once: it is an impersonated identity and it carries an offer to pay. When a door answered it differently, nothing in the data could say which of the two did it. The sixth is our own CrawlPriceIndexBot/1.0 carrying that same header at that same value, so the price signal is isolated from who is asking. It adds no further impersonation: it is the crawler we already declare, with one header added. It starts four days after it was ruled, deliberately. The identity matrix keeps no archived per-identity distribution, so a leg switched on without a five-leg edition beside it has no control to be read against — and a before-and-after that is failing for other reasons in both runs is not a control. The 13 September census runs the five legs and records their distribution; the 17 September census starts the sixth. It costs one edition to have a real comparison.
Its readings are archived from the day it starts. Its level is published from its first edition, with the number of editions it has accrued stated beside it; no edition-over-edition movement is claimed for it until its own series is long enough to support one. Measuring a site twice a week from September and telling that site what was found in October is not a defensible posture toward a measured party, and hiding a level that is already precise buys nothing.
The five, named in full, because “several identities” does not tell a site operator what knocked on their door — one request per identity per domain per edition:
CrawlPriceIndexBot/1.0, signed as above — the honest-identity control;crawler-max-price: 0.001.The fifth identity states a price a crawler would be prepared to pay, in the units that header defines. It is sent to observe whether a door that refuses a named crawler answers the same request differently when it carries that header; where the two answers differ, the difference is the measurement, and it is published in the Terminal as the max-price probe. No payment is offered in any binding sense and none has ever been made. Nothing in this pipeline can settle a payment: there is no wallet, no credential and no code path that could complete one. The header is a declaration a door may answer or ignore.
Two things follow that we would rather state than be asked. Presenting another operator’s user-agent string means the response we record is the response that string received, not one that operator received — we are not those crawlers and we do not claim their traffic. And the request originates from an ordinary consumer connection, not from any published crawler address range, so a door that verifies its callers by address will treat these knocks as unverified, which is itself part of what the matrix measures. Named on this page 10 September 2026; the matrix itself is older, and the disclosure should have arrived with it.
robots.txt and the homepage only. We do not crawl article pages and we store no page content.We measure crawler policy two ways, on purpose. The wide probe visits the top 2,000 domains by rank plus every domain that blocks any tracked crawler in robots.txt — it asked 8,002 domains in the current edition and reached 7,491 of them, meaning in the frame and answering; a timeout or dropped socket on either leg is recorded as its own outcome. Asked and reached are the two words this site uses for those two quantities, on every page and beside every feed field that carries one; every wire count below is over the reached set. Each target receives two homepage requests in one sweep: one presenting the published user-agent string of a named AI crawler (GPTBot in this edition), and one as our own self-identified crawler (the honest-identity control). It records the HTTP status of each, payment or licensing headers, and Cloudflare fronting. It retains no content. Since it sends another crawler's user-agent string it is a disclosed measurement study: one request per identity per domain, from one residential network position, once per edition, never repeated to obtain access. The census sends its knock once per door per edition, redirects followed, whether or not the door disallows the crawler it presents, and never repeats it; a second instrument, the 331-domain identity matrix below, presents the same string once more at the doors it shares with the census. What it yields is a pair of answers per domain, never a cause — a firewall, a rate limiter, a missing browser signature and a deliberate decision all produce the same pair, and the product says so wherever the numbers appear.
CrawlPriceIndexBot/1.0 — our own self-identified crawler, the honest-identity control. On this instrument that control request is unsigned: it carries the truthful user-agent string and no RFC 9421 signature (the census fetcher and the reachability sweep are the signed ones, see how we identify ourselves). Until 4 September 2026 this page, the home page, Explore and the Terminal described that control as “an ordinary browser”. It was not one, and it never had been; the count of domains that refused the crawler while serving the control was always computed against the identified-crawler request, so no published count changed. The correction is logged in the changelog and in the machine-readable corrections array at /corrections.json. A third request per domain carrying a current Chrome user-agent string — not a real browser: no JavaScript, no browser TLS fingerprint — has run since the 2026-09-06 edition, the first edition whose archived probe records a browser identity. It runs where the first knock was answered, and its count is printed as its own line, never substituted into the pair above.
The identity matrix is the older, smaller instrument: a capped panel of 331 domains asked under several identities — a request carrying a current Chrome user-agent string (not a real browser: no JavaScript, no browser TLS fingerprint), our own identified crawler, and the published user-agent strings of several AI crawlers — to observe whether a site answers different crawlers differently. It remains the source for the per-crawler exhibits (posted prices, token walls, payment headers) and is reported with its n every time.
The panel is assembled, not hand-picked. It has three parts: a fixed spine carried unchanged for continuity; a signal tier, where any domain the wide probe flags as showing payment or blocking behaviour is promoted automatically the following edition and dropped after eight silent editions; and a rotating audit, a small random sample from the full scan each edition, so that a wall which fires only at named AI crawlers is measured on a set that was not selected for it. (Until the 2026-09-02 edition the audit tier was the only instrument that could see such a wall at all, because the wide probe then ran under our own identity only; the wide probe now meets those walls itself, and the audit tier's job is the unselected denominator.) The panel's full composition is published at /panel.json, which names the identities each instrument presents.
*.in-addr.arpa, 167 rows) and nothing else — CDN, cloud and parked domains stay in, because how much of a popularity ranking is not a website is a finding, not a defect to hide. Full provenance, including every removed row, is published at frame-cpi50k-v1.json. Alongside the frame we run two probes: the wide probe, which this edition paired the 7,491 domains it reached with our own self-identified crawler (the honest-identity control) under one presented identity, and the identity-matrix panel of 331 domains asked under several (see above).robots.txt; that parsed count is the denominator behind every rate we publish, and it is the same figure the dashboard uses. We never quote a rate as a share of “the web”. The remainder is not a rounding error and we census it explicitly — see Frame reachability below.The census runs twice a week — the run starts at 01:00 Europe/Luxembourg on Sunday and Wednesday (23:00 UTC the evening before) and finishes around 03:00 UTC; the finish time is stamped on each edition and printed on status. Only scheduled census runs are published as editions. A measurement can be real and still not be an edition: on 24 August 2026 a full scan was taken at 21:29 UTC on a Monday to validate the pipeline’s move onto its new census machine (31,562 domains parsed). The data is sound and the file and its timestamp anchor are retained, but publishing it alongside the scheduled editions would imply a Monday census that does not exist — so it is excluded from the published series, from the trend chart, from the comparison baselines and from every change set. The exclusion is recorded in series.json under unscheduled, with its reason, rather than being applied silently.
A ranked list keeps domains long after the website behind them stops answering. Because every rate we publish carries a denominator, we measure that directly rather than assume it. Every edition we attempt robots.txt and the homepage for all 50,000 domains, trying the apex over HTTPS, then www, then HTTP, and record the outcome by reason. Measured on edition 2026-09-02 across the full frame:
Anatomy of the frame breaks the domains that served no readable robots.txt into a different set of reasons (unresolved, TLS, timed out, refused, no file, not a robots file). That decomposition comes from the robots.txt fetch itself; the census above comes from the reachability sweep, which also tries the homepage and the www host. The two are different sweeps with different failure classes, so their buckets are not meant to match one another; each sums to its own stated total.
Declared policy and enforced access are different measurements, and we keep them apart. 4,993 domains (10.0% of the frame) answered and then returned 403 or 429 to the same identified crawler at the homepage. That is a divergence between what a site declares to crawlers and what its edge actually does. It is not a robots.txt policy, it is not evidence of intent, and it is never folded into a block rate. A further 1,041 domains disallow our crawler in robots.txt. On the reachability sweep we obey that: the sweep reads robots.txt first, and where it disallows us the homepage is not fetched and the domain is recorded as excluded by its own instruction — a publisher of a compliance index does not get to make exceptions for itself.
Correction, 4 September 2026 — the wide probe does not apply that exclusion. Until today this page described the exclusion as if it held across the product. It does not. The wide probe builds its target list from the census as “the top 2,000 by rank, plus every domain that blocks a tracked crawler in robots.txt”, and that rule contains no robots.txt check of its own: a domain that disallows our crawler can be a probe target, and its homepage is requested once on each leg. So the wide-probe figures on this site — the reached set, the refusals, and the 402s — are computed over a target list from which those domains were not removed in advance. The reachability sweep and the wide probe are different instruments with different politeness rules, and this page now says which is which. Logged in corrections.json.
Timeout is defined as no response within 8 seconds, with a 22-second absolute ceiling per domain. Thresholds are stated because the choice changes the answer.
Alongside robots.txt we record a small set of self-declared signals a domain serves publicly on its own homepage. These are facts about observable artefacts, dated and attributable, in the same spirit as reading a robots.txt. Measured on each reachability sweep across the full 50,000-domain frame (the sweep can lag the robots edition by one run); the figures below are the latest sweep’s, read from the public feed (first measured 20 August 2026, before the published series began):
ads.txt file — an IAB standard declaring authorised digital sellers.What we do not do with them. We publish what a domain declared and when we observed it — never an inferred label about what kind of organisation it is. We do not publish machine-generated category guesses. We do not publish inferred labels in sensitive categories. Where a site declares an adult content rating (RTA or an equivalent meta tag), that is reported only in aggregate, never as a per-domain field. Absence of a signal means we did not observe one; it is not evidence that the thing is absent.
Every rate elsewhere on this site is stated over a population we defined and can defend. A cohort is the one exception: a subscriber picks the domains, and from the Terminal a cohort can be turned into an anonymous link anyone can open. Three rules apply to a cohort’s figures that the rest of the product does not need, and they are enforced in one place rather than per surface.
The arithmetic runs on the census machine, against the same per-edition archive the rest of this page describes, and the API serves what it computed — so a cohort created between editions has no reading until the next census, and says so rather than showing a number computed a different way.
Content-Signal line and a License: directive, which places us
inside the small group of domains that declare terms about AI use — a group we also
count. We disclose that rather than exempting ourselves, because a measurer who removes
inconvenient observations has already conceded the point. No domain is ever removed from
the frame or the counts on request; a domain can ask not to be named publicly (see crawl etiquette above)..io domain is rarely British Indian Ocean Territory, and .com/.org/.io carry no geography at all. Every published group, with the reasons a group’s rate moves set out before the numbers: suffix groups.vantage.json, which the pre-flight check asserts before every edition.If we get something wrong, we fix it in the next edition and record it in the machine-readable corrections array with the date, what changed and why. Three renderings of one file: the table on /status, the JSON at /corrections.json, and the corrections key in the free feed at /index.json and in the licensed dataset. We do not silently amend past editions. Report an error to hello@crawlpriceindex.com and it will be acknowledged in the next edition.
Every figure we publish is regenerated from a full sweep, every edition. To check any headline number yourself:
# the aggregates behind the homepage curl -s https://crawlpriceindex.com/index.json # how we measure, machine-readable curl -s https://api.crawlpriceindex.com/v1/methodology # any in-frame domain's stance, free curl -s "https://api.crawlpriceindex.com/v1/check?domain=example.com" # and the site's own robots.txt, to check us against the source curl -s https://example.com/robots.txt
If our stance for a domain disagrees with that site’s live robots.txt, that is a bug and we want to hear about it. Subscribers get the full per-domain dataset as JSON or CSV, with a per-customer watermark.
The Crawl Price Index is funded by subscriptions. We take no payment from AI companies, publishers, CDNs or licensing platforms in exchange for coverage, ranking or inclusion. We link to Cloudflare, TollBit and RSL as setup routes for site owners with no commercial relationship to any of them. Our own site publishes machine-readable licensing terms and names a price for AI training crawlers in rsl.xml. That figure is our own commercial choice for our own pages; it is not modelled from the index, and the index never quotes it.