← The Crawl Price Index

Methodology

How the numbers are made — and what they do and don't mean.

What we collect

Twice a week, we request two things from each domain in scope: its robots.txt file, and its homepage. We record the HTTP status code, response headers, and — only for 402 and 403 responses — a short snippet of the body used to classify the paywall type. We store no article content. The homepage request is made under several documented identities so we can see whether a site treats a browser, a named AI crawler, and a self-identified research agent differently.

Scope

The identities we present

Bulk robots.txt collection uses a single honest, self-identified user-agent with contact information — robots.txt is public by design and meant to be read by agents. The small, fixed publisher panel additionally probes with documented crawler user-agents to observe differential treatment. We do not attempt to defeat bot detection, solve challenges, or impersonate verified crawlers to obtain access we aren't entitled to. A future verified-crawler registration will let us receive genuine price quotes through official channels.

How we classify a signal

Priced

A 402 response carrying an explicit price, e.g. Cloudflare pay-per-crawl's crawler-price header and an x402-format body naming an amount and asset.

Marketplace-gated

A 402 demanding a marketplace token, fingerprinted by headers such as x-tollbit-forwarded.

Licensing 402

A payment-required response routing to a licensing sales contact rather than an automated price.

Declared-free / bot-block / robots.txt opt-out

Respectively: an explicit header stating no charge; content served to browsers but refused to AI user-agents; and the declarative Disallow directives in robots.txt.

Limits & honest caveats

A robots.txt block is a request, not a wall. It records a site's stated preference; it does not prove the site enforces it, and independent research shows a meaningful share of stated blocks go unenforced. Our block-rate figures measure declarations, and we label them as such.

Corrections & contact

We publish a dated changelog for methodology changes and correct errors openly. To report an issue, request bulk access, or ask us to verify a signal on your domain: hello@ (domain TBD).