Twice a week, we request two things from each domain in scope: its robots.txt file, and its homepage. We record the HTTP status code, response headers, and — only for 402 and 403 responses — a short snippet of the body used to classify the paywall type. We store no article content. The homepage request is made under several documented identities so we can see whether a site treats a browser, a named AI crawler, and a self-identified research agent differently.
Bulk robots.txt collection uses a single honest, self-identified user-agent with contact information — robots.txt is public by design and meant to be read by agents. The small, fixed publisher panel additionally probes with documented crawler user-agents to observe differential treatment. We do not attempt to defeat bot detection, solve challenges, or impersonate verified crawlers to obtain access we aren't entitled to. A future verified-crawler registration will let us receive genuine price quotes through official channels.
A 402 response carrying an explicit price, e.g. Cloudflare pay-per-crawl's crawler-price header and an x402-format body naming an amount and asset.
A 402 demanding a marketplace token, fingerprinted by headers such as x-tollbit-forwarded.
A payment-required response routing to a licensing sales contact rather than an automated price.
Respectively: an explicit header stating no charge; content served to browsers but refused to AI user-agents; and the declarative Disallow directives in robots.txt.
robots.txt block is a request, not a wall. It records a site's stated preference; it does not prove the site enforces it, and independent research shows a meaningful share of stated blocks go unenforced. Our block-rate figures measure declarations, and we label them as such.We publish a dated changelog for methodology changes and correct errors openly. To report an issue, request bulk access, or ask us to verify a signal on your domain: hello@ (domain TBD).