The thesis
Software is becoming the web’s primary reader — and the first reader that can be charged at the door.
Three things follow. Publishers will stop treating “bots” as one thing and start pricing the crawler that trains a model differently from the one that sends a reader. Access will become a transaction, denominated in something a machine can settle. And what a site declares will drift away from what its front door actually does, because charging means stopping people at the gate rather than asking them politely not to come in.
The rest of this page is the evidence — including the evidence that cuts against it. The index, twice a week, is how we find out. Our dated bets on where this goes, registered in advance and scored every edition by a script, are on /futures.
What becomes answerable
Questions that had no data behind them, now do.
Every claim on this page is somebody else’s measurement until this one: a pair of requests to the same door in the same sweep, one presenting a named AI crawler and one our own self-identified crawler.
- Is my sector walling up faster than the rest of the web?Block rates by crawler and by domain suffix, re-read twice a week across the top 50,000 domains.
- Which sites refuse a named AI crawler at the door, and which charge instead?A pair of answers — a request presenting a named AI crawler, and our own self-identified crawler (the honest-identity control) — for every domain in the wide probe (the top 2,000 by rank plus every domain blocking a tracked crawler; 7,491 reached this edition), every edition; every observed machine-readable price, and every payment code that quotes no price, recorded with the date it appeared.
- Did this domain change its AI policy the week that deal was announced?A dated history that cannot be back-filled — the one thing a competitor who starts later can never reconstruct.
- Is a licensing deal visible on the wire, or only in the press release?Identity-conditional access — sites that serve one AI crawler and refuse another — observed directly on a capped panel of 331 domains asked under several identities, and reported with its n.
- How many refusals come from the same front door?The count of wire refusals that arrive from Cloudflare-fronted hosts, edition by edition, beside the count that do not — on the Wire evidence tab, with the probed set as its base.
A reporter finds the story; an analyst finds the trend line; a publisher finds out where they stand; an AI team finds out what access will cost. The data is the same — the questions are yours. Three separate instruments see a price, and they are never summed: all three, side by side, with their n →
The shift
Reading the web is becoming a transaction, not a right.
From 15 September 2026, Cloudflare blocks training and agent crawlers by default on ad-bearing pages for newly onboarding domains, and offers publishers a pay-per-crawl toll in the same dashboard. TollBit already runs a bot paywall on thousands of publisher sites. Major outlets have signed private licensing deals with AI companies; others return a payment demand to any crawler that knocks.
None of this is speculative, and none of it is visible the way markets normally are. There is no ticker. A price quoted to an AI crawler appears for an instant in an HTTP response and is gone. A publisher's policy can change overnight and nobody outside that publisher would know.
The gate being built
Aug 2023
OpenAI publishes GPTBot with a robots.txt opt-out. Refusing an AI crawler becomes something a site can do.
Jul 2024
A study of 14,000 domains finds terms and robots.txt
contradicting each other.
Source
Jul 2025
Cloudflare begins blocking AI crawlers by default for new domains, with a pay-per-crawl toll in the same dashboard.
Sep 2025
Cloudflare says it serves
over a billion HTTP 402 responses a day, and auto-applies
ai-train=no to
3.8 million domains at once.
Source
Feb–Jun 2026
Microsoft opens a marketplace paying publishers for content used in answers; Mastercard ships payments built for fractions of a cent at machine speed.
15 Sep 2026
Cloudflare blocks training and agent crawlers by default on ad-bearing pages, leaving search crawlers through. It has not happened yet.
Each is a dated public announcement. Together they describe a gate, a price and a rail to settle on. None of them records what any individual site declares.
The numbers
The pool grows slowly. What is in it changes fast.
Three infrastructure companies measured the automated share of web traffic in overlapping periods and landed twenty points apart — Fastly 37% (Apr–Jul 2025), Thales/Imperva 53% (FY2025), Cloudflare ~57% (Jul 2026). All three sell bot mitigation. We show the spread rather than pick a headline, because the level is genuinely not settled.
What is consistent is direction, and there are two speeds in it. Measured by one source with one method across three annual reports, the automated share creeps up about a point and a half a year. Over roughly the same window, AI training crawlers went from 22% to 52% of all crawler activity.
That gap is the finding. The volume of machine traffic is not exploding. Its composition is changing quickly — from indexing that sent you readers, to training that does not.
How the two measurements fit together an analytical bridge, not a clean hierarchy
All web traffic100%
of which automated53%◆
of which crawler activityno source publishes this — the layer we estimateestimate ~52%45–60% by definition
of which requests whose purpose is AI traininga purpose, not a bot category — dual-purpose crawlers like Googlebot appear here too52%
These are
not a tidy set of nested populations, and we no longer draw them as one. The top figure is Imperva measuring all traffic including APIs; the bottom is Cloudflare classifying crawler requests
by purpose on its own network. Different companies, different baskets, different dimensions.
The middle layer is ours. Nobody publishes crawler activity as a share of automated traffic, so we estimate it from Cloudflare’s bot categories under three readings of the word “crawler”: search-engine plus dedicated AI crawlers gives ~45%; adding AI-search fetching gives ~52%; adding AI assistants gives ~60%. That spread is a
definition range, not a confidence interval — it reflects where the line is drawn, and better measurement will not narrow it.
We stop before multiplying the chain out. It is tempting to compound these into a single headline share of all web traffic for AI-training crawling. We do not publish that number, because the three layers do not share a denominator and the product would be false precision.
The workings, the sources, and a correction we had to make →
Automated share of traffic same question, three networks
FastlyApr–Jul 202537%
Thales / ImpervaFY202553%◆
CloudflareJul 2026~57%
A twenty-point spread on the same question. All three sell bot mitigation.
◆ This one figure also appears in the panel beside it — it is the last point of the Thales/Imperva series, the only one of the three measured the same way three years running.
Two speeds how fast each one moves
Automated share of all traffic
49.6%2023
51%2024
53%◆2025
+3.4pp in two years
Training-purpose share of crawler requests
+30pp in ~14 months
Both axes run 0–100%, so the heights are comparable even though the denominators are not. Left: Thales/Imperva, 2024–2026 reports. Right:
Cloudflare, Jul 2026. Spread sources:
Cloudflare ·
Thales/Imperva ·
Fastly.
Two cautions. In September 2025 Cloudflare expected bot traffic to pass human traffic “by the end of 2029”; in July 2026 it reported this had already happened — so either the shift accelerated by three years in nine months, or the basis changed. And growth is not uniform: Bytespider fell about 85% year on year, and ChatGPT-User volume fell quarter on quarter in 2026. The category grows; individual crawlers rise and fall sharply.
The motive
The trade that paid for the open web has stopped paying.
Everything above describes a gate being built. This is why anyone would want one. Search crawling was barter: you index my pages, you send me readers, I am paid in traffic. Nobody signed anything, but the exchange rate held for twenty years — Google still returns a visitor for roughly every 5 pages it takes.
Anthropic’s crawler takes 38,065. That is not a shifted exchange rate, it is a collapsed one. When the implicit payment stops arriving, a publisher has two options left: refuse, or charge. The refusal is written into robots.txt or answered at the door; the charge is set on the wire, in a 402 and its headers — both are the behaviour this index counts, twice a week.
The second series matters as much as the first. Restricted to news and publications, the same ratios fall by one to two orders of magnitude — still lopsided, nothing like the headline. Both are shown because quoting only the larger one would be exactly the kind of thing this index exists to avoid.
And a reader who arrives via an AI answer mostly does not click through: Pew measured 8% click-through on searches showing an AI summary against 15% without, with 1% clicking a link inside the summary.
Pages crawled per visitor referred · log scale
Anthropicall industries38,065:1
of which news & publications2,500:1
OpenAIall industries1,091:1
of which news & publications152:1
Perplexityall industries195:1
of which news & publications32.7:1
Googleall industries5:1
DuckDuckGoall industries0.3:1
Source: Cloudflare,
“The crawl-to-click gap” (29 Aug 2025, data Jan–Jul 2025) and
by purpose and industry (28 Aug 2025).
Axis: log scale, clamped at 0.1:1 at the left edge and 38,065:1 at full width; DuckDuckGo’s 0.3:1 is drawn against that clamp rather than at a sliver, so every bar is readable at its own value. The figures are the measurement; the lengths are a log of them.
Cloudflare’s own caveat: referrals from native apps carry no
Referer header, so these ratios overstate the gap by an unknown amount. Pew:
AI summaries and click-through (22 Jul 2025).
Evidence against
The parts of this thesis the public data does not support.
If we only published what agreed with us, none of the rest of this site would be worth reading. Three things cut the other way, and the weakest leg is the one the whole idea depends on.
- Machine-to-machine payment is barely real yet. Analysis of the x402 rail found roughly half of transactions were artificial — self-dealing and wash trading — with genuine daily volume around $28,000. Visa, a year into its agent-payments programme, reported agent-initiated transactions in the hundreds. The plumbing is being laid; almost nothing is flowing through it. (Chainalysis, Jun 2026; Artemis via CoinDesk, Mar 2026; Visa, Dec 2025)
- Blocking is a minority behaviour outside news. Only 10.6% of the top million sites fully block GPTBot, and as of mid-2025 only 14% of the top ten thousand domains carried any AI-crawler directive at all. The high numbers you read about are a news-sector story, not a web-wide one. (Bouchaud & Ramaciotti, arXiv 2025; Cloudflare, Jun 2025)
- Licensing revenue is real but small, and growing slower than the business it is meant to replace. Reddit’s entire data-licensing line — with the largest AI buyers as customers — was $43M in Q2 2026, about 5% of revenue, growing 24% while the company overall grew 61%. That is the clearest audited price signal in this market, and it is not yet paying anyone’s rent. (Reddit Q2 2026, SEC filing)
Not every crawler is growing either: Bytespider fell 85% year over year to May 2025, and training-bot volume on one publisher network declined 15% across the second half of 2025 even as retrieval and search bots grew. The category is expanding; individual crawlers are volatile, and some are shrinking.
The gap
A market that is mostly refusals still needs a neutral record. This one doesn't have one yet.
Oil had no reference price until Platts began publishing assessments; fertiliser had none until Argus did. The reference publisher takes money from no participant — which is exactly what lets everyone cite it. The crawl economy is at that pre-reference moment now: real transactions, no independent record.
What this is
An independent observatory, not a participant.
The Crawl Price Index reads the machine-readable policy of the top 50,000 domains twice a week and records what it finds: who blocks which AI crawlers, who refuses one at the door, who charges instead, what the price is, and what changed since the previous edition. It takes no money from anyone it measures and brokers no deals — the discipline that lets a price index be cited by the market it observes.
Headline figures are free to cite with attribution. The full per-domain record and the accumulating history are the subscription. The method is documented in full, the crawler is cryptographically identifiable, and every edition is dated and reproducible.