2026-09-10 · product
The browser-UA challenge count is corrected, and its cut now sums
The Wire tab's browser-UA cut prints three rows under the heading “the same 3,223, cut by what the browser-UA leg got”. The CHALLENGED row was counted over every in-frame row rather than over the refusals it is drawn inside, so it published 23.8% where 22.8% is true — 768 rows against 735 on the 2026-09-09 edition. The counter now increments inside the refusal branch, and a dated correction sits beside the figure. An earlier internal review proposed the other repair — re-basing the row on the reached set, 768 of 7,496 = 10.2% — and it is not taken. Measured on the census machine, the browser-UA groups partition the refusals exactly: 1,944 served, 735 challenged, 529 refused too and 15 with no browser status, summing to 3,223. Re-basing would take the row out of the set it is drawn inside and the cut would sum to nothing. The label was right; the numerator was not. The fourth group had no row at all, so the cut never visibly summed to the number in its own heading. It is now drawn — derived rather than emitted, exactly as the control cut derives its own third rung — with the arithmetic printed beside it. Nothing is newly measured and the method is unchanged, so the methodology version does not move; a corrected miscount carries a correction, not a version.
2026-09-10 · product
401 is counted and exported: the authentication wall gets a row, two columns, and a back-computed series
The wide probe has recorded a 401 on every row since the series began and nothing ever counted it. That is now fixed in three places, and the counts are NOT new: wire-history.json stores the raw status per domain per edition, so the whole published series was back-computed rather than started from today — seven domains on the control leg in every edition from 23 August, stable to the row, and five on the named-crawler leg from 2 September, the edition that leg began. Exactly the position wire_control_p402 was in on 5 September: in the free history file since the beginning and never exported. (1) The Wire tab's refusal cut now has a 401 row. It was not a missing panel, it was a broken partition: the cut listed 403, 402, 429, 451, 406 and 404, so every authentication challenge fell into the bucket labelled “other 4xx, 5xx and non-standard codes (e.g. 999)” — a label that tells a reader the opposite of what those rows are. The row states the distinction the number exists for: 403 is a refusal, 401 is a condition of entry, and the two payment-shaped codes are never added, because 402 states a price and 401 states a credential. (2) The aggregate series CSV emits wire_ai_p401 and wire_control_p401, both over wire_reached, the same denominator as wire_ai_p402. The ai-leg column is EMPTY, never 0, in an edition whose probe carried no named-crawler leg: a zero there would claim a measurement nobody took. (3) COLUMN POSITIONS MOVE, and this is the part worth reading twice. The two new columns sit beside their 402 siblings rather than at the end of the file, so the export reads as two pairs and the Account page's denominator paragraph renders in that order. Every column after wire_control_p402 therefore shifts right by two — as adding wire_control_p402 itself did on 5 September. The file is read by header name by this product, by the exporter and by the deploy guard; a parser that reads it by column position will need updating, and this entry exists so that is found here rather than in a broken script. Nothing is newly measured — the statuses were already in the archive and no published figure changes — so the methodology version does not move. The separate 401 capture that keeps the WWW-Authenticate challenge, and the methodology version it introduces, arrives with the 13 September edition and is its own entry below.
2026-09-13 · methodology 2026-09-13.1
The authentication wall is measured: 401 counted on both wire legs, and the challenge header kept
401 was already arriving and nothing was looking at it. Both wire instruments recorded the status on the row — the 2026-09-02 edition holds seven domains answering 401 to the control leg and five to the named crawler — and no count, column or page said so, while the one header that says what a 401 wants was discarded before the row was written. A 403 and a 401 are different objects: 403 is a refusal, 401 is a door saying not until you identify yourself, and this is the way, and the second half of that sentence arrives in WWW-Authenticate. (1) The wide probe and the identity matrix now count 401 on each leg, separately from 402 and never folded into it: 402 states a price, 401 states a condition of entry, and a reader who adds them has measured nothing. (2) The WWW-Authenticate and Proxy-Authenticate challenge is captured off the response, under its own pattern. It is deliberately NOT added to the existing signal patterns: those fill counts the archive already holds earlier values of, and widening one would have silently moved a published series in a way no downstream guard could tell from a real change in the data. A new guard, check-capture.cjs, holds the two patterns apart from here on, and each prober refuses to run if its own signal pattern grows to match a challenge header. (3) The counts are published in the free per-edition archive from the first edition that carries them, and are written as null — never as zero — for every edition probed before the capture existed, because “we did not measure this” and “we measured it and found none” are different facts and the archive is the one artefact that cannot be rebuilt if it conflates them. Nothing about the request changed: the same identities, the same headers, the same timeout, the same target URL, so the probe version does not move and the series stays comparable across the change. What changed is what is kept from the answer. This entry is dated to the first edition that carries the capture — the Sunday sweep of 13 September — and not to the day the code was written, so that the methodology version it introduces cannot stamp a deploy made before that edition ran. It was first written dated 11 September, which is a Friday and not a census day at all; the date was corrected before either version could be published, and check-capture.cjs now asserts the rule rather than trusting anyone to remember the sweep calendar.
2026-09-10 · product
The identity matrix is named in full, on the page and in the machine-readable methodology
The wire instruments were described asymmetrically. /methodology said only that a 331-domain identity matrix asked “under several identities”, and the string crawler-max-price appeared on no public page at all. Both halves are now written out. (1) /methodology names the five identities the matrix presents, one request per identity per domain per edition: our own CrawlPriceIndexBot/1.0, the honest-identity control; a current Chrome user-agent string, which is a string and not a browser; GPTBot’s published user-agent string; ClaudeBot’s published user-agent string; and ClaudeBot’s published user-agent string carrying the request header crawler-max-price: 0.001. That fifth identity states a price a crawler would be prepared to pay, in the units that header defines, and exists to observe whether a door that refuses a named crawler answers differently when the request carries it. No payment is offered in any binding sense and none has ever been made: there is no wallet, no credential and no code path in this pipeline that could settle one. Two riders are stated rather than waited for — presenting another operator’s user-agent string means the response recorded is the response that string received and not one that operator received, and the requests originate from an ordinary consumer connection rather than any published crawler address range, so a door that verifies its callers by address treats these knocks as unverified, which is part of what the matrix measures. (2) The same facts are now in /v1/methodology’s crawler_identity field. They were not, for a day: the disclosure shipped to the page while the machine-readable methodology still described the census, the reachability sweep and the wide probe and stopped there, so a reader with a browser was told about the matrix and a machine reading the methodology was not. On a site whose premise is that a claim is worth what a machine can check, that was the wrong way round, and a new rule in check-public.cjs now fails the deploy when an instrument named on /methodology is named nowhere in crawler_identity. (3) A paragraph between the two wire sections states the trade they make rather than leaving it to be inferred: the wide probe covers the whole reached frame and presents one named crawler, which is why a wall that fires at ClaudeBot and not at GPTBot is invisible to it; the matrix covers 331 domains under five identities apiece. The two are a pair, not one claim. (4) The honest identity’s contact string was the placeholder “hello@crawlpriceindex TBD”, shipping in real requests to 50,000 domains on the one identity whose whole purpose is to be reachable; it is now hello@crawlpriceindex.com, which is where the operations report already goes. Nothing here is newly measured and no published figure changes — the matrix has run since before this entry and every reading it produced stands — so the methodology version does not move. What changed is that the method now says what it does.
2026-09-05 · product
The control leg’s 402 count ships as a CSV column; a false correction is withdrawn; the corrections file states its own order; four published figures get the corrections they were owed
This entry was amended before publication on 5 September: it was drafted in the morning, while the wire_control_p402 column was still missing and before four figure corrections were written, and it was never deployed in its earlier form. What it said then, and what is true now, are both below. THE DELIVERABLE. (1) The aggregate series CSV now emits wire_control_p402 — the control leg’s own 402 count, per edition: 46, 45, 47, 47 for the four editions to 2026-09-02. The header row carries fourteen wire_* columns. Its denominator is that edition’s wire_reached, and it is not a subset of wire_ai_p402 and is never subtracted from it: a domain can answer 402 to both legs. Two pages were saying the opposite while it shipped — the Wire tab’s footnote told a reader the count “is not yet a column in the aggregate series CSV”, and the Account page’s paragraph defining every wire_ column was a typed list of thirteen beside a file with fourteen. Both are now rendered from the exporter’s own column list, and a guard asserts that list against the real header row, so neither page can state a delivery state it does not read. THE TRUST ARTEFACTS. (2) The 4 September correction announcing that column is WITHDRAWN by a dated entry and carries a dated note in place: the column was not in the CSV when the entry was published, and the export emitted thirteen wire_* columns without it. The count itself was real and already published in the free history file as wide_probe.p402; what was announced was the deliverable, and the deliverable did not exist. The entry text stands as published, per the rule on /status, and a third dated entry now records that the column ships. (3) /corrections.json is an object that states its own ordering rule — { artefact, order, count, entries } — sorted newest first by date and then by a stable id of the form -, because 28 of the entries share one date and a date sort alone leaves a 28-way tie. Every entry carries that id, so a correction can be cited. The free feed’s corrections key is unchanged: still a plain array, now sorted, with corrections_order beside it. (4) /index.json’s wire_csv_columns list is reconciled against the exporter’s own column list before it is published, so the feed states the columns the CSV actually emits, in the order it emits them; it had been naming wire_payment_headers for a column the CSV emits as wire_any_signal_headers, and wire_control_p402 for a column the CSV did not emit at all. check-corrections.cjs resolves every field named in a correction against the built feed and the real CSV header — produced outside the browser from the exporter’s own code — and fails the deploy when a correction names something the artefacts do not carry. That guard first exempted a whole entry from resolution when it was withdrawn or declared it asserted nothing, which is why it read green while four surfaces denied a column that had shipped; the exemption is now a single named field inside an entry, never the entry. FOUR PUBLISHED FIGURES, WITH CORRECTIONS BESIDE THEM. This round changes four figures that were published, so — against the rule printed on /status, that we do not silently amend past editions — each one now has a dated correction. (5) The Watchlist’s reach figure, published on 4 September as “20 of the list’s 100 were reached”, divided by a 121-row exhibit set instead of by the list; it reads 82 of the 83 the GPTBot leg was run against, of 100 asked — the control leg was run against all 100 and reached 83, and every rate on the panel divides by the base printed beside it. In the same pass the big table’s Wire column had titled 80 of its 100 rows “not reached by this edition’s probe” for domains the probe reached; all 100 now carry a status, and the foot names the leg order it is read in. (6) For one day, two paid bar panels drew a 393px axis strip over a 177px bar track, offset 186px: a bar printed as 25.8% appeared to sit at about 32%, and every bar on those panels read roughly 2.2 times its printed value. The printed numbers were right; the ruler was not. Every tick strip is now written from the measured track after each render and resize, and a new guard measures all 28 axis sites at 1440, 1280, 800 and 640 and fails the deploy on more than 1px at either edge. (7) Full detail’s unread-reasons panel asserted that two counts from two different files “both sum to” one total — 18,500 is the key count of the 2026-08-23 per-domain file, 18,512 is the aggregate’s tail.tail_total — and offered “nine published buckets” over five drawn groups. Neither sentence renders; every number names the file and edition it is read from. AND THE EXPORTS SAY WHAT THEY CANNOT SAY. (8) The new per-panel CSV exports were shipped on the claim that every file leads with its finding, denominator, scale and n. Nineteen of the thirty-three could not fill three of those lines and printed “Denominator: not stated on this panel”, which reads as though the panel has none. Those files now write only the lines they can fill and name the ones they cannot, and why.
2026-09-04 · product
Every panel states its scale, its denominator and its finding — and a geometry guard that can see what text guards cannot
This round is about how the numbers are drawn rather than what they say: no published figure changes. (1) Every bar panel now states its scale — an axis with ticks, or the sentence saying the bars are scaled to the widest bar in the panel and not to 100%. Three bar conventions ran side by side here and none was labelled, which is how a max-scaled bar gets read as a share of the whole. A guard now fails the build on any panel that states neither. (2) The by-edition charts default to Level with the axis anchored at zero, and draw one unconnected dot per published edition. With four editions a joined line claims a direction the data cannot support; the line returns automatically when the series reaches eight. The Explore chart drew three points under a heading promising every edition, because it plots the change between editions — four editions give three changes. It now says so. (3) Panels are titled with their finding and its number instead of a category label, and every tab opens with a finding rather than a control. (4) Two 18×18 matrices are gone: the co-treatment square, where every value sat between 77% and 100% and the picture read as ‘all crawlers are treated alike’, a claim nobody made; and the selective-exclusion square, now a ranked list of the twenty pairs it existed to show. The Wire refusal counts are drawn as a ladder whose branches do not pretend to nest, the crawler map's two inverted axes are stated the way round a reader reads them, and the reachability ladder is split so that no bar divides by a base other than its own panel's. (5) New per-domain pictures for subscribers: the wire wall (one row per domain, one column per edition, the honest control above the AI identity), a dumbbell of a saved list against the rest of the frame, and a cohort series with the 15 September Cloudflare default drawn on it. (6) A geometry guard now renders every public page and every dashboard tab headless before a deploy and fails on overlapping labels, text outside its frame, and any chart whose marks disagree with the edition count its own caption prints. Two defects on this site were found by a reader and not by six rounds of review, because those reviews read pages as text: two labels that occupy the same pixels are two correct strings in the markup, and a three-point chart is three correct numbers. That is the gap this guard closes.
2026-09-04 · product
corrections.json published; the hash says what it certifies; one start date per instrument
Four things the site said that were not checkable are now checkable. (1) The machine-readable corrections array is published at /corrections.json, carried in the free feed at /index.json under `corrections`, and named by URL in /v1/methodology, on /methodology and on /status. It was promised in six places and answered 404. (2) The per-edition SHA-256 table on /status now hashes each private file as it ships — the app/private copy the API streams back — prints the route beside every row, and says what a reader can compare it against: the bytes a route returns, hashed as they arrive. It previously read 'it certifies the file you downloaded', which was not true of the dashboard's download buttons or of any CSV: those rebuild the file in the browser from rows already parsed there, so they carry the same data as different bytes. The table also prints every edition this machine has hashed rather than promising five and printing one. (3) /status gains one line per instrument with its own first edition — robots census, the wire probe's honest control leg, its named-crawler leg, the browser-UA leg that has not run, exhibit tracking, the registry's per-domain capture — because five surfaces were answering 'when did this begin' with five different dates for five different instruments. (4) One word per quantity on the wide probe: asked (the domains the probe requested) and reached (those in the frame that answered). The feed keys refusal_headline.probed and probe_panel_observations.probed keep their names for continuity and every sentence beside them now says which of the two it is. Also: /panel.json's selection string is rebuilt from the probe's own output, so it no longer says the wide probe 'never impersonates' — a sentence written before the named-crawler leg existed; /status describes /world as it is (three groups drawn for a free reader, the rest counted); and the 2026-09-02 entry below carries the dated note the 4 September correction owed it. Nothing is newly measured and no published figure changes, so the methodology version does not move.
2026-09-04 · product
Correction: the wide probe's control leg is our own identified crawler, not a browser
Since edition 2026-09-02 the wide probe has sent two homepage requests per target: one presenting GPTBot's published user-agent string and one as CrawlPriceIndexBot/1.0, our own self-identified crawler (the honest-identity control; unsigned on this instrument — the census fetcher is the signed one). Until today the home page, Explore, the methodology page, the Terminal and the 2026-09-02 entry below described that control as an ordinary browser. It was not one. Every published count that reads 'refused the crawler while serving …' was always computed against the identified-crawler request (compute-dashboard's basis and caveat_control fields have said 'honest-identity request as the control' throughout), so no figure changes; the wording does, everywhere the leg is named. The methodology version does not move: nothing measured changed. A third request per domain carrying a current Chrome user-agent string — not a real browser: no JavaScript, no browser TLS fingerprint — is planned from the 6 September edition and will be entered here, with a version, once an edition carrying it has published. The census identity is unchanged. Logged in corrections.json.
2026-09-04 · product
Watchlist & cohorts on Terminal; the aggregate series as one CSV; API cap 5,000
A new Terminal tab, Watchlist, takes a pasted list of domains or a cohort defined by rank band, ads.txt, suffix, declared state and changed-since-baseline, and shows its per-crawler block rate beside the rest of the frame (not the whole frame), with movement against the header baseline, the wire pair for every probed row, unread reasons, and two CSV exports (current edition, and every edition in long format). Lists are saved in the browser only. Nothing is newly measured. Account & data gains an aggregate series CSV, one row per published edition, read from the archived snapshots. The API cap rises from 1,000 to 5,000 requests a month per key. Change alerts stay at three confirmed domains per address.
Note, 2026-09-04: Correction, same day: the Watchlist has three CSV exports, not two. The third is the cohort’s wire history — the per-edition pair of answers (our own self-identified crawler, and the named AI crawler where that leg ran) for every probed row in the cohort, one row per domain per edition, with a leg that did not run left empty and never written as a zero. It shipped with the tab and this entry did not name it.
2026-09-04 · methodology 2026-09-04.1
One denominator per tab; the feed describes the wide probe; two issues withdrawn
The selective-exclusion matrix on the Crawlers tab is now a share of the domains with a readable robots.txt (31,488), like every other rate on that tab; it had been a share of the 50,000-domain frame. The public feed's probe_panel_observations now describes the wide probe (one presented identity, purposive set) instead of the retired 650-request panel, and carries no inferred price bands. The identity-matrix panel size, the wide probe's reachable count and the edition count are printed from one source each. Two issues of The Weekly Crawl (23 and 24 August) are withdrawn from the archive: one was computed across the frame change, the other on the 24 August validation scan; the emails stand and the withdrawal is logged on status. /estimate is retired: the site measures and does not model a market.
2026-09-03 · product
Shelf simplified: Terminal only
Withdrawn: the €29 one-off snapshot (Terminal is monthly, cancel any time; snapshot links bought before today keep their 48-hour window) and the machine tier (one edition for USD 20 in USDC on Base), which was described on this page, in the privacy notice and in the security page but never sold. Retired: the token-gated free sample page. The public site and the dashboard now use one direction rule for policy changes (into an explicit block, out of one, or reclassified), one edition count, one rate floor for suffix groups (n ≥ 100), and label unread robots.txt by what happened (refused, no file, unreadable, unreachable) instead of calling it absent.
2026-09-02 · methodology 2026-09-02.1
Wide probe presents a named crawler
The wide probe now sends two homepage requests per target in one sweep: one presenting GPTBot's published user-agent string and one as an ordinary browser, and records both answers. This replaces the single identified identity used since 2026-08-18 and makes identity-conditional refusal observable for the one identity presented. It is a disclosed measurement study: one request per identity per domain per edition, no content retained, never retried to get past a refusal. The Wire evidence tab, the Briefing's CPI-11 and CPI-12 and the home page read from this probe. The 2026-08-18 entry's 'never impersonating' described the earlier instrument and no longer applies to the wide probe.
Note, 2026-09-04: The second leg described above as “one as an ordinary browser” is not a browser and never was. It is our own self-identified crawler, CrawlPriceIndexBot/1.0, sent unsigned on this instrument — the honest-identity control. Every count computed from the pair was always computed against that identified-crawler request, so no published figure changes; the description does. The full correction is the 4 September entry above, and it is recorded on /status and in /corrections.json.
2026-08-23 · methodology 2026-08-23.2
Every issue gets a public permalink
The Weekly Crawl is now published to /issues/ as well as to subscribers, rendered from the same content model as the email so the two cannot describe different weeks. The issue leads with the named domains that changed AI-crawler policy that week, ranked by the shape of the change rather than by Tranco position: an edit that blocks some crawler roles while opening others is a decision, a blanket switch across every crawler is a setting. The five to seven domains named on the page are the same ones the free email names.
Note, 2026-09-04: Two issues (23 and 24 August) were withdrawn from the archive on 4 September 2026, with the reasons recorded on /status and in the archive index. The rule as it stands: every issue gets a public permalink, or a dated withdrawal note where the permalink was.
2026-08-23 · methodology 2026-08-23.1
Edition published on the CPI-50K v1 frame; corrections issued
31,469 of 50,000 domains served a readable robots.txt. Six corrections are recorded at /status and in the machine-readable corrections array of the dataset, covering the public feed lag, three newsletter defects, a freshness block counted over the wrong population, and two forbidden framings found when the public copy guard was pointed at all seventeen pages instead of six. The adaptive crawl throttle was also recalibrated: its trigger was hardcoded at 8% of the domains in each 2,500-domain chunk returning 403 or 429, against a measured floor of 7.7% of the 50,000-domain frame that refuses an identified crawler however gently it knocks. It therefore fired on a property of the web rather than of our behaviour, and it could only ratchet downward. It now measures the floor per run and can recover.
2026-08-21 · methodology 2026-08-21.1
CPI-50K v1 frame adopted; published series restarts
The sampling frame is now a custom Tranco list built from Cisco Umbrella and Majestic only. Cloudflare Radar (CC BY-NC) and the Chrome UX Report (CC BY-SA) are deliberately excluded, because a non-commercial licence restricts use and not merely redistribution. The published series therefore restarts at the 2026-08-23 edition; earlier editions remain on disk but are excluded from the series, the charts and the dataset. On the 17,179 domains readable in both frames the measurement did not move: any-block 20.50% before against 20.57% after, and a training-versus-search asymmetry of 18.9 to 1 in both.
2026-08-18 · methodology 2026-08-18.1
Wide honest probe + three-tier panel
Added a wide, honest-identity probe covering the top-ranked domains plus every domain observed blocking a crawler — one signed request each, never impersonating — recording payment headers, 402 walls, Cloudflare fronting, X-Robots-Tag and llms.txt. The identity-matrix panel (which does probe under other crawlers' names, and is deliberately capped) is now assembled from three tiers: a fixed continuity spine, domains auto-promoted from the honest probe when they show payment or blocking behaviour, and a rotating random audit that samples the full scan so AI-only walls invisible to honest probing are still measured. Panel composition is published at /panel.json. [Superseded on 2026-09-02: the wide probe now presents GPTBot's user-agent as one of two identities — see the 2026-09-02 entry.]
2026-08-16 · methodology 2026-08-16.2
Observed-price series corrected
Fixed a parsing error that stored a null price for the first two editions; the observed top price is now recorded as a continuous series from the 8 Aug baseline, with the early points backfilled from the value observed at the time and marked as such. The price shown has been a flat $0.50 (stackoverflow.com, via Cloudflare pay-per-crawl) at every observation to date.
2026-08-16 · methodology 2026-08-16.1
www fallback + reachability census (coverage restated)
The scanner now retries on www. when a bare domain fails at the network level, recovering roughly a thousand sites that only answer on their www host. Because this changes what counts as 'reached', coverage steps up between the 15 Aug and 16 Aug editions for methodological reasons, not because the web changed. Every unreachable domain is now itemised by reason (no DNS, refused, timeout, broken TLS, HTTP error) in a separate census, so coverage is reported against the reachable web rather than the raw 50,000. A late-sweep retry pass, capped in time, recovers transient failures.
2026-08-11 · methodology 2026-08-11.1
Provenance model, public methodology, change alerts
Every published signal now carries an explicit evidence type — observed, derived or inferred — plus a pinned methodology version, an explicit list of what we do not measure, and a corrections policy. Added a machine-readable methodology endpoint at /v1/methodology and a full human methodology page. Added per-domain change alerts: any site owner can be told when a tracked AI crawler's access to their domain changes.
2026-08-10
Free per-domain checker, CSV feed, machine payment tier
Launched /check — a free tool showing any domain's AI-crawler policy fingerprint, its percentile against every parsed domain, and ready-to-paste robots.txt and RSL files. Added a CSV flavour of the paid dataset. Added a machine tier: one weekly edition for USD 20 in USDC on Base, verified on-chain, with history withheld from machine buyers.
2026-08-09
First full 50,000-domain sweep published
27,067 domains parsed across the Tranco top 50,000, 53 ccTLD editions, and the first enforcement measurement comparing declared blocks against what an identified crawler actually receives. Weekly history begins accumulating from this edition.
2026-08-09
Crawler identity signed under Web Bot Auth
Our crawler now signs every request using HTTP Message Signatures (RFC 9421) under the Web Bot Auth profile, and publishes its public key directory, so any site can verify that traffic claiming to be ours really is. Submitted to Cloudflare's verified bots programme.