{
 "artefact": "Crawl Price Index — corrections",
 "order": "Newest first: date descending, then id descending. id is <date>-<nn>, nn being publication order within that date — 28 of these entries share 2026-09-04, so a date sort alone leaves a 28-way tie and no defined order. Entries are never rewritten: an entry that was wrong when published stands as published, carries a dated note, and is withdrawn by a later dated entry.",
 "count": 47,
 "entries": [
  {
   "id": "2026-09-10-04",
   "date": "2026-09-10",
   "what": "A published percentage on the Terminal's Wire tab moves: the browser-UA leg's CHALLENGED count is corrected from 23.8% to 22.8% of refusals, on the 2026-09-09 edition 768 rows to 735. The row belongs to a cut that opens “the same 3,223, cut by what the browser-UA leg got”, so its denominator is the refusals — but compute-dashboard.cjs incremented the counter outside the refusal branch, counting a challenge on every in-frame row while the page rendered it over the refusals. The count is now taken inside that branch, so the row counts the set its own label names. The three drawn rows plus a fourth the page did not draw partition the refusals exactly — 1,944 served, 735 challenged, 529 refused too, 15 with no browser status, summing to 3,223 — and that fourth group is now a row of its own with the arithmetic printed beside it, so the cut visibly sums to the number in its heading. An earlier internal review proposed re-basing this row on the reached set instead (768 of 7,496 = 10.2%). That repair is not taken: it would remove the row from the set it is drawn inside, and the cut would then sum to nothing. The label was correct; the numerator was not. No other figure on the panel changes, and the browser-UA leg's served and refused-too counts were already taken inside the refusal branch and are unaffected."
  },
  {
   "id": "2026-09-10-03",
   "date": "2026-09-10",
   "what": "The identity matrix's five identities are now named on /methodology, and one of them is named for the first time anywhere on this site: ClaudeBot's published user-agent string carrying the header crawler-max-price: 0.001. The page previously said the matrix was 'asked under several identities' and named none of them; the header appeared on no public page. The crawl-etiquette section did disclose that named crawlers' user-agent strings are presented in the matrix, so what was missing was which identities, and that one of them states a price.",
   "why": "Presenting a named crawler's user-agent is an act this site had disclosed. Sending a header that states a price a crawler would pay is a different act, and a reader cannot evaluate the max-price finding without knowing it was made. The measurement is worth keeping - it is how the 403-becomes-200 flip was observed - so it is disclosed rather than withdrawn. Also corrected in the same commit: the honest crawler's user-agent sent a placeholder contact string ending 'TBD', while this page published the working address. The page was right and the wire was wrong; the wire now sends the address /methodology has always shown.",
   "asserts_note": "The matrix is older than this disclosure. The earliest record held of the max-price identity running is the 25 August 2026 review, which quotes a flip observed that edition."
  },
  {
   "id": "2026-09-10-02",
   "date": "2026-09-10",
   "what": "Two rungs of destination A on /futures were registered inconsistent with themselves and are now marked so on the page. A7 (“The browser is the stranger”) reads dashboard.wire_join.control_status_mix.200 — the identified-crawler control — while its registered name, condition and denominator population all said “the browser control”. Its population read “reached wire domains (the browser control answered 200)”, which is false about the field: the count is the identified crawler's, and the browser leg is a separate reading. The population is corrected in place to “reached wire domains”. The rung's name and condition are BYTE-IDENTICAL to the registered text and stay that way: correcting them would change what the destination claims while leaving the goalposts hash untouched, which is precisely the conflation the freeze exists to stop. Instead both A7 and A8 carry a dated amendments[] entry, printed on the rung's own row, saying that which control the rung means is unresolved and that the rung is unscorable under either reading until it is re-registered. A rung with an open amendment is not scored even where its fields resolve. Two unhashed fields on one rung, one dated entry; no goalpost moves, and the register's sha256 is unchanged.",
   "why": "A7's field is emitted by nothing, so the rung is unscorable today whichever control is meant and nothing on any page depended on the answer. That will not stay true: on this edition either reading already clears the rung's 10% bound (the identified leg at 20.9% or more, the browser leg at 74%), so the day the field ships, a rung whose registered sentence names one control and whose expression reads the other would light — and the page would print a reached rung on evidence its own words do not describe. Re-pointing the expression is a goalpost move and needs a retirement and a re-registration, not an amendment; naming the control the expression actually reads would make the frozen sentence mean something new with the hash unchanged. Saying, dated, that the register is inconsistent with itself and scoring nothing until it is fixed is the only one of the three that claims less than before.",
   "asserts_absent": [
    "dashboard.wire_join.control_status_mix"
   ],
   "asserts_note": "The field A7 reads is emitted by no census script and appears in no published feed; the assertion is that it is absent, which is why the rung is unscorable."
  },
  {
   "id": "2026-09-10-01",
   "date": "2026-09-10",
   "what": "/futures scores a destination's kill condition over a stated window, and the window is now written down: the FIRST n editions published on or after the collapse date, where n is the rule's registered `consecutive` (one for The Toll Web and The Great Lockout, three for The End of the Internet of Humans). Until today the window was one edition wide whatever n said, so destination A's three-reading rule could not fire at all — it was registered with a condition its own scorer could never satisfy. The obvious repair, letting the run be found on any edition from the date until the destination arrives, was rejected: it leaves A right and silently widens the other two from one edition to roughly three hundred, and widening when a kill condition may fire claims MORE than the register does. Destinations D and C are scored exactly as they were before this entry. Where an edition inside the window carries no reading for the rule, the window is not treated as false: the card says the rule is not determinable at its date. Note that `consecutive` now carries two meanings on this register, and both are honoured as registered: at a checkpoint it means n editions ENDING at the checkpoint, and in a collapse rule it means the first n editions STARTING at the collapse date.",
   "why": "Every collapse text on this register begins “at 18 Aug 2027”, so the evidence has to be at that date and not merely somewhere in the archive. An interpretation of a registered word may narrow what the page can claim and may not widen it; of the two readings that fix destination A, only this one leaves the other two destinations where the register put them. The page prints the window's state per destination — before the date, open with a partial run, collapsed on a named edition, or closed — and the guard recomputes it from the same function the page uses."
  },
  {
   "id": "2026-09-05-08",
   "date": "2026-09-05",
   "what": "For roughly five hours on 5 September 2026 — from about 06:03 UTC until the correction below was deployed — the public feed at /index.json and the data payload behind the Explore page served figures from a run that is not a census edition and has been withdrawn. It reported 29,672 domains with a readable robots.txt, and block rates about a third of their true level: GPTBot at 4.6% where the 2 September census measured 15.9%, CCBot at 4.6% against 15.5%, and every one of the eighteen tracked crawlers moved the same way and by a similar amount. Anything read off the public feed or the Explore charts in that window is wrong and should be re-read. The correct current figures are those of the 2 September 2026 edition: 31,488 domains with a readable robots.txt, GPTBot 15.9%. The run was not a census. A daily job left over from the pre-Raspberry-Pi setup had been advancing one sweep in slices on a laptop since 11 August; when the final slice completed it treated the sweep as finished and ran the whole publish chain. What it published is therefore a measurement smeared across twenty-five days at a pace and reachability the census does not use, not a snapshot of one night, and 5 September is in any case a Saturday, while the census runs Sunday and Wednesday. Three things were NOT affected, and the distinction matters. The paid Terminal was never touched: it still reads the 2 September edition, because that job does not deploy the app. No newsletter was sent — the newsletter guard refused, because the issue page the email links to did not exist. No change alert was sent, and none could have been: there were no active watches. The withdrawn edition has been removed from the published series rather than filtered out of it, so no chart, delta or comparison baseline can reach it; its raw files are retained as evidence outside the tree and in git history.",
   "why": "Every guard this project runs compares one published surface against another, and a sweep that ran to completion at the wrong pace is internally consistent — uniformly wrong, and so invisible to all of them. The denominators agreed with the rates, the rates agreed with the charts, the charts agreed with the copy. Nothing asked whether the measurement had happened the way the method says it happens. That guard now exists and runs before anything is published, sent or committed: it compares an edition against the one before it on the coverage denominator, on whether the record carries every section its predecessor carried, on whether every crawler with a published rate also has a state breakdown, and on whether the whole board of crawler rates moved one way at once. Three of its rules fail on this run, and two of those could not be argued away — a missing measurement leg, and a state map holding one key where the previous edition holds eighteen. It can be overridden, but only with a written reason that is recorded. The job that published this has been deleted rather than disabled. It had been disabled once already, and a disabled job that still exists on disk is a job that runs again. One figure in the internal account of this incident was also overstated before it reached anyone: an early reading held that the truncated scan would have mailed tens of thousands of false alerts. It would not have. The alert job only examines watched domains and skips any domain missing from either side of its comparison, so a short scan under-reports rather than over-reports. The real cost was a corrupted comparison baseline, and it is recorded here at its true size rather than its first-draft one."
  },
  {
   "id": "2026-09-05-07",
   "date": "2026-09-05",
   "what": "The machine-payment registry count was published against the wrong base on two surfaces and has been changed. The free Machine-payments gate said the in-frame domains were a share of the 50,000-domain frame, and the paid how-to-read line said “of the 50,000 domains CPI scans”. The registry captures are joined against the frame as it stood at capture time, which is 49,156 domains, and that is now the base printed on every surface — the free gate, the Terminal and /explore, which had each divided by a different number. The count of in-frame domains did not change; only the base it is quoted against.",
   "why": "One instrument, one denominator. The buyer panel named this figure as rendered three ways in three consecutive reviews, and it was the crack in a page whose whole argument is that denominators are stated. The base was corrected in the same round that wrote three other corrections for changed figures, and it needed its own — a figure quietly replaced on a free page is exactly what this file exists to refuse, and it was caught by the verifier rather than by us."
  },
  {
   "id": "2026-09-05-06",
   "date": "2026-09-05",
   "what": "Two claims are withdrawn from Full detail's unread-reasons panel. Its legend asserted that two counts 'both sum to' the same total when they come from different files: 18,500 is the key count of the 2026-08-23 per-domain file and 18,512 is the aggregate's own tail.tail_total for this edition. They are different quantities, so the sentence was false whichever number a reader carried away. Above it, a caption offered 'nine published buckets' over five drawn groups. Neither sentence renders now: every number on the panel names the file and the edition it is read from, and the group count is the number of groups actually drawn.",
   "why": "A cross-file identity asserted in a caption is the hardest error to catch, because both numbers are correct and only the equals sign between them is wrong — a vocabulary fix would have left it in place. The rule adopted this round is that a sentence names which file each number comes from, or it does not make the claim."
  },
  {
   "id": "2026-09-05-05",
   "date": "2026-09-05",
   "what": "For one day — the edition published on 4 September — two paid bar panels drew an axis that did not rule their own bars. On Full detail's rank-band panel and on Policy layer's rank-frame panel, the tick strip ran from 212px to 605px across a bar track that ran from 398px to 575px: a 393px ruler over a 177px track, offset 186px to the left. Read off that ruler, a bar printed as 25.8% appeared to sit at about 31%. The misreading was not a constant multiple: the ruler was both shifted and stretched, so the error is an offset plus a compression — read ≈ 18.9 + 0.45 × printed. Small bars were exaggerated most and a bar at the top of the scale read slightly LOW; across the seven bars actually on the panel the overstatement ran from 1.08× to 1.44×, and a bar at the axis maximum read 0.92×. An earlier draft of this entry said “roughly 2.2 times”, which was wrong — 2.2 is the ratio of the ruler’s length to the track’s, not the error in any reading. That draft was corrected before publication and never deployed. The number beside each bar was right the whole time; the ruler above them was not. A third panel's strip was 60px out on its right edge. Every tick strip is now written from the measured bar track after each render and each resize, and a new guard measures all 28 axis sites at 1440, 1280, 800 and 640 and fails the deploy when either edge is more than 1px out.",
   "why": "A figure a reader takes off a chart is a published figure, even when the chart's printed labels are correct — and this was the flagship panel of the round that introduced axes, so reading values off the ruler is exactly what it invited. Nothing in the suite could see it: the scale guard asks only whether a scale is stated, and the overlap guard only whether text collides. Both reported green for a day. Cause: the tick strip laid its columns on pixel constants passed by the caller while the row beneath it took its columns from CSS, and four media queries change that CSS again. Two grids that were supposed to be one."
  },
  {
   "id": "2026-09-05-04",
   "date": "2026-09-05",
   "what": "The Watchlist's reach figure, published on 4 September as '20 of the list's 100 were reached', divided by the wrong set. It was computed over a 121-row exhibit set — the rows that produced a payment exhibit — rather than over the list, so a reader saw a 20% reach rate for a cohort whose doors were almost all answered. It now reads 82 of the 83 the GPTBot leg was run against, out of 100 asked — the control leg was run against all 100 and reached 83, and the GPTBot leg was run against those 83 — and every rate on the panel divides by the base printed beside it: 22 refused (26.8% of reached), 11 refused the crawler while serving the control (13.4%), 3 answered 402 (3.7%), 60 served (73.2%), 16 Cloudflare-fronted (19.5%). In the same pass, the big table's Wire column had titled 80 of its 100 rows 'not reached by this edition's probe' for domains the probe did reach. All 100 rows now carry a wire status and the table's foot names the leg order that status is read in.",
   "why": "The exhibit set is not the reached set, and dividing by it turns a reach rate into a share of a file. The panel named its denominator honestly and it was still the wrong one: labelling a false denominator does not rescue the figure, which is the ruling the buyer panel made unanimously this round. This is a denominator on a paid surface, so it is a published figure and it gets a correction rather than a quiet rebuild."
  },
  {
   "id": "2026-09-05-03",
   "date": "2026-09-05",
   "what": "The aggregate series CSV now emits the wire_control_p402 column. The 4 September entry that announced it was withdrawn this morning because the column did not exist on the day that entry was published; this entry records that it exists. The export's header row carries fourteen wire_* columns with control_p402 among them, and the four published editions read 46, 45, 47 and 47. Its denominator is that edition's wire_reached — the probe targets that returned an HTTP status on the control leg — and it is not a subset of wire_ai_p402 and is never subtracted from it, because a domain can answer 402 to both legs. Two pages had been saying the opposite. The Wire tab's paired-identity footnote told a reader the count 'is not yet a column in the aggregate series CSV', and the Account page's paragraph defining every wire_ column was a hard-typed list of thirteen beside a file with fourteen. Both are now rendered from the exporter's own column list, so neither page can state a delivery state it does not read.",
   "why": "A withdrawal with no deliverable behind it leaves a subscriber with two corrections and no column. What let the falsehood live a day is that the guard written for exactly this class, check-corrections.cjs, exempted an entire entry from field resolution when it carried withdrawn_by or asserts: false — and both 5 September entries carried one, so their present-tense claims about a subscriber's file were never resolved against the file. The exemption is now a named field inside an entry (asserts_absent, with its reason), never the entry; every other claim in every entry, withdrawn ones included, is resolved. The Account page's column table is asserted against the real header row on every deploy, so a column can no longer reach the download and stay undefined on the page."
  },
  {
   "id": "2026-09-05-02",
   "date": "2026-09-05",
   "what": "corrections.json was published on 4 September as a bare array in no order — entry [0] dated 2026-08-23, the newest entry at index 13. It is now an object that states its own ordering rule: { artefact, order, count, entries }, sorted newest first by date and then by a stable id (<date>-<nn>, publication order within that date). Every entry carries that id. The free feed’s corrections key is unchanged — still a plain array, now sorted, with corrections_order beside it. (Corrected before publication, 5 September: the count and the entry list this envelope states are the ones in the file as it ships, which now carries the 5 September entries added after this one was drafted. It was never deployed in its earlier form.)",
   "why": "28 of the 35 entries share one date, so sorting by date alone leaves a 28-way tie and no defined order; a machine-readable record whose order carries no meaning will be read as chronological by whoever consumes it. The ids also make a single correction citable, which is what the dated note under the 4 September entry needed in order to attach to one entry and not to a date.",
   "asserts": false,
   "asserts_note": "Describes the shape of this file itself, not a payload field."
  },
  {
   "id": "2026-09-05-01",
   "date": "2026-09-05",
   "what": "The 4 September entry below announcing the control leg’s 402 count as “the wire_control_p402 column of the aggregate series CSV” is WITHDRAWN. That column was not in the CSV on the day the entry was published: the exporter emitted thirteen wire_* columns and control_p402 was not among them. The same day, /index.json’s wire_csv_columns list advertised the column to every reader of the free feed, and separately named wire_payment_headers for a column the CSV has emitted as wire_any_signal_headers since 4 September — two published statements about a subscriber file that the file does not support. (Corrected before publication, 5 September: this entry was drafted while the column was still missing and said in the present tense that the export emits thirteen columns. It was never deployed in that form. The column shipped the same day — see 2026-09-05-03 — and the clause is now in the past tense, where it belongs. The 4 September entry it withdraws is untouched: its date, what and why are byte-identical to the published file.)",
   "why": "Corrections are the artefact this product asks to be trusted on, so a false one costs more than the error it describes. What was true: the value existed and was published, in the free history file, as wide_probe.p402 — 46, 45, 47, 47 for the editions of 23, 26, 30 August and 2 September. What was not true: that it had reached the CSV. The feed’s column list is now reconciled against the exporter’s own column list before it is published, so the feed cannot advertise a column the export does not emit, and check-corrections.cjs resolves every field named in a correction against the built artefacts and fails the deploy when one does not exist. That guard, run against this entry on the day it was published, would have refused it.",
   "asserts": false,
   "asserts_note": "This entry describes the state of the CSV and the feed on 4 and 5 September 2026. Its field names are a record of what was and was not there on those days, not a claim about the current build. wire_payment_headers is named in asserts_absent because it is named here precisely BECAUSE the file does not carry it: the CSV emits that column as wire_any_signal_headers. Every other field this entry names is resolved against the build like any other.",
   "asserts_absent": [
    "wire_payment_headers"
   ]
  },
  {
   "id": "2026-09-04-28",
   "date": "2026-09-04",
   "what": "The Wire tab's paired-identity panel printed three different denominators for one leg — a title dividing by 7,428, a column labelled n=7,491 and a caption naming 8,002 — and its title used 'refused' to mean 4xx or 5xx EXCLUDING 402, while the refusal ladder three panels above used the same word including them. The published rates were 39.1% for the AI leg and 21.3% for the control. The title is now computed from the same per-column tally the n labels are drawn from, and folds 402 into the refusal as the ladder and the Watchlist do: 43.1% of the 7,491 doors the AI leg was run against, against 20.5% of the 8,002 the control was run against. The underlying counts are unchanged; what changed is which base each rate divides by and what the word means.",
   "why": "One word with two definitions on one tab is a figure that cannot be quoted safely, and a caption that points the reader at the number printed under the column while the title divides by a different one is worse than no caption. The panel now has one definition and one base per column, stated in the frame. In the same pass the panel stopped closing with a note about why POINTS are drawn unconnected: it is a stacked-column chart and there were never any points."
  },
  {
   "id": "2026-09-04-27",
   "date": "2026-09-04",
   "what": "The Wire refusal ladder closed with 'Seven numbers, each a share of the rung above it'. It draws six rungs on a payload without wire_402 and eight with one, never seven; and 1,643, 1,577 and 321 are all shares of 3,227 rather than of each other. The footer now prints the number of rungs actually drawn and says that the first three are steps while the rest are two separate cuts of the same 3,227, whose rows do not sum. The two cuts are drawn in their own bracketed blocks with different coloured rails, so the picture no longer shows three parallel children of one number. Where the payload carries no wire_402.control the panel now says that the split is not published rather than asserting the crossing with no evidence. No count changed.",
   "why": "The count was typed into a string while the rungs were built from the payload, so it was wrong on every edition. The clause contradicted the panel's own branch-B header two lines above it, and the indentation agreed with the clause rather than the header — a caption losing to a rail."
  },
  {
   "id": "2026-09-04-26",
   "date": "2026-09-04",
   "what": "Two panels published titles asserting a count above a body that showed none: the Briefing's 'Named examples' printed '102 domains moved 719 crawler states' over 'No named changes in this interval', and Policy changes' feed printed 'every one of the 719 changes, by domain and crawler' over 'No per-domain changes in this interval'. Both titles are now written by whichever loader filled the body, so neither can claim more than the panel shows, and both empty states print the reconciliation: the aggregate figure, the per-domain file's own edition stamp, and the fact that its newest pair carries no diff for the interval. The aggregate counts are unchanged and the panels that draw them are unaffected.",
   "why": "The gap between the aggregate change count and the per-domain file is an open product defect under investigation. Putting the aggregate figure in a headline the panel beneath it cannot substantiate turns an unresolved reconciliation into a promise, and leaves a subscriber with a number and nothing behind it. An empty leg is empty and says why; it is never a zero and never a claim."
  },
  {
   "id": "2026-09-04-25",
   "date": "2026-09-04",
   "what": "The Wire evidence tab's declared-versus-observed matrix said, of the lower-left cell, that 'a domain with no readable robots.txt is not in this matrix at all', two sentences after naming six such domains inside that cell. The exclusion was never made: the four cells are drawn from buckets that route every declared state other than a full disallow into the bottom row, no-readable-robots included. The panel now says which row those domains fold into and prints how many, split by the cell each lands in, and states that the Watchlist applies the opposite convention to its own cohort rates. No cell count changed. In the same pass the panel stopped apologising for a count it was reading under a key the builder does not emit (wire_join.excluded_no_readable_robots); the key is wire_join.declared, and the builder's own parts now sum to the cells they describe rather than to a base 105 rows larger.",
   "why": "A matrix that names six domains in a cell and then says that class of domain is not in it is not a caveat, it is a contradiction, and a reader who trusts either sentence is misled by the other. Keeping the rows and naming the fold is the reading the data supports; excluding them would have changed published cell counts to match a sentence rather than the other way round."
  },
  {
   "id": "2026-09-04-24",
   "date": "2026-09-04",
   "what": "A second place on the Crawlers tab was still dividing selective exclusion by the 50,000-domain frame after the matrix itself was fixed: the crawler drill-down's 'Selective exclusion' card, which read the precomputed exclusion_matrix percentages. It printed 'most often against ChatGPT-User (0.37% of the parsed set)' one click away from a list printing the same pair as 186 domains, 0.59% of 31,488 — two figures for one pair, both labelled 'of the parsed set'. Both surfaces now read one pair of helpers over exclusion_matrix_counts and exclusion_matrix_denominator, and the drill-down prints the domain count beside the share. No count changed; one of the two shares did, to the one the panel's own denominator produces.",
   "why": "Two computing sites for one figure is the arrangement that made the first error possible, and fixing only the site that was noticed would have left the two agreeing by coincidence after the next census rather than by construction. There is now one site."
  },
  {
   "id": "2026-09-04-23",
   "date": "2026-09-04",
   "what": "The correction note above, as first published on 4 September, said the GPTBot-blocked / OAI-SearchBot-allowed cell moved from 0.30% to 0.48%. That pair is 149 domains and moves to 0.47%; 0.48% is the GPTBot / ChatGPT-User cell, which is 151. The note now names the pair each figure belongs to and quotes the domain count beside it, from exclusion_matrix_counts. No published cell moved: the error was in the correction, not in the matrix.",
   "why": "0.48 is 0.30 x 1.588 rounded — the figure obtained by multiplying a rounded percentage back out, which is exactly the practice that note's own 'why' was written to stop. A correction reconstructed the same way as the defect it corrects is not a correction. Every figure in it is now read from a published count over a published denominator, and each is attached to the pair it describes."
  },
  {
   "id": "2026-09-04-22",
   "date": "2026-09-04",
   "what": "The 4 September changelog entry announcing the Watchlist described \"two CSV exports\" when the tab shipped three; the third, the cohort's per-edition wire history, was not mentioned at all. The entry text stands as published, with a dated note beneath it naming the third export.",
   "why": "A release note that undercounts its own release understates what was shipped and makes the tab's own copy (which says the same) harder to correct. Published entries are not rewritten here, so the count is corrected in a dated note attached to that entry."
  },
  {
   "id": "2026-09-04-21",
   "date": "2026-09-04",
   "what": "/panel.json's `selection` string — linked from /methodology as the panel's full composition — said \"the honest wide probe covers thousands of domains and never impersonates\". That described the instrument as it ran before the 2026-09-02 edition. Since that edition the wide probe sends a second request presenting a named AI crawler's published user-agent string, which /methodology has said since the same day. The string is now generated from the probe's own output each run and names both legs, and the file carries an `instruments` block stating the identity each instrument presents.",
   "why": "The disclosure file contradicted the page that links it as the disclosure, for two days, on the one question the file exists to answer. Rebuilding the sentence from the probe means the next change of instrument rewrites it rather than outliving it. The /methodology paragraph that justified the panel's audit tier by the wide probe's honesty was rewritten in the same pass, for the same reason. No count changed."
  },
  {
   "id": "2026-09-04-20",
   "date": "2026-09-04",
   "what": "The word \"probed\" was used for two different quantities: the 8,002 domains the wide probe requested (on /status and /methodology) and the 7,491 it reached (on the home page, /explore and in the feed field refusal_headline.probed). Public copy now uses one word per quantity — asked for the first, reached for the second — and every sentence beside a feed field named `probed` says which of the two that field carries. No count changed; the two numbers were always printed correctly and always agreed across surfaces.",
   "why": "Two quantities under one word is how a figure gets quoted against the wrong denominator. The feed keys keep their names so nothing citing them breaks; the naming is fixed in the prose that surrounds them."
  },
  {
   "id": "2026-09-04-19",
   "date": "2026-09-04",
   "what": "/status said 53 ccTLD suffix groups were \"listed on /world\". /world draws three for a free reader and names the remaining fifty as a count. The sentence now says what is public: how many groups are in the edition's dataset, how many carry a rate at n >= 100, how many /world draws, and how many the Terminal's Segments tab rates.",
   "why": "The count was right and the verb was wrong, on the page whose own header says it cannot drift from reality. The free draw is now read out of /world itself at build time, so the two cannot disagree again."
  },
  {
   "id": "2026-09-04-18",
   "date": "2026-09-04",
   "what": "The per-edition hash table on /status said the published SHA-256 \"certifies the file you downloaded\". It did not. The digest was taken from the census machine's own copy of each private file, and the artefacts a subscriber can obtain from the dashboard — the per-domain JSON download and every CSV — are rebuilt in the browser out of rows already parsed there, so they carry the same data as a different sequence of bytes and hash differently. A reader who tested the claim on the per-domain download got the published byte count exactly and a different digest. The files are now hashed as served (the app/private copies the API returns), each row prints the route it is served from, and the paragraph states what a reader can compare the digest against: the bytes a route returns, hashed as they arrive.",
   "why": "A number printed as a proof that fails to verify is worse than no number, because the reader is the one who has to stand behind it. No hash value changed meaning; what changed is which copy is hashed and what the page tells you to check it against. In the same pass the table stopped promising \"the last five\" editions while printing one: it prints every edition this machine has hashed, and says how many that is."
  },
  {
   "id": "2026-09-04-17",
   "date": "2026-09-04",
   "what": "The 30 August issue put two dated correction notes ahead of its own first figure.",
   "why": "A published issue is never rewritten, so the notes stay, verbatim. They are now a two-line collapsed block beneath the issue's lead sentence, and open by themselves only for a reader who arrived on the /status link, where the correction is what they came for."
  },
  {
   "id": "2026-09-04-16",
   "date": "2026-09-04",
   "what": "/status stat tiles had no gutter and were glued to the ITEM/VALUE table beneath them, and the page had no single source for when each instrument started.",
   "why": "A 12px gutter and 20px of separation, and a new 'Since when' table derived from the built files: the census start from series.json, the wire legs from the per-edition metadata in the wire history, exhibit tracking and the registry capture from their own archives. An instrument that has not run says 'not yet run' and never carries a date. Every other surface quotes this table rather than typing a start date."
  },
  {
   "id": "2026-09-04-15",
   "date": "2026-09-04",
   "what": "/anatomy's rank-band table was headed 'Tranco band'.",
   "why": "The frame has been CPI-50K v1 — a custom Tranco list built from Cisco Umbrella and Majestic only — since 2026-08-23, and every other surface names it that way. The header now reads 'CPI-50K v1 rank band'. The bands and the figures are unchanged."
  },
  {
   "id": "2026-09-04-14",
   "date": "2026-09-04",
   "what": "The home page's 'N times harder on training' multiples could not be derived from the bars beside them, and no bar named the role that is the card's whole claim.",
   "why": "The multiple is the mean of the vendor's training-role crawlers over the mean of its others. Each bar now carries its role, both means are drawn as rules across the bars, and the arithmetic is printed under the group. The multiples themselves are unchanged."
  },
  {
   "id": "2026-09-04-13",
   "date": "2026-09-04",
   "what": "/world opened with roughly 700 words of preface before its first figure, and its distribution card was titled 'Declared blocking by suffix group' while drawing three of the fifty-three groups it keys.",
   "why": "The card is now titled for what it draws, the finding is stated above the bars, and the preface is two sentences. The suffix-group counts (groups keyed, groups rated, the Terminal's larger count, and the number drawn here) are rendered from the payload rather than typed, so the three surfaces cannot disagree."
  },
  {
   "id": "2026-09-04-12",
   "date": "2026-09-04",
   "what": "/check drew the domain's percentile as a 60px solid block with a marker and no scale.",
   "why": "The page's own stylesheet set that bar to 8px, and theme.css — which loads after it — defines a class of the same name as a 60px coverage bar, so the percentile bar had silently been rendered at the coverage bar's height. The percentile now has its own class, a ticked 0-100 axis, and one sentence naming what the percentile counts."
  },
  {
   "id": "2026-09-04-11",
   "date": "2026-09-04",
   "what": "/why's crawl-to-click bars carried hand-typed widths that did not follow the ratios printed beside them: the 2,500:1 bar was drawn shorter than the 1,091:1 bar, and DuckDuckGo's 0.3:1 was a 4% sliver.",
   "why": "log(x) is negative below 1, so a ratio under 1:1 has no honest length on an unclamped log axis. The widths are now computed from the ratios on a log axis clamped at 0.1:1, and the note beside the chart states the clamp and the value drawn at full width. The ratios themselves are Cloudflare's published figures and are unchanged."
  },
  {
   "id": "2026-09-04-10",
   "date": "2026-09-04",
   "what": "/explore's price panel drew four bars on three scales — two prices in dollars and two shares of two different populations — with one bar drawn at minimum width under the caption 'the true share is too small to see'.",
   "why": "A panel needs one denominator, and a value too small to draw is a number in a list, not a bar with an apology. The four figures are now four labelled rows. Every figure is unchanged; only the form is."
  },
  {
   "id": "2026-09-04-09",
   "date": "2026-09-04",
   "what": "/anatomy overlapped two of the three segment labels under the 'blocks at least one' bar whenever the third segment was small, which it is in every edition published so far (94 of 6,337).",
   "why": "The three labels were placed at x positions derived from the segment widths with no check that a label fits under its own segment. A fix on 3 September moved one label down a line, which corrected that instance and not the cause. The three labels now go through a collision pass that centres each on its segment, separates them, falls back to a stacked legend where they cannot fit on one line, and draws a leader from any label that had to move. Tested over 20,000 randomised width combinations with no overlap."
  },
  {
   "id": "2026-09-04-08",
   "date": "2026-09-04",
   "what": "/explore's 'Edition by edition' chart drew three points under a heading promising every edition since the baseline, and drew no point at all for the first edition.",
   "why": "It plotted the change in each crawler's block rate against the edition before it, so four published editions produced three changes and the baseline edition had no predecessor to be measured against. It now defaults to Level — one dot per published edition, on an axis anchored at zero — with Change one click away and stating in the plot how many changes it is drawing. Both modes draw unconnected dots: with four points a line claims a direction the data cannot support. No figure changes; the chart now draws all of them."
  },
  {
   "id": "2026-09-04-07",
   "date": "2026-09-04",
   "what": "The control leg's own 402 count was published in the free history file from the first edition and exported in no CSV and on no page. It is now the wire_control_p402 column of the aggregate series CSV.",
   "why": "The column existed as a number and not as a deliverable. Its denominator is the reached set of that edition — the probe targets that returned an HTTP status on the control leg — the same denominator as every other wire_* count in the row.",
   "note": {
    "date": "2026-09-05",
    "text": "This entry was false when it was published. The `wire_control_p402` column was not in the aggregate series CSV on 4 September: the export emitted thirteen wire_* columns and this was not one of them. The count itself is real and always was — the free history file carries it as `wide_probe.p402` (46, 45, 47, 47 for the four editions to 2026-09-02) and wire-aggregate.cjs computes it — but a correction that announces a deliverable which does not exist is worse than the missing deliverable. The entry text stands as published. It is withdrawn by the 5 September entry above. The column shipped later the same day: 2026-09-05-03 records it, and the export now emits fourteen wire_* columns. This note was itself corrected before publication — drafted in the present tense while the column was still missing, never deployed in that form, and repaired rather than left to contradict the file a subscriber downloads."
   },
   "withdrawn_by": "2026-09-05-01"
  },
  {
   "id": "2026-09-04-06",
   "date": "2026-09-04",
   "what": "All thirteen wire_* columns of the aggregate series CSV have been blank on every row since D28 added them.",
   "why": "The columns read history.json's per-edition wide_probe block. That block is written into history/<edition>.json by the census for the current edition and back-filled into archived editions by build-wire-history.cjs, and neither had run in the order that carries it into app/data/history.json, so no edition carried one. compute-dashboard.cjs now computes the per-edition aggregate from private/wire-history.json — the same archived probe rows, counts only — and fills any edition whose own snapshot has no block. Census-written values are never overwritten. Nothing is newly measured: every column is a count over rows already published in the wire history."
  },
  {
   "id": "2026-09-04-05",
   "date": "2026-09-04",
   "what": "/methodology said the 1,041 domains that disallow our crawler in robots.txt are excluded before their homepage is fetched, without saying which instrument that applies to. It is true of the reachability sweep and false of the wide probe.",
   "why": "probe-wide.cjs builds its target list as `rank <= RANK_N || blocksAny` read from the census CSV, and applies no robots.txt check of its own before requesting each target's homepage. sweep-reachability.cjs does read robots.txt first and records 'disallowed_by_robots' without fetching. Two instruments, two politeness rules, one sentence covering both. /methodology now states the split, and every wide-probe figure on the site (reached, refusals, 402s) is computed over a target list from which those domains were not removed in advance. No published figure changes; what changes is what the page claims about how the set was built."
  },
  {
   "id": "2026-09-04-04",
   "date": "2026-09-04",
   "what": "The reachability sentence on /explore omitted the 260 domains whose fetch was refused, reset or failed TLS (\"other\") from the never-answered share; the share now includes them and moves by about half a percentage point. The counts were always on the page; the sentence was the thing that dropped them.",
   "why": "A rounding-and-omission error in the page copy, found in the 4 September ledger walk (L-014, L-213). No published count changed."
  },
  {
   "id": "2026-09-04-03",
   "date": "2026-09-04",
   "what": "From edition 2026-09-02 the wide probe's control request was described on the home page, Explore, the methodology page and in the Terminal as 'an ordinary browser'. It is our own self-identified crawler, CrawlPriceIndexBot/1.0, sent unsigned on this instrument — the honest-identity control. Every count that reads 'refused the crawler while serving …' is now worded 'while serving our own self-identified crawler (the honest-identity control)'. No count changed.",
   "why": "The counts were computed against the identified-crawler request from the first day (the dashboard builder's own basis field says so), and the pages above it drifted to a description the instrument never matched. A control that is a named crawler and a control that is a browser answer different questions — a site can refuse both a named AI crawler and our crawler while still serving a person — so the mislabel overstated what the pair shows. The description is corrected on every surface; the numbers are left as published. A request carrying a Chrome user-agent string is planned as a third, separately counted leg from the 6 September edition, and will be dated on the methodology page when the first edition carrying it publishes."
  },
  {
   "id": "2026-09-04-02",
   "date": "2026-09-04",
   "what": "The selective-exclusion matrix on the Crawlers tab (row crawler blocked while column crawler allowed) was a share of the 50,000-domain frame while every other rate on the tab, and the sub-line printed directly beneath it, is a share of the domains with a readable robots.txt. It is now computed on the parsed set. No cell count changed and no pair changed rank; every cell is multiplied by 50,000/31,488 = 1.588. Naming each cell by the pair it belongs to: GPTBot blocked while OAI-SearchBot is explicitly allowed — 149 domains — moves from 0.30% to 0.47%; GPTBot blocked while ChatGPT-User is explicitly allowed — 151 domains — moves from 0.30% to 0.48%; and the largest cell on the matrix, Bytespider blocked while ChatGPT-User is explicitly allowed — 186 domains — moves from 0.37% to 0.59%. The cell counts and the denominator are now published in dashboard.json as exclusion_matrix_counts and exclusion_matrix_denominator, so every surface reads a domain count from a key instead of reconstructing one from a rounded percentage.",
   "why": "One denominator per tab. The cell counts were right; the base was the wrong one, and the caption did not say so. An earlier version of this note named a domain count for that cell which no built file carried — it had been reconstructed by multiplying a rounded percentage back out. Publishing the counts is what stops that happening again."
  },
  {
   "id": "2026-09-04-01",
   "date": "2026-09-04",
   "what": "Two issues of The Weekly Crawl are withdrawn from the public archive: 23 August (its change set was computed across the frame change, on the 17,179 domains readable in both frames) and 24 August (computed on the 24 August validation scan, which is not a published edition). The 30 August issue stays, with a note that its change window ran from that same scan.",
   "why": "The method says frame churn is never counted as a policy change and that the 24 August scan is excluded from every change set; the archive said otherwise. The emails cannot be recalled, so the withdrawal is recorded here rather than the pages quietly amended. From the next issue, the window runs from one Sunday edition to the next — two published editions — and the issue says so."
  },
  {
   "id": "2026-09-01-01",
   "date": "2026-09-01",
   "what": "Selective treatment, the crawlers-blocked histogram and the policy-diversity histogram were computed over the 50,000-domain frame while every other rate published here is a share of the domains with a readable robots.txt. Selective treatment was stated as 1.62%; on the parsed set the same 811 domains are 2.57%. The histograms are recomputed on the parsed set, so the zero-blocked bucket no longer includes the 18,500 domains that served no robots.txt.",
   "why": "Two denominators in one product invite a figure being quoted against the wrong base, and counting a domain that yielded no reading as one that chose to block nobody is an inference about intent this index does not make. Both figures are now published side by side: 2.57% of the parsed set, 1.62% of the frame."
  },
  {
   "id": "2026-08-23-04",
   "date": "2026-08-23",
   "what": "The free public feed carried the previous edition's asymmetry ratio, change interval and reachability figures for roughly two days.",
   "why": "rebuild.cjs read the dashboard file, published from it, and regenerated it 1.6 seconds later, so the citable tier lagged a full edition. A new step runs last and refuses to write when any denominator disagrees with the parsed count."
  },
  {
   "id": "2026-08-23-03",
   "date": "2026-08-23",
   "what": "The weekly email shipped three defects: every crawler's week-over-week move printed as “0 pts”, the ccTLD count printed as zero and under a geographic label, and the subject described a rate as a share of the top web.",
   "why": "A delta object was read as a number, a renamed key was not updated, and the newsletter had no copy guard at all while the public pages and the paid dashboard both had one. An email cannot be un-sent, so the guard added in response is blocking rather than advisory."
  },
  {
   "id": "2026-08-23-02",
   "date": "2026-08-23",
   "what": "Two pages described a rate as a share of “the scanned web”, and four used a geographic label for what is a domain-suffix grouping.",
   "why": "Neither is accurate. The population is domains with a readable robots.txt inside a 50,000-domain frame, which is not the web; and a domain suffix indicates neither operator location nor audience. The copy guard that forbids both framings had only ever been pointed at six of the seventeen public pages, so the pages carrying them were never checked. It now enumerates the directory."
  },
  {
   "id": "2026-08-23-01",
   "date": "2026-08-23",
   "what": "The freshness block reported 31,506 rows with a reading while coverage reported 31,469 parsed, and its note described the larger number as the base of every published rate.",
   "why": "rebuild.cjs counted the 50,000-domain frame plus 37 research-panel domains that sit outside it. The rates were, and are, computed over the 31,469 frame rows. The count now excludes the panel and the note names what it is. In the same block, sweep_at reported the time rebuild ran rather than the time the web was read; it is now the sweep timestamp."
  },
  {
   "id": "2026-08-22-01",
   "date": "2026-08-22",
   "what": "A chart on /why was relabelled to describe training share as a proportion of AI-crawler requests, and has been reverted to a proportion of all crawler requests.",
   "why": "The argument used to justify the relabel compared a purpose split against a category split as though they nested. They do not: a multi-purpose crawler can be counted in one category and a different purpose. Four reviewers rejected the reasoning and were right. Separately, any compounded “share of all web traffic” figure for AI-training crawling is withdrawn entirely, because the layers do not share a denominator. Full account at /methodology#crawler-share. Nothing measured changed."
  },
  {
   "id": "2026-08-16-01",
   "date": "2026-08-16",
   "what": "The observed-price series stored a null price for the first two editions.",
   "why": "A parsing error. The early points are backfilled from the value observed at the time and marked as backfilled. The observed top price has been $0.50 at every observation to date."
  }
 ]
}
