The Crawl Price Index.
Suffix groups

Suffix groups — what a domain ending
does and does not tell you.

Declared AI-crawler blocking varies widely across domain suffixes, and it is tempting to read that as a map of national attitudes. It is not one: a suffix is a string at the end of a name, and the three reasons a group’s rate moves — set out below, before the numbers — are none of them national.

Edition — · grouped by domain suffix within the CPI-50K frame
Read this first

Three reasons a suffix group’s rate moves — none of them national.

  • A suffix is not a country. It is a string at the end of a domain name. It does not establish where an operator is, who owns the site, where its audience is, or where it is hosted. Plenty of organisations never use the suffix of the place they operate from, and most suffixes can be registered by anyone, anywhere.
  • Suffix groups differ in rank composition. Rank correlates with declared blocking in our own data. A group weighted toward higher-ranked domains will show a higher rate than one weighted toward the long tail, before any question of policy arises. Part of any gap you see below is rank, not attitude.
  • Sample sizes differ by an order of magnitude. Groups here range from a few dozen domains to over a thousand. A small group moves several points when a handful of sites change a single line in a file.

The dashboard’s Segments tab compares each group against the whole-index rate and against rank bands, which is the cut that begins to separate these effects. This page shows the raw distribution only.

The distribution

Declared blocking by suffix group.

—

—

Share of each group’s domains that explicitly block at least one of the 18 tracked AI crawlers, among domains serving a readable robots.txt. Sample size (n) is shown for every group, and a rate is published only above the floor named above — the same floor the Terminal’s Segments tab uses, so the two pages name the same highest and lowest groups.

Two suffix conventions, and why the n for .uk differs between pages

This page keys a domain by its country-code suffix and lists a two-label public suffix as its own group (.co.uk is a group here, separate from .uk). The Terminal’s Segments tab and the band table on Explore key by the last label only, so a .co.uk domain appears under .uk there. The two pages’ n for .uk therefore differ, and neither is wrong: they are two conventions, stated here.

Highest observed: —. Lowest: —. Neither is a ranking of anything but this measurement. Groups below the floor are listed with their n and no rate: —.

The baseline

The generic-suffix baseline.

Generic suffixes (.com, .org and similar), which carry no geography at all: —
—
Per-crawler detail

Per-crawler detail by suffix group.

The headline rate hides which crawler is being blocked. Some groups show a wider gap between training-role and search-role crawlers than others. The same three confounds above apply to every column.

ccTLDGPTBotClaudeBotGoogle-Ext.Any AI