What AI crawlers actually do on small-business websites
Correction note 12 September 2026, before publication
An earlier internal draft of this study said 64.4% of the traffic we could check failed the identity check. That was wrong. Our crawler list checked Google's general-purpose crawler, GoogleOther, against the wrong one of Google's published address files, so real Google fetches were filed as unconfirmed. We fixed the list, added Amazon's and Mistral's lists, which it had missed, and re-checked the history on 12 September 2026. The figures below use the corrected data. The 64.4% was never published.
bernard hosts websites for small businesses. For 39 days we logged every request to those sites that gave the name of an AI crawler, then checked each one against the addresses its operator says that crawler uses. Nine of the 22 sites were never confirmed as read by any AI crawler. Of the traffic we could check, half was not who it said it was.
The question
Small-business owners hear that AI crawlers read everything, and that they should be blocked. Neither claim usually comes with a log. We have the logs for the sites we host, so we asked: which AI crawlers really fetched these sites, what for, and how much of the traffic using their names came from them?
How we measured
Every request to a hosted site passes our edge server, the machine in front of the site, including requests answered from cache. When a request's user agent (the name a visitor gives itself) matches one of the 16 AI crawler names in our registry, we log it. We then check its address against the operator's published list. OpenAI, Anthropic, Perplexity, Google, Amazon and Mistral each publish one. In Anthropic's words: "If a crawler has a source IP address on this list, it indicates that the crawler is coming from Anthropic."
That check puts every request in one of three groups, and we never add them together:
- Verified: the address is on the operator's list.
- Claimed: the operator publishes a list and the address is not on it.
- Unverifiable: our registry holds no list for that operator, so the name can be neither confirmed nor refuted.
We keep purposes apart too. Training collects text for future models. A search index is what an assistant's search looks through. A live fetch means a person asked an assistant something and it fetched the page to answer. The three purposes are never summed.
The study covers 22 customer sites published with at least one page before 1 September 2026, from 4 August to 11 September 2026 inclusive. Three kinds of request are left out of every customer figure: requests to our own sites, reported separately below; requests to made-up hostnames a scanner guessed at; and 13,337 requests from the last two days not yet checked. No site is named, and any figure resting on fewer than five sites is withheld.
What we found
Many sites were barely read
| The 22 sites | Sites |
|---|---|
| At least one verified fetch, any crawler in our registry | 13 |
| At least one verified fetch, leaving out Google's general-purpose crawler | 10 |
| Reached by a verified live fetch | 10 |
| Reached by a verified search-index crawl | 9 |
| Reached by a verified training crawl | 9 |
| Never requested under any registered name | 8 |
Counting Google's general-purpose crawler, the median site had 2 verified fetches in 39 days; leave it out and it had none. Among the 13 sites that were read at all, the median was 654. Google calls GoogleOther "the generic crawler that may be used by various product teams", so we do not count it as an AI crawler.
Who was verified, and for what
| Operator | Purpose | Verified fetches | Sites reached | Distinct pages |
|---|---|---|---|---|
| Amazon | Training | 7,839 | 8 | 3,178 |
| Anthropic | Training | 6,454 | 9 | 4,005 |
| Google (GoogleOther) | Not stated | 3,424 | 13 | 2,281 |
| OpenAI | Live fetch | 3,173 | 8 | 393 |
| Perplexity | Search index | 1,959 | 7 | 868 |
| OpenAI | Training | 1,785 | 8 | 1,314 |
| OpenAI | Search index | 1,633 | 8 | 906 |
| Anthropic | Live fetch | 197 | 9 | 108 |
| Anthropic | Search index | Withheld: fewer than 5 sites | ||
Perplexity crawled these sites for its index 1,959 times and made no verified live fetch. All 518 requests using the name of its live fetcher, Perplexity-User, failed the check. Anthropic publishes one list for its three crawlers, so a verified Anthropic fetch proves the operator but not which crawler; its split by purpose rests on the name alone.
Half the checkable traffic was not who it claimed to be
Against the four long-standing lists (OpenAI, Anthropic, Perplexity and Google), 19,357 of 38,092 requests failed: 50.8%. Leave out Google's general-purpose crawler and it is 55.8%. Add Amazon and Mistral and it is 51.2%. The failures do not look like crawlers that moved address.
| The 19,357 that failed | |
|---|---|
| From Google Cloud's customer address ranges | 96.6% |
| On the operator's list when re-checked on 12 September | 0 |
| Asked for a page that does not exist | 97.7% (verified requests: 15.8%) |
| Addresses behind them | 286 |
| Addresses that used more than one operator's name | 69 |
| Most crawler names used by one address | 15 |
| Verified requests from any of those 69 addresses | 0 |
Google Cloud publishes those ranges as “customer-usable” address space, in a list anyone can download: cloud.json. That tells you where the machines were rented, not who rented them, and it says nothing against Google. An address that calls itself OpenAI's crawler, then Anthropic's, then Perplexity's, is not a crawler with a configuration fault. The 7,537 failing Amazonbot requests look the same: 99.6% came from Google Cloud's customer ranges and 99.0% asked for missing pages.
Some of it was not crawling at all. 11,994 requests asked for the files where websites keep passwords and keys, such as environment files and cloud credential files. They came from 226 addresses, under 16 different crawler names, and reached nine of the sites. Not one was verified. None of those files exists on these sites, so every request got a "not found".
Most traffic cannot be checked either way
80.9% of checked requests came under names for which our registry holds no address list. Nearly all of that is one training crawler fetching the same pages on one site again and again; leave that site out and the share is 21.8%. 1,764 of the credential requests above landed here, which is why "unverifiable" does not mean "probably fine".
What a log inside the site would miss
13.6% of all requests, and 29.0% of verified ones, were answered from our edge cache and never reached the site's application. See reading your own crawler logs.
Our own sites, separately
The bernard team runs three of the sites we observe: the platform's own domain, this site and a help site. They are left out above: a site about AI discoverability is not a typical small business. On those three, 6,533 requests were verified and 49.3% of checkable requests failed, much the same failure rate as the customer sites.
What this cannot show
- It is a floor, not a census. Only the 16 registered names are recognised, a crawler that gives no name leaves no row, and a site behind its owner's own caching service can be answered without the request reaching us.
- One platform's customers are not the web. These 22 sites share one host, one robots.txt policy and one page structure.
- A crawl is not a citation. A verified fetch proves a page was read. It says nothing about whether the page was quoted or recommended. See why citation counts aren't enough.
- Amazon's failures are a ceiling. Amazon's list names individual machines and was regenerated on 8 September 2026, so an older request from a machine since dropped reads as failed. Treat the 7,537 as the most that could be forged, not a count.
- Mistral's list may be stale. It holds four addresses and has not changed since February 2025. All 943 MistralAI-User requests failed it and every one asked for a missing page, but we read them as unverifiable or stale, not forged.
- Nothing here measures Google AI Overviews. They come from the ordinary search index, fetched by Googlebot, with no separate name to observe.
- The window is short. Collection began on 29 July 2026, so no trend is drawn.
The aggregated tables behind every figure were archived on 12 September 2026. If any figure turns out to be wrong, the correction will appear here with a date.