What we have measured so far, and what we have not
This site opened by promising measurement rather than opinion. Five weeks in, this is the account of what has been measured on our own pages, what those numbers can and cannot support, and what is banked but not yet reportable. Most of it is smaller than a reader might expect. We are publishing it anyway. Holding numbers back until they look better is not measurement.
This page is re-issued when the numbers change. The date above is the date of the numbers, not the date the page was written.
1. Who has read this site
Every request is recorded at the edge with the name the fetcher gave, and that name is then checked against the IP ranges the operator publishes. Only requests from inside those ranges are counted. The window queried runs 3 August to 10 September 2026; the site went live on 4 August, so there was nothing to observe before then. The site has 29 pages.
| Crawler | Operator | Purpose | Verified fetches | Distinct URLs |
|---|---|---|---|---|
| GPTBot | OpenAI | Training | 115 | 29 |
| OAI-SearchBot | OpenAI | Search index | 39 | 30 |
| ClaudeBot | Anthropic | Training | 30 | 29 |
| ChatGPT-User | OpenAI | Live fetch for a person's question | 26 | 3 |
| PerplexityBot | Perplexity | Search index | 6 | 5 |
| Claude-User, Claude-SearchBot, Perplexity-User | Anthropic, Perplexity | Live fetch, search index | 0 | 0 |
Read those as floors, and read the three purposes separately. A training scrape says nothing about whether we are surfaced today; a search-index crawl is the one that gates citation; a live fetch happens because a person asked an assistant something at that moment. Summing them makes one blended number that answers no question, so we do not. The most-fetched pages by verified crawlers, in order, were the homepage, the folklore ledger, Tracking ChatGPT referrals, What an entity is, and Author identity and provenance.
Far more traffic wears these names than earns them
Requests carrying an AI crawler's name from addresses outside the operator's published ranges are reported and never counted. Some of that is forgery, some is misconfiguration, and we cannot tell which. In the same window:
| Name presented | Requests from outside the published ranges |
|---|---|
| ChatGPT-User | 609 |
| ClaudeBot | 354 |
| OAI-SearchBot | 264 |
| GPTBot | 231 |
| PerplexityBot | 208 |
| Google-Extended | 204 |
| Claude-User | 107 |
In a 200-row sample of the unconfirmed ChatGPT-User requests, five addresses accounted for all 200, the top two for 117 and 70 of them, every one on a rented cloud server. That is one operator or a few, not a crowd. We report the shape and stop there. We have not identified who it is and will not guess.
A third group can be neither confirmed nor called forged, because their operators publish no ranges to check against: Amazonbot at 937 requests, Meta-ExternalAgent at 61, Bytespider at 35. They sit in their own column.
What this cannot show
- It is a floor, not a census. A response served from a cache we do not see, and any fetcher we do not recognise, leaves no row.
- Only known names are recognised. The registry of AI user agents was last reviewed on 29 July 2026. A crawler launched since then is invisible to it.
- Anthropic publishes one shared address list for its three crawlers, so a verified Anthropic fetch proves the operator, not which crawler. Its purpose split rests on the user-agent string alone.
- Nothing here measures Google AI Overviews. Those are served from the ordinary search index by Googlebot; no distinct agent exists to observe.
- A crawl proves a page was read, never that it was cited. Reading is a precondition: an absence explains a zero, and presence explains nothing.
2. Google Search
From Search Console, 6 August to 6 September 2026. Google's reporting runs about three days behind, so the window stops short of today. Across the site: 666 impressions, 2 clicks, 65 distinct queries.
| Query | Impressions | Average position |
|---|---|---|
| ai discoverability | 252 | 83 |
| ai content discoverability | 66 | 76 |
| ai findability | 21 | 90 |
| entities | 20 | 85 |
| content discoverability | 19 | 78 |
| commodity vs non commodity | 16 | 41 |
| non-commodity | 15 | 42 |
| Page | Impressions |
|---|---|
| AI discoverability | 375 |
| Commodity vs non-commodity | 115 |
| What an entity is (the 1 click) | 55 |
| How AI search finds sources | 32 |
| Author identity & provenance | 23 |
| Tracking ChatGPT referrals | 23 |
The two pages carrying most of that sit at average position 82 and 57. The reading is unremarkable. Google has indexed the pages and is showing them for the phrases they are actually about, a long way down. For a five-week-old site with no external links pointing at it, that is where you would expect to be.
3. The two things we measure and cannot report
This site's analytics counts only visitors who accept cookies, as the consent banner promises. Over the last 28 days it recorded too few to report, so no visitor number appears here and none will until there is one worth printing.
The same applies to clicks arriving from an AI assistant. The measurement exists, in the referrer an assistant attaches when it sends someone to a page, and is described in Tracking ChatGPT referrals. The count is not reportable for the same reason. Anyone selling a small business an AI-referral dashboard should sit with that: the instrument is real and can still have nothing to say.
4. Already published
Who is open to AI crawlers? (7 August 2026) is the one finding this site has published. It read the robots.txt of the 32 most-cited domains in one client's cited set, evaluated against nine AI agents: 28 of the 32 permit all nine, only four block anything at all, and the most-cited non-video domain blocks everybody and is cited anyway. That page carries its own sample limits and a correction issued hours after publication.
5. Banked, and deliberately not reported
The Small-Business AI Citation Benchmark was pre-registered in August 2026. No AI answers have been collected and nothing from it has published as a result. It was paused on 20 August 2026 while the measuring instrument was built out as a product tool, and restarted on 10 September 2026.
What is banked: 479 evidence bundles captured on 8 August 2026, one per page that AI engines had cited in answers to an earlier question set, each holding the page's bytes, a full-page screenshot, the response headers and a robots verdict. The page-measurement instrument, version b.9, declares 113 features and currently reads 83 of them; the other 30 are declared unbuilt rather than left out.
We could print readings off those bundles today. We will not, because no per-feature accuracy figure exists yet, and this site's standard forbids publishing a measurement whose error is unknown. An unvalidated instrument still produces confident numbers, which is the failure the folklore ledger exists to police.
The next step is the validation round: one human coder, a seeded and reproducible 60-page sample, blind, so the scoring sheet never shows the machine's reading, eight calibration cases agreed in advance, and the cut point for commodity language written down before scoring starts. The bars were frozen on 9 August 2026, before a coder saw a page:
| Class of measurement | Agreement required |
|---|---|
| Parsed features (the fact is in the markup or it is not) | 0.98 or better |
| robots.txt verdicts | 1.00 |
| Judgement features (a human has to decide) | 0.85 or better |
A measurement that misses its bar is published as not finding-grade, rather than improved until it passes. A bar moved after the result is not a bar.
What none of this shows
- Nothing here connects crawling to citation. We have no evidence that being fetched more often makes a page more likely to be cited, and we make no such claim. See why citation counts aren't enough.
- One site, five weeks. These are our own pages, on our own platform, over a short window. They describe this site. They are not a sample of the web, and they are not a baseline for yours.
- The interesting number is missing. Whether AI systems cite this site is the question the benchmark exists to answer, and it has not been answered.
If any figure above turns out to be wrong, the correction appears here with a date, the way the census correction did.