The Small-Business AI Citation Benchmark: method and predictions
This page is our commitment device. It records how the study will run and what we expect it to find, before we collect a single answer. When the results publish, judge them against this page.
The five questions
- Which sources do AI search systems cite for small-business questions?
- What distinguishes cited pages from comparable pages that are not cited?
- How much do citations vary between systems, between phrasings of the same question, and between repeated runs of the identical question?
- Does a cited source shape the answer, or does the system merely list it?
- Which changes can a small business make that are associated with better visibility?
Design
| Element | Pilot value |
|---|---|
| Canonical questions | 50, across five sectors (local professional services; home and local services; artists and makers; coaches and course businesses; small ecommerce) |
| Phrasings per question | 4: neutral, conversational, constraint-rich, evidence-seeking |
| Systems | 3: ChatGPT search, Google AI Mode, Perplexity |
| Repetitions | 3 per phrasing per system, each in a fresh session |
| Total recorded answers | 1,800 |
| Environment | Fixed United Kingdom setting: English, controlled accounts, no prior conversation, desktop |
We balance question intents across five kinds: discovery ("who makes handmade ceramic dinnerware in Yorkshire?"), comparison, validation, informational expertise ("what clay is best for outdoor plant pots?"), and transactional. The study avoids over-representing "who is best" questions, because many real opportunities are specialist questions a knowledgeable business can answer.
What we record
For every answer: the full text, the citation list and order, where each citation sits in the answer, and whether the system visibly searched. We fetch and snapshot every cited page at observation time, because pages change after answers are generated, and we measure each page for the same set of observable features.
The comparison group is the point. Citation lists alone produce survivor bias. For each question we also collect the pages that were left uncited: the top conventional search results, relevant local businesses, and pages one system cited while another ignored them. We measure cited and uncited pages identically, across technical condition, structure, evidence content, trust signals, and topical fit. That turns "cited pages were more likely to contain X" into a measured statement rather than an impression.
The ten predictions
Registered in advance. Where the data disagrees with us, we will publish the disagreement.
What would make the pilot fail
We state this in advance as well. The pilot succeeds only if collection can be repeated consistently, citations can be extracted and their pages fetched, human reviewers can agree on whether a citation supports its claim, and the cost of one observation is sustainable. If any of these fail, we will say so and revise the method in public.
Limits
This study observes. It does not see inside any ranking system, and it cannot prove cause. We will word observational findings as associations, and reserve causal language for controlled experiments we have not yet run. No result from this study will become a score that claims to predict citation.