Content is Everything
The folklore ledger

Claims we reject, and why

Maintained by the Content is Everything team · 16 entries (8 contradicted, 5 unsupported, 3 context-dependent) · revision 6, 2026-08-21

AI search has produced a fast-growing body of advice, and much of it has never been tested. Some of it conflicts with what the platforms themselves publish. This page is a public ledger of popular claims about ranking and AI visibility, each with a verdict, the evidence behind it, and the date we last checked that the evidence still says what we claim it says.

Our measurement study forms the backbone of this ledger. Today's verdicts rest on what the platforms have published. As the study runs, entries gain a further section recording what we measured: the same real small-business questions, put to the major AI systems again and again, with cited and uncited pages measured side by side. Where a claim is testable, we say so. Where our data contradicts our own verdict, we change the verdict. Anchoring the ledger to an ongoing study is what separates evidence from criticism.

Three rules govern every entry. Every verdict cites its evidence, with a link and a date. Verdicts change when evidence changes. We publish reversals, because a ledger that records only the times we were right would amount to marketing.

How to read a verdict

Contradicted conflicts with authoritative platform guidance or reliable evidence. Unsupported commonly repeated without adequate evidence. Context-dependent holds only under conditions the advice usually omits. Experimental worth testing and not yet established.

The entries

Every claim in the ledger, with its current verdict. Each row links to the full entry and its sources.

1 · “Create an llms.txt file to rank in Google's AI results” #

Contradicted (for Google) · last reviewed 2026-08-07

Google states that it does not use llms.txt for Search or for its generative features, and that publishing one will neither harm nor help a site's visibility or rankings in Google Search. Creating the file cannot improve visibility in AI Overviews or AI Mode, and advice that sells it as a Google ranking intervention conflicts with the operator's own statement.

You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn't use them.

Google Search Central, Guide to optimizing for generative AI features on Google Search

What we measured 2026-08-03

Across 2,572 recorded AI-crawler fetches of our own sites — identified by user agent and verified against operator-published address ranges — the number of requests for llms.txt was zero. Not a small number: none. The file was present and reachable throughout.

This does not prove no system anywhere reads one. It does mean that on the sites we can observe directly, over the period we can observe them, no AI crawler asked for the file even once.

Method, sample and stated limits: Reading your crawler logs: the evidence you already own

What is defensible: the file costs little to publish, and if a system you care about documents that it reads one, publishing it does no harm. Note what the evidence does not yet contain: OpenAI's own publisher guidance describes robots.txt and crawler access, and says nothing about llms.txt. Treat it as an optional experiment, not a mechanism. Advice that sells it as an AI-ranking fix belongs in this ledger.

In the manual: Technical discoverability.

Sources

2 · “Break your articles into small, AI-friendly chunks” #

Contradicted (as a universal rule) · last reviewed 2026-08-06

Google states that there is no requirement to break content into tiny pieces for AI to understand it, that its systems can understand multiple topics on a page and surface the relevant piece, and that there is no ideal page length. Sections and headings should follow the needs of the reader, which good writing required long before AI search existed.

There's no requirement to break your content into tiny pieces for AI to better understand it. Google systems are able to understand the nuance of multiple topics on a page and show the relevant piece to users.

Google Search Central, Guide to optimizing for generative AI features on Google Search

What is defensible: clear structure helps readers and may help a system isolate a relevant passage. Google's own wording allows that shorter or longer pages can work well depending on audience and subject. That argues for well-organised documents written for a reader. It does not argue for slicing content to a rumoured machine-preferred size.

In the manual: What makes a page citable.

Source

3 · “Add special AI schema to your pages” #

Contradicted · last reviewed 2026-08-06

No recognised structured-data vocabulary grants access to Google's AI Overviews or AI Mode. Google states plainly that structured data is not required for generative AI search and that there is no special schema.org markup to add. Any product selling "AI schema" is selling a thing the platform says does not exist.

Structured data isn't required for generative AI search, and there's no special schema.org markup you need to add. However, it's a good idea to continue using it as part of your overall SEO strategy.

Google Search Central, Guide to optimizing for generative AI features on Google Search

What is defensible: ordinary structured data, describing what is visibly true on the page, supports explicit interpretation and keeps a page eligible for rich results. Google recommends continuing to use it as part of an overall SEO strategy. It is a useful foundation and not a citation mechanism.

In the manual: Technical discoverability.

Source

4 · “Write a separate page for every long-tail prompt” #

Contradicted (as a strategy) · last reviewed 2026-08-06

The strategy is built on a real mechanism. Google's AI features do issue a set of concurrent related queries for one question, which Google calls query fan-out. The inference drawn from it is where the strategy fails: Google states in the same guidance that creating separate content for every possible variation of how people might search, primarily to manipulate rankings, violates its scaled content abuse spam policy. The spam policy names generating many pages without adding value for users as an example. The strategy aims at the mechanism and walks into the penalty.

Creating separate content for every possible variation of how people might search primarily to manipulate rankings violates Google's scaled content abuse spam policy. This is also an ineffective long-term strategy.

Google Search Central, Guide to optimizing for generative AI features on Google Search

What is defensible: covering the legitimate sub-questions of a topic, where each section has genuine value to a reader. One useful resource beats a hundred near-duplicates, by Google's stated policy as well as by taste.

In the manual: Query fan-out.

Sources

5 · “AI-generated content is automatically penalised” #

Contradicted · last reviewed 2026-08-07

Google's published position since February 2023 is that appropriate use of AI or automation is not against its guidelines, and that its systems reward quality however content is produced. What it polices is the result. Using automation, including AI, to generate content with the primary purpose of manipulating ranking is a violation of the spam policies, and the current spam policy names generative AI tools explicitly in its scaled content abuse examples. The line is drawn at purpose and quality, not at authorship.

An independent authority describes the same thing from the other side. Wikipedia's editors maintain a field guide to the writing conventions typical of AI chatbots, and their explanation of why the writing is recognisable is the most useful account of the problem we have read anywhere. A language model regresses to the mean: it drops the specific, unusual, nuanced facts — which are statistically rare — and replaces them with generic, positive description, which is statistically common. Their example: the highly specific “inventor of the first train-coupling device” becomes “a revolutionary titan of industry”. The subject ends up simultaneously less specific and more exaggerated. That is exactly the failure that makes a page not worth citing, and it is a property of unsupervised generation rather than of authorship.

Appropriate use of AI or automation is not against our guidelines. This means that it is not used to generate content primarily to manipulate search rankings, which is against our spam policies.

Google Search Central Blog, Google Search's guidance about AI-generated content

What is defensible: the caution underneath the myth. AI-assisted commodity content, the generic summary any system could produce, is what the current guidance says Google does not want to surface. The problem is commodity, whoever the author. Google's own FAQ also suggests disclosing AI involvement where a reader would reasonably wonder how something was made. Our position on our own process is on the About page: drafted with AI, refined and signed by a named human, evidence required throughout.

In the manual: Commodity and non-commodity content.

Sources

6 · “Structured data makes AI systems cite you” #

Unsupported (as a citation lever) · last reviewed 2026-08-06

Structured data can make pages eligible for enhanced search displays, and it gives machines an explicit statement of what a page describes. No platform states, and no adequate evidence shows, that it causes citation in a generated answer. Google's guidance lists overfocusing on structured data among the things a site does not need to do for generative AI search.

What is defensible: structured data as foundation: accurate, matching the visible page, maintained. This claim is one our measurement study can inform, because we record schema presence and type on cited and uncited pages alike. We will update this entry with what we find, in either direction.

Under measurement: Signals dataset records schema presence and type on cited and uncited comparison pages. When the benchmark reports, this entry gains a "what we measured" section, whichever way the result falls.

In the manual: Measurement.

Source

7 · “More backlinks means more AI citations” #

Unsupported (as stated) · last reviewed 2026-08-06

Links and external reputation support discovery and authority, and no platform disputes that. Citation in a generated answer also depends on relevance to the specific question, on what the page contains, and on selection behaviour that varies by system and by run. A general link-building campaign is a different thing from a citation strategy. Google's spam policies define link spam as creating links to or from a site primarily to manipulate rankings, and name link exchanges, bought links and low-quality directory links among the examples, so the industrial version of this advice carries a real downside.

Link spam is the practice of creating links to or from a site primarily for the purpose of manipulating search rankings.

Google Search Central, Spam policies for Google web search

What is defensible: genuine external corroboration, meaning independent sources your customers already trust describing you accurately. Which sources matter, and how much, is a question our study is designed to measure.

Under measurement: Signals dataset records referring-domain profiles for cited and uncited comparison pages. When the benchmark reports, this entry gains a "what we measured" section, whichever way the result falls.

In the manual: Entity and trust.

Sources

8 · “This tool's AI visibility score predicts your citations” #

Unsupported · last reviewed 2026-08-07

No third party has access to the internal systems of Google, OpenAI or Perplexity, and Google says so directly. Citation output varies between engines, between phrasings of the same question, and between repeated runs of the identical question. We measure that variation directly. A single score claiming to predict citation probability compresses away the uncertainty that defines the territory.

Be wary of third-party tools that promise ranking success or claim to use "internal" Google metrics. No third-party tool has access to our internal ranking or AI systems.

Google Search Central, Guide to optimizing for generative AI features on Google Search

What we measured 2026-08-01

The most useful evidence we have for this entry is a failure of our own. On its first run, our recalibrated signal test reported a confident result: a jump in entity resolution on one engine, with a confidence interval that excluded zero. It was wrong. The interval had zero width — the procedure resamples prompts, only two branded prompts existed, and when both moved the same way every resample returned an identical figure. The interval excluded zero by arithmetic, not by evidence.

We caught it, published it, and added two guards: a zero-width interval now fails the test rather than passing it, and below five prompts the verdict is withheld as underpowered. We record it here because it is exactly the failure this ledger exists to police, and it appeared in our own instrument within minutes of it being built. A tool that has never reported its own false positive has not been looked at hard enough.

Method, sample and stated limits: Prompt-set testing: how we measure, and what breaks

What is defensible: measuring observable readiness. Can systems fetch your pages, is the content extractable, is the entity clear? Report those as separate dimensions with their limits stated. Where a first-party number exists, prefer it: Google now reports generative-AI impressions in Search Console, which is a real measurement rather than an inference, within its own limits (see entry 12). Measurement can offer that honestly. It cannot offer a prediction.

Under measurement: Run-to-run and engine-to-engine citation variance is measured directly by the benchmark. When the benchmark reports, this entry gains a "what we measured" section, whichever way the result falls.

In the manual: Why citation counts are not enough.

Source

9 · “Rewrite your pages in AI-friendly language so the models understand you” #

Contradicted · last reviewed 2026-08-06 · added 2026-08-06

Google states that you do not need to write in a specific way just for generative AI search, because AI systems understand synonyms and general meaning and can connect a question to content that does not use the same precise words. The corollary matters for small businesses being sold keyword work: Google says explicitly that you do not have to worry about lacking long-tail keywords or about failing to capture every variation of how someone might search.

This entry is new in this revision. It is added because the platform guidance that supports it is recent, and because the advice it rejects is being sold now.

You don't need to write in a specific way just for generative AI search. AI systems can understand synonyms and general meanings of what someone is seeking, in order to connect them with content that might not use the same precise words.

Google Search Central, Guide to optimizing for generative AI features on Google Search

What is defensible: writing plainly, in the words your customers actually use, because it serves the reader. That is ordinary good writing, and it was already the recommendation before AI search. It is a different act from rewording a page to satisfy a rumoured machine preference.

In the manual: Original content.

Source

10 · “Get your brand mentioned in Reddit threads and forums and the AI systems will cite you” #

Context-dependent (the mechanism is real, the tactic is not) · last reviewed 2026-08-06 · added 2026-08-06

Both halves of this claim need separating. Google confirms the mechanism: its generative AI features can show what is being said about products and services across the web, including blogs, videos and forum discussions. Google then rejects the tactic in the same breath, saying that seeking inauthentic mentions is not as helpful as it might seem, because its core ranking systems focus on high-quality content while other systems block spam, and its generative AI features depend on both.

So genuine discussion of a business by people who actually used it is part of the picture. Manufacturing that discussion is a spam-policed activity aimed at the systems designed to catch it. The advice is usually sold without that distinction, which is what puts it on this ledger.

However, seeking inauthentic "mentions" across the web isn't as helpful as it might seem. Our core ranking systems focus on high-quality content while other systems block spam; our generative AI features depend on both.

Google Search Central, Guide to optimizing for generative AI features on Google Search

What is defensible: being genuinely worth discussing, and being present where your customers already talk, as a participant rather than a planter. The distinction Google draws is authenticity, and it is not a distinction a purchased mention survives.

Under measurement: The benchmark records which source types are cited for each question class, including community and forum sources. When the benchmark reports, this entry gains a "what we measured" section, whichever way the result falls.

In the manual: Entity and trust.

Sources

11 · “Block the AI crawlers to protect your content” #

Context-dependent (it depends which crawler) · last reviewed 2026-08-07 · added 2026-08-06

The advice collapses two different crawlers into one decision, and the two have opposite consequences. OpenAI documents GPTBot as the agent a publisher disallows to exclude content from potential training, and OAI-SearchBot as the agent that must be allowed for a site's content to be included in summaries and snippets in ChatGPT. Blocking the training crawler does not remove a site from ChatGPT search. Blocking the search crawler does.

OpenAI also notes that a disallowed page may still surface as a bare link and title in ChatGPT Atlas if the URL is obtained elsewhere, and that suppressing that requires a noindex meta tag on a page the crawler is permitted to read. A blanket block therefore does not give a publisher the control the advice implies, in either direction.

For your site content to be included in summaries and snippets in ChatGPT, make sure you aren't blocking OAI-SearchBot.

OpenAI Help Center, Publishers and Developers, FAQ

What we measured 2026-08-07

Corrected 2026-08-07, hours after first publication — see the note below. Publishers are writing policy along exactly this split, and the clearest case in our sample is Medium: it disallows GPTBot and ClaudeBot — the training collectors — while permitting OAI-SearchBot, PerplexityBot, Claude-SearchBot and the live user-fetch agents. Index me for answers; do not train on me. That is the decision this entry describes, taken deliberately and written down.

The contrast case is Yelp, which blocks all nine major AI agents outright, with narrow Allow carve-outs for its editorial /article paths. Its business listings — the user-generated content that is the actual asset — are closed to every one of them.

What we published first, and why it was wrong. Our initial reading had Yelp blocking the search crawlers while permitting the training ones. It was the opposite of the truth, and the cause was our own tooling: a hand-written script read the rules between one User-Agent: line and the next, and Yelp's file lists several agents consecutively above a single shared rule block. The script saw an empty block and read it as permission. Our proper evaluator, which merges consecutive agent lines as the standard requires, disagreed within hours and was right. We have replaced the ad-hoc check with the evaluator everywhere.

Method, sample and stated limits: Who is actually open to AI crawlers? (census, 32 domains)

What is defensible: deciding training and search separately, on the record, and writing that decision into robots.txt deliberately. That is a legitimate choice with a real trade-off. What is not defensible is a blanket block sold as protection, or a blanket allow sold as a visibility tactic, with neither party naming which crawler does what.

In the manual: Crawler controls: search vs training.

Source

12 · “There is no way to see whether AI systems are surfacing your site, so you need a third-party tool” #

Contradicted (for Google, with limits) · last reviewed 2026-08-06 · added 2026-08-06

This one has changed, which is the point of keeping a ledger. Google now publishes a Generative AI performance report in Search Console, covering AI Overviews and AI Mode, and OpenAI states that ChatGPT appends utm_source=chatgpt.com to referral URLs so publishers can track that traffic in ordinary analytics. Both are first-party and free. A tool is not required to know something.

The limits belong in the same entry, or this becomes folklore of its own. Google's report shows impressions, meaning how many times links to your site were shown in a generative AI feature. It does not show the questions that produced them, and it excludes Search Labs experiments. A referral parameter counts the people who clicked through, not the far larger number who read an answer about you and never clicked.

how many times links to your site were shown to a user in a generative AI feature on Google Search

Google Search Console Help, Generative AI performance report (Search)

What is defensible: using the first-party surfaces first, and being precise about what they measure. Impressions are not citations, and clicks are not influence. The gap between what these reports show and what a business actually wants to know, which is whether AI systems recommend it and on what basis, is the gap our benchmark exists to measure.

Under measurement: The benchmark measures answer-level citation, which no first-party report currently exposes. When the benchmark reports, this entry gains a "what we measured" section, whichever way the result falls.

In the manual: Measurement, ChatGPT referrals.

Sources

13 · “Open your robots.txt to the AI crawlers and you will get cited” #

Context-dependent (necessary at best, and not the lever it is sold as) · last reviewed 2026-08-07 · added 2026-08-07

Being fetchable is a floor, not a strategy. The advice treats crawl permission as the thing standing between a business and a citation, and the sources AI engines actually reach for do not behave that way.

Access to the largest sources has moved off the public protocol altogether. Where a publisher has a commercial agreement with an AI company, that company does not need permission in a text file, and the text file tells you nothing about the arrangement. Robots.txt still governs everyone without a contract, which is most businesses — so it remains worth getting right, for the reason it always was: a page a crawler cannot fetch cannot be chosen. It is a precondition, and preconditions do not cause outcomes.

What we measured 2026-08-07

We read the robots.txt of the 32 most-cited domains across two full cycles of our own measurement, and evaluated each one against the nine major AI agents with a proper robots evaluator.

Blocking is rare. Twenty-eight of the 32 permit all nine. Only four block anything at all: Reddit and Yelp block all nine; Medium blocks the two training collectors; one directory blocks a single crawler. Whatever is deciding which sources get cited, it is mostly not robots.txt, because almost nothing in the cited set is closed.

And the most-cited source of all is one of the exceptions. Reddit, at 216 citations, is four lines: User-agent: *, Disallow: /. No named AI agents, no allowlist for the AI companies it has publicly signed agreements with — they do not need one. A site that forbids all automated access is cited more than any other publisher in our data.

Method, sample and stated limits: Who is actually open to AI crawlers? (census, 32 domains)

What is defensible: keeping your site fetchable, and deciding crawler access deliberately by purpose (see entry 11). What is not defensible is selling an open robots.txt as a route to being cited. It removes an obstacle. It does not create a reason to choose you, and the evidence above shows how weak the relationship between the two can be.

Under measurement: Our benchmark measures cited and uncited pages side by side, including their crawler policies, so the association between openness and citation is directly testable rather than assumed. When the benchmark reports, this entry gains a "what we measured" section, whichever way the result falls.

In the manual: Crawler controls: search versus training.

Source

14 · “An AI-visibility tool can measure every page that gets cited” #

Unsupported · last reviewed 2026-08-07 · added 2026-08-07

Every citation-monitoring product implies complete coverage: give it your questions and it will tell you what the engines cited. Nobody publishes the ceiling on that promise, so here is ours.

A measurement instrument that respects robots.txt cannot fetch a page whose robots file forbids it — and some of the most-cited sources on the web forbid it. Separately, a substantial share of highly-cited domains refuse an honestly-identified research client at the edge, before any policy file is even read. Both gaps are invisible in a tool's output: an unmeasurable page and an unremarkable page look identical in a dashboard.

There are only three ways to respond. Disclose the gap, as we do. Disguise the instrument as an ordinary browser and take the page anyway. Or say nothing and let the number imply a completeness it does not have. The second is the one the industry does not discuss, and it is a choice about honesty rather than engineering.

What we measured 2026-08-07

Of the 32 most-cited domains in our own dataset, 2 forbid our instrument by robots directive — Reddit and Yelp, between them among the most-cited sources we record. Those pages are named in our data, recorded as blocked, and excluded from analysis. We do not spoof a user agent to reach them.

Withdrawn on the day of publication: we first reported that a further six domains would not serve an honestly-identified research client at all. On re-measurement every one of them responded normally, so those six were transient failures in our collection script and not a policy of any kind. The claim is withdrawn rather than quietly amended. A separate and real effect remains — some sites refuse a plain client and serve a browser — but we will not put a number on it until we have measured it properly.

Method, sample and stated limits: Who is actually open to AI crawlers? (census, 32 domains)

What is defensible: measuring what can be measured honestly and publishing the boundary. Coverage figures should name their exclusions. A tool that cannot tell you which pages it failed to reach, and why, is reporting a sample while implying a census.

In the manual: Why citation counts alone are inadequate.

Source

15 · “Strip the AI tells — the em-dashes, the "stands as a testament to" — and your content will be fine” #

Contradicted (the source of the tell-lists says the opposite) · last reviewed 2026-08-21 · added 2026-08-07

Lists of AI writing tells circulate widely, and a whole genre of advice has grown on top of them: run your draft through a checklist, remove the giveaway phrases and the em-dashes, and the text is fixed. The most careful and widely used of those lists tells you directly that this does not work.

Wikipedia's editors maintain a field guide to AI writing conventions, built from real examples across articles and drafts. It is explicit that the patterns it catalogues are symptoms rather than the disease, and it asks readers not to treat the signs as the things to be fixed — because doing so removes the evidence while leaving the problem, and makes the problem harder to find.

The deeper faults are the ones worth your attention: claims nothing supports, sources that do not say what the text says they say, promotional framing where specifics should be, and generic description standing in for first-hand knowledge. A page can pass every stylistic checklist ever written and still contain none of the specific, verifiable substance that would give anyone — a reader, an editor, or an engine — a reason to rely on it.

The patterns listed here are also only potential signs of a problem, not the problem itself. [...] Please do not merely treat these signs as the problems to be fixed; that could just make detection harder.

Wikipedia · WikiProject AI Cleanup, Wikipedia:Signs of AI writing (advice page, revision 1368101818)

What is defensible: editing to remove empty phrasing because the writing is genuinely better without it. That is ordinary good editing and it needs no AI framing. What is not defensible is treating a tell-list as a remediation checklist, or selling that pass as making content safe. The list's own authors warn against exactly this.

Where we stand: A note on our own tool, because we ship one. bernard's site review counts these tells and shows a check called “Looks like AI slop”. That is a statement about how a page READS, which anyone can verify by reading it, and it is not this entry's claim in disguise: the check never tells an owner to delete their em-dashes or their curly quotes, and the remedy it offers is rewriting the words, not removing the evidence. Counting a tell and prescribing tell-stripping are different acts, and only the second is what this entry rejects.

In the manual: Commodity versus non-commodity content.

Sources

16 · “An AI-detection tool can tell you whether a page was written by AI” #

Unsupported (better than chance, not good enough to act on alone) · last reviewed 2026-08-21 · added 2026-08-07

Detector scores are increasingly quoted as though they settle the question, in SEO audits, in content procurement, and in arguments about whether a competitor's pages are “real”. The organisation with the largest practical incentive to detect undisclosed AI text — and years of doing it at scale — declines to rely on them.

Wikipedia's guidance is that detection tools perform better than random chance but carry non-trivial error rates, and are susceptible to paraphrasing, markup and spacing changes, and to models they were not trained on. A high detector score is explicitly not accepted there as grounds for deleting a page.

Human judgement fares no better and often worse. The same guidance cites research finding that people distinguish AI text from human text at close to chance, and notes the confound that makes this permanent: human writing is itself being shaped by these models, so the two populations are converging. A detector is measuring a moving target with a blurring boundary.

Do not solely rely on artificial intelligence content detection tools (such as GPTZero and Pangram). While they perform better than random chance, these tools have non-trivial error rates.

Wikipedia · WikiProject AI Cleanup, Wikipedia:Signs of AI writing (advice page, revision 1368101818)

What is defensible: using a detector as one weak signal among several, in a process where a human reads the actual text and checks whether its claims hold up. What is not defensible is a number treated as a verdict — on a supplier, a competitor, or a page — when the field's most experienced practical users refuse to let it decide anything on its own.

Where we stand: A note on our own tool, because we ship one. bernard's site review counts AI writing tells and never claims to know who wrote a page. It reports how a page reads and quotes the phrasing back; it produces no probability, no authorship verdict and no accusation, and it makes no authorship claim anywhere in the product. Nothing we sell rests on the thing this entry rates unsupported.

In the manual: Original and citable content.

Source

Claims under test

This is where the ledger grows. The following claims sit in neither the rejected nor the endorsed column. They are registered hypotheses in our study, published before data collection, and each will receive its evidence-backed verdict as results come in:

  • AI systems cite third-party editorial sources more than business-owned pages for "best provider" questions, and the reverse holds for specific factual and transactional questions.
  • Pages containing original evidence, such as figures, first-hand examples and stated methodology, are cited more often than comparable pages without it.
  • Citation sets are less stable than conventional top-ten rankings.
  • The sources cited most often are not always the sources that most shape the answer's content.
  • Measuring through a developer API tells you the same thing as measuring through the app a customer actually uses.

The full list of ten registered hypotheses, and the method for testing them, is on the methodology page.

What changed, and when

A ledger that quietly rewrites itself is worth nothing. Every revision is recorded here.

Revision 5 · 2026-08-07

A correction, published the same day as the claim it corrects. Our first reading of Yelp's crawler policy was the opposite of the truth; the cause was our own ad-hoc script rather than the source, and our proper evaluator caught it within hours.

  • Entry 11: the Yelp example is corrected and replaced. Yelp does not permit the training crawlers while blocking the search ones — it blocks all nine major AI agents, with narrow carve-outs for its editorial paths. The genuine purpose-differentiated case in our sample is Medium, which blocks GPTBot and ClaudeBot while permitting the search and user-fetch agents.
  • The cause is named rather than glossed: a hand-written check read the rules between one User-Agent line and the next, and Yelp lists several agents consecutively above one shared rule block, so the check saw an empty block and read it as permission. The evaluator we built for the study merges consecutive agent lines as the standard requires, and disagreed.
  • Entry 14: the claim that six further domains would not serve an honestly-identified research client is WITHDRAWN. All six responded normally on re-measurement; they were transient failures in our collection script. The robots-blocked figure of two — Reddit and Yelp — stands.
  • Entry 13: restated against the re-measured census. Blocking is rarer than we implied: 28 of 32 domains permit all nine major AI agents, which strengthens rather than weakens the entry — almost nothing in the cited set is closed, so robots.txt is not what is deciding citation.

Revision 4 · 2026-08-07. Two entries added (15, 16) from Wikipedia's field guide to AI writing — an editor-maintained advice page, cited at a pinned revision because it changes daily. Entry 5 gains the mechanism that explains why AI-assisted commodity content fails.

Revision 3 · 2026-08-07. The first revision carrying our own measurements. Two entries added (13, 14) and three existing entries (1, 8, 11) gain a “what we measured” section — the section this ledger was designed to grow, now populated for the first time.

Revision 2 · 2026-08-06. Every source re-verified against the live document and turned into a dated citation with a URL and a verbatim quote. Four entries added (9 to 12). Entries 1, 4, 5, 7 and 8 revised where the re-read changed what the evidence supports. Per-entry permalinks and an entry index added.

Revision 1 · 2026-08-04. First publication. Eight entries, verdicts resting on Google's published guidance, sources named but not linked.

Revision 6 · 2026-08-21. No verdict changes. Entries 15 and 16 gain a note stating where our own product stands against them, because bernard's site review now carries a check that counts AI writing tells and the distinction has to be published rather than assumed.

Submit a claim

Seen advice that belongs on this ledger? Send it to us, with a link to where the claim is being made. We add entries when a claim is widespread enough to matter and specific enough to test.

Publication of an entry means the claim as commonly stated lacks adequate support. It does not mean the person repeating it acts in bad faith. Most folklore spreads in good faith, which is why a ledger is needed.

made with bernard

Cookie settings