Research · Read 7 October 2026 · One file, two lines

Two lines in your robots.txt. Only one has a reader on record.

A robots.txt file can now carry two kinds of line about AI. One is thirty years old: a User-agent name and an Allow or Disallow rule, and the big crawlers document how they read it. The other is new: a "content-signal" directive. A Google Search Advocate's public comment on it says that, as far as he knows, no crawler uses it.

By Rish Sadh, founderUpdated

The line nobody has said they read

In July 2026, Google Search Advocate John Mueller answered a question about this in a Reddit comment. Search Engine Roundtable reported it on 6 July. His words, hedges included:

I guess a few things ... Google doesn't use llms.txt or llms-author.txt. I don't know of any other crawler / llm confirming they're using these (other than SEO tools).

AFAIK none of the crawlers / llms use the "content-signal" robots.txt directives. It was made up by a CDN, afaik it has no effects whatsoever for any crawler or llm. Using it just adds bloat & future maintenance to your robots.txt file. You can also add other arbitrary things to your robots.txt file, crawlers just use the directives that they support and ignore the rest.

That is one person, in a comment, saying "as far as I know". It is not a Google announcement, and the page should not be read as one.

The same goes for llms.txt. This site publishes one. On this evidence it is a courtesy note, not a lever. More on that in What is llms.txt.

The line that works

The old line is the one vendors write about. OpenAI's crawler page: "To help ensure your site appears in search results, we recommend allowing OAI-SearchBot in your site's robots.txt file and allowing requests from our published IP ranges below." And in Cloudflare's summary of Google's controls: "Googlebot allows site owners to opt out of training by adding a Disallow rule to robots.txt for "Google-Extended"".

It is also being written for people now. Cloudflare, 21 August 2026: "Bot Preference Sync reflects what you've set in your AI bot configuration by updating corresponding preferences to your robots.txt". And: "For all new customers, Bot Preference Sync will be on by default".

What it writes depends on the site. In Cloudflare's recommended settings for new domains, a site that does not earn from ads allows training. Only a site that runs ads gets "Disallow AI Training". A small business site on Cloudflare does not block AI training by default.

Saying no to training, and keeping search

The worry with a training opt-out is losing search. On 15 September Cloudflare set a bar for crawlers that do both, which includes "Assurance that opting out of AI training will not affect traditional search results". Then it reported what each company had said, and the three statements are not the same:

  • "Apple has also stated that disallowing training does not impact search ranking."
  • "Google has also stated that disallowing Google-Extended does not impact search ranking."
  • "Microsoft has also stated that using NOARCHIVE will not impact search ranking."

Microsoft's is about a meta tag, not robots.txt. Cloudflare says Microsoft is "currently building the mechanism to also respect a "no training" preference in robots.txt at the domain/site level, targeted for early 2027", and "Until that support launches, selecting Disallow AI Training will not automatically convey a no-training preference to Bing through robots.txt."

Even the line that works doesn't reach everything

OpenAI's own documentation: "ChatGPT-User is not used for crawling the web in an automatic fashion. Because these actions are initiated by a user, robots.txt rules may not apply."

That is one user agent, and OpenAI says "may not", not "do not". The same page tells you to allow its search crawler in robots.txt. So robots.txt is not dead. It just has edges, and a live fetch triggered by a person is one of them. Do AI crawlers respect robots.txt goes through the other crawlers one by one.

Ours

reidify.design's robots.txt, 7 October 2026: five User-agent blocks, all Allow: /, one for all crawlers and one each for GPTBot, PerplexityBot, Google-Extended and ClaudeBot, plus a sitemap line. No content-signal line. It is served as written; nothing between the server and the reader rewrites it.

What this page doesn't say

  • That a content-signal line does harm. The comment above calls it bloat, nothing more.
  • That no crawler will ever read one. Only that nobody has said they do.
  • That OpenAI ignores robots.txt, or that robots.txt is dead.
  • That small business sites on Cloudflare now block AI by default. Cloudflare's own defaults say the opposite.
  • Anything about any particular site's settings but our own.

Check yours

Open yourdomain.com/robots.txt in a browser. Look for two things. Any line containing content-signal: on the evidence above, nothing is known to read it. And any Disallow under a crawler you want to be found by, such as OAI-SearchBot or Googlebot: that one is read, and it means what it says.