AI Usage Preferences: Telling AI Systems How They May Use Your Content
Content-Signal lines in robots.txt say yes or no to search, AI answers and AI training. What we check, where the line goes, and the mistake most sites make.
What we check
We read your robots.txt and look for Content-Signal lines. A line looks like this:
User-agent: *
Content-Signal: search=yes, ai-input=yes, ai-train=no
Allow: /
Each key takes yes or no:
search: building a search index and showing your pages in search results.ai-input: giving your page to an AI model so it can answer a question, as ChatGPT search and Perplexity do.ai-train: training or fine-tuning a model on your content.
A key you leave out means you state no preference for that use. We also read the IETF draft form, Content-Usage: train-ai=n, and count it the same way.
This check does not change your score. Crawlers follow these signals by choice, and the standard is still a draft, so we report it and do not reward or punish it.
Where the line goes
A Content-Signal line belongs to the User-agent group above it. Two placement mistakes are common.
First, a crawler with its own group does not read the User-agent: * group. If your file has a User-agent: GPTBot group, GPTBot ignores the signal under *. Copy the line into that group too.
Second, never put the line between two User-agent lines:
User-agent: GPTBot
Content-Signal: ai-train=no
User-agent: Googlebot
Disallow: /private
Google reads GPTBot and Googlebot as one group here, so both are kept out of /private. Some other parsers end the group at the Content-Signal line and let GPTBot in. Put the line after the last Allow or Disallow rule of its group and the problem goes away.
Our default, and why
The robots.txt fix file in your report adds search=yes, ai-input=yes, ai-train=no. GenReady exists to help AI systems cite you. ai-input=no asks AI answer engines not to use your page, and that works against being cited. Training is a separate decision, so the default says no. Change it to yes if you want your content in training data.
If ai-input=yes while robots.txt blocks a crawler that fetches pages for AI answers, such as PerplexityBot or ChatGPT-User, the two rules contradict each other. We flag that.
Cloudflare
If Cloudflare manages your robots.txt, it adds its own block at the top of the file, with its own Content-signal line. You cannot change that block by editing the file on your server. Change the signals in the Cloudflare dashboard, in the managed robots.txt setting. Our fix file leaves the block alone and tells you so.
๐ก Quick win
Open your robots.txt and search for User-agent. Every group that names an AI crawler and does not block it completely needs its own Content-Signal line.
