Should you block AI crawlers? GPTBot, training bots and search bots, in robots.txt and Cloudflare · Citable

View this page

Whether to block AI bots: the three kinds of AI crawler, what blocking each one costs, how to block AI crawlers in robots.txt and Cloudflare, what to do about bots that ignore robots.txt, and a rule that keeps you in AI answers.

Price
USD 59
Organisation
Citable

Description

written for teams no score out of 100 no ten-step checklist written for teams no score out of 100 no ten-step checklist written for teams no score out of 100 no ten-step checklist

The three kinds of AI crawler

“AI bots” are treated as one thing. They are three, they do different jobs, and the decision is different for each.

Kind What it does Examples If you block it

Training Collects pages ahead of time, to train a model. GPTBot , ClaudeBot , CCBot , Bytespider Your pages stay out of what the model learns from. Visitors and search results do not change.

Search index Builds the index an AI answer is looked up in. OAI-SearchBot , Claude-SearchBot , PerplexityBot You stop appearing as a source in that engine’s answers.

Live fetch Opens a page at the moment a person asks about it. ChatGPT-User , Claude-User , Perplexity-User The person who asked gets an answer written without your page.

Should I block GPTBot?

It depends on what you publish. If your content is the product — journalism, research, a paid course — keeping it out of training is a reasonable position, and blocking GPTBot does that. If your site exists to be found — a store, a service, a software product — being part of what a model knows is mostly in your favour, and blocking it gains you nothing you can measure.

Either way, blocking GPTBot does not take you out of ChatGPT’s answers. That is decided by OAI-SearchBot and ChatGPT-User, which are separate names with separate rules. You can refuse training and still be cited.

The block most sites make by accident

The expensive mistake is a rule that was meant for training bots and catches all three kinds: a blanket Disallow for every AI name copied from a blog post, or one switch in a CDN that blocks “AI bots” as a group. Cloudflare now offers to block AI crawlers when a domain is added, so many site owners have this on without having chosen it. The site keeps working, Google keeps ranking it, and it quietly stops appearing in AI answers. Nothing reports the loss.

How to block AI crawlers in robots.txt

Name each one. This keeps training bots out and leaves search and live fetch alone:

User-agent: GPTBot

Disallow: /

User-agent: ClaudeBot

User-agent: CCBot

User-agent: OAI-SearchBot

Allow: /

User-agent: ChatGPT-User

Allow: / A group that names a crawler replaces the * group for that crawler; it does not add to it. Every name, and where the rule goes on Shopify, WordPress, Wix, Next.js and the rest, is in the AI crawler list .

AI bots ignoring robots.txt, and bots that hammer a server

robots.txt is a request. A well-behaved crawler follows it; nothing in the file stops one that does not. If a bot is ignoring your rules, or fetching so fast it loads the server, the rule has to be enforced where requests arrive: block or rate-limit that user agent at your CDN or web server. Cloudflare , Nginx and Apache each have a page for how.

One more thing worth knowing: a user agent is a claim. Anyone can send a request that says it is GPTBot. The real ones come from addresses their operators publish, which is how a genuine crawler is told from an imitation.

A sensible default

Allow search-index and live-fetch crawlers. They are how you appear in answers.

Decide training on purpose — per crawler, not as a group.

Check the layer in front of your site , not just robots.txt. The edge wins.

Rate-limit rather than block a crawler that is merely too fast.

Look at what they are given. Allowing a crawler in is half of it.

Blocking is a decision about your content. Being blocked by a default you never saw is not a decision at all.

See who is reading you before you decide

It is hard to choose well without knowing who comes. The free scan shows which AI engines your robots.txt lets in today, and which it refuses. Connected to Citable, every read is listed as it happens — the crawler, the page, whether its address was really the operator’s — and the crawlers you allow are given the content alone instead of the whole page. What each engine does with a site is under AI engines .

The chain, drawn Access, then comprehension, then the choice.

Access — an engine you refuse never reads anything.

Comprehension — the facts have to survive the markup.

Choice — which stays the model’s, whatever anyone sells you.

Read next The rest of the set.

6 min read AI search optimization

How an assistant decides what to say about your site, and which of those steps a team can actually influence.

7 min read What GEO can and cannot do

What generative engine optimisation means, what it borrowed from SEO, and which parts of it are sold as certainty but aren't.

5 min read How AI crawlers read a website

Rendering, fetch budgets, robots.txt, and why a JavaScript-drawn price is invisible to almost every engine reading you.

8 min read Generative engine optimization (GEO)

What it means, how an AI engine chooses its sources, the strategies with evidence behind them, and what GEO tools actually do.

7 min read Answer engine optimization (AEO)

What it means, three before-and-after examples, the practices that matter — and the step that comes before any of them.

6 min read How to get cited by ChatGPT

How ChatGPT picks the pages it cites, and the six things a site can actually do to show up in its answers.

5 min read Can ChatGPT read my website?

How it reads a site, why it does not run JavaScript, the five things that stop it, and how to check in a minute.

5 min read llms.txt: what it is, and what it does not do

What goes in the file, a working example, best practices — and whether any AI search engine actually reads it.

5 min read utm_source=chatgpt.com, and tracking AI traffic in GA4

What the tag on your URLs means, why ChatGPT adds it, and how to see every AI engine in Google Analytics.

5 min read GEO vs AEO vs AIO vs LLMO

The full form and meaning of each, where the terms came from, and the four steps underneath all of them.

Or just look at it One page, served both ways.

What survives each trip, with the token count for both.

About Citable Who is behind this page, and what we stand on.

When an AI agent — ChatGPT, Claude, Perplexity or any of the 81 that Citable recognises — opens a page, Citable removes the theme, scripts, navigation and layout and serves only the content: the same facts, at the same URL, on the site's own domain. People and Google still get the full designed page.

Who opens the page What they are given

A person in a browser The full designed page, untouched.

Googlebot, Bingbot and other search engines The full designed page, untouched.

An AI agent — GPTBot, ChatGPT-User, ClaudeBot, PerplexityBot and the rest The content only: no UI, no scripts, the same facts at the same URL.

Measured, not estimated. A real product page of 386.8 KB is served to AI agents as 4.6 KB — 98.8% smaller, with the same facts. Every figure on this site comes from a fetch anybody can repeat.

One list. The 81 AI agents named on these pages are the list the product itself runs on. A page here cannot describe a crawler Citable does not serve.

Search is never touched. Googlebot and Bingbot always get your real page. The clean copy is for AI agents only — at the same address, with the same facts.

Checked, not trusted. A user agent is a claim, and anyone can send one. Where an operator publishes its addresses, every read is checked against them and recorded as verified or not.

Nothing about your visitors. Citable sees which crawler read which page. It does not see who your customers are, and their details never reach us.

Honest about limits. Nobody can promise a citation or a ranking, and we do not. What we show is what an AI agent was given, and whether it sent somebody back.

How it works · The measurements · What it touches and stores · Why Citable · Pricing · Questions, answered plainly · Talk to us

Free · about 20 seconds · no signup Find out whether AI has ever read your site.

Type your address. We fetch a page exactly as GPTBot does, count the tokens, and score fifteen things an answer engine needs. If the log then stays empty for a month, that is worth knowing too.

No contract · cancel any time · Googlebot untouched