Allow or block TerraCotta on Cloudflare: AI Crawl Control, the managed robots.txt, and a custom rule

Open AI Crawl Control for the domain, find TerraCotta in the Crawlers table and set its Action to Block. If it is not listed, add a custom rule that blocks requests whose user agent contains “TerraCotta”.

This response

You are ClaudeBot (Anthropic). Citable recognised you as an AI agent, so this is https://getcitable.in/crawlers/terracotta/cloudflare with the UI removed — the content only. A browser asking for the same address gets the full designed page.

TerraCotta, and where Cloudflare keeps the rule

TerraCotta is a crawler run by Ceramic AI. It reads pages ahead of time, to train a model or to build an index.

Cloudflare does not hold a site’s robots.txt; the origin does. But it can add to that file, and it can refuse a crawler before the origin is ever asked.

Step by step

What to paste

A group that names a crawler replaces the * group for that crawler; it does not add to it. Paths the * group closes are open to a crawler with its own group unless they are repeated there.

A custom rule that blocks it (Security → WAF → Custom rules; action: Block):

(http.user_agent contains "TerraCotta") or (http.user_agent contains "Terra Cotta")

robots.txt at the origin, refusing TerraCotta:

User-agent: TerraCotta
Disallow: /

User-agent: Terra Cotta
Disallow: /

robots.txt at the origin, allowing TerraCotta:

User-agent: TerraCotta
Allow: /

User-agent: Terra Cotta
Allow: /

What trips people up

The edge answers first. If Cloudflare blocks TerraCotta, an Allow in robots.txt changes nothing: the crawler is refused before the origin is asked. A site that allows a crawler on paper and blocks it at the edge is a site that crawler never reads.

Asked, or actually refused?

Yes. A Block in AI Crawl Control or in a custom rule is answered by Cloudflare itself, and the request never reaches the origin.

Check it yourself

A 200 means the name is let through; a 403 means something in front of the page refuses it. This tests the name from your own address. A platform that checks a crawler’s address as well may treat the real TerraCotta differently.

The request:

curl -I -A "TerraCotta" https://your-site.example/

Questions

How do I block TerraCotta on Cloudflare?

Open AI Crawl Control for the domain, find TerraCotta in the Crawlers table and set its Action to Block. If it is not listed, add a custom rule that blocks requests whose user agent contains “TerraCotta”.

How do I allow TerraCotta on Cloudflare?

In AI Crawl Control set TerraCotta to Allow, and check under Security → Settings that the AI bot setting is not blocking it.

Can Cloudflare actually stop TerraCotta, or only ask it?

Yes. A Block in AI Crawl Control or in a custom rule is answered by Cloudflare itself, and the request never reaches the origin.

Does a rule for TerraCotta affect Googlebot or my Google rankings?

No. Googlebot goes by its own name and follows its own rules; a group or a firewall rule for TerraCotta does not apply to it.

Checked against Cloudflare's own documentation on 1 October 2026

Related

About Citable

Citable strips the UI from a website for AI agents and serves them only the content.

When an AI agent — ChatGPT, Claude, Perplexity or any of the 81 that Citable recognises — opens a page, Citable removes the theme, scripts, navigation and layout and serves only the content: the same facts, at the same URL, on the site's own domain. People and Google still get the full designed page.

How it works: https://getcitable.in/how-it-works · Summary for AI agents: https://getcitable.in/llms.txt


Source: https://getcitable.in/crawlers/terracotta/cloudflare · Citable · Run the free audit · llms.txt · getcitable@gmail.com