Your CDN may be blocking AI crawlers — how to allow GPTBot on Cloudflare, Akamai, AWS and more · Citable

View this page

Why ChatGPT, Perplexity and Claude cannot read a site behind bot protection, how to tell which CDN is refusing them, and the one setting on each — Cloudflare, Akamai, AWS WAF, Fastly, Imperva, DataDome — that lets AI crawlers through.

Price
USD 59
Organisation
Citable

Description

written for teams no score out of 100 no ten-step checklist written for teams no score out of 100 no ten-step checklist written for teams no score out of 100 no ten-step checklist

What it looks like

We fetched popular Indian commerce and media sites as GPTBot and as Chrome. Most refusals fell into three shapes, and the free scan now names each one:

What GPTBot got What it means

HTTP 403 or 999 , while a browser got the page The edge refuses AI crawlers by name. Answer engines cannot read the site.

HTTP 403 for every automated request Bot protection refuses anything that is not a browser. AI crawlers are almost certainly refused too.

No answer at all The connection is held open and nothing is sent. Akamai and Imperva do this to automated requests.

robots.txt does not decide any of this. It is a request a crawler reads; the edge answers before the crawler gets that far.

Which CDN is refusing them

Every edge signs its responses — cf-ray for Cloudflare, akamai-grn or an AkamaiGHost server for Akamai, x-amz-cf-id for CloudFront. When a site sends nothing back, its DNS still names the edge: a CNAME ending in edgekey.net is Akamai. The free scan reads both and tells you.

How to let AI crawlers through, per CDN

CDN or bot service The setting

Cloudflare Security → Bots: allow verified AI crawlers and turn off “Block AI bots”.

Amazon CloudFront AWS WAF Bot Control: allow the AI-crawler category (CategoryAI).

Fastly Fastly Bot Management: allow verified AI crawlers.

Akamai Bot Manager: set the AI-crawler category to Allow.

Azure Front Door WAF policy: exclude the AI crawlers' user-agents from the bot-protection rule set.

Sucuri Firewall → Access control: allow-list the AI crawlers' user-agents.

Imperva Bot protection: add the AI crawlers to the allow-list.

DataDome Bot policy: allow AI agents and LLM crawlers.

HUMAN (PerimeterX) Bot Defender policy: allow the AI crawlers.

Kasada Bot policy: allow the AI crawlers.

Allow verified crawlers where the CDN offers it: the major engines publish their IP ranges, and a verified-bot list lets GPTBot in without letting in everything that claims to be GPTBot.

Then check it

Run the free scan again. It fetches as GPTBot, so a site that is still refusing it says so in the first line — and a site that lets it in shows what it is now served.

The chain, drawn Access, then comprehension, then the choice.

Access — an engine you refuse never reads anything.

Comprehension — the facts have to survive the markup.

Choice — which stays the model’s, whatever anyone sells you.

Read next The rest of the set.

6 min read AI search optimization

How an assistant decides what to say about your site, and which of those steps a team can actually influence.

7 min read What GEO can and cannot do

What generative engine optimisation means, what it borrowed from SEO, and which parts of it are sold as certainty but aren't.

5 min read How AI crawlers read a website

Rendering, fetch budgets, robots.txt, and why a JavaScript-drawn price is invisible to almost every engine reading you.

8 min read Generative engine optimization (GEO)

What it means, how an AI engine chooses its sources, the strategies with evidence behind them, and what GEO tools actually do.

7 min read Answer engine optimization (AEO)

What it means, three before-and-after examples, the practices that matter — and the step that comes before any of them.

6 min read How to get cited by ChatGPT

How ChatGPT picks the pages it cites, and the six things a site can actually do to show up in its answers.

5 min read Can ChatGPT read my website?

How it reads a site, why it does not run JavaScript, the five things that stop it, and how to check in a minute.

5 min read llms.txt: what it is, and what it does not do

What goes in the file, a working example, best practices — and whether any AI search engine actually reads it.

5 min read utm_source=chatgpt.com, and tracking AI traffic in GA4

What the tag on your URLs means, why ChatGPT adds it, and how to see every AI engine in Google Analytics.

6 min read Should you block AI crawlers?

Training bots, search bots and live fetch are three different decisions — and the one blocked by accident costs the most.

5 min read GEO vs AEO vs AIO vs LLMO

The full form and meaning of each, where the terms came from, and the four steps underneath all of them.

Or just look at it One page, served both ways.

What survives each trip, with the token count for both.

About Citable Who is behind this page, and what we stand on.

When an AI agent — ChatGPT, Claude, Perplexity or any of the 81 that Citable recognises — opens a page, Citable removes the theme, scripts, navigation and layout and serves only the content: the same facts, at the same URL, on the site's own domain. People and Google still get the full designed page.

Who opens the page What they are given

A person in a browser The full designed page, untouched.

Googlebot, Bingbot and other search engines The full designed page, untouched.

An AI agent — GPTBot, ChatGPT-User, ClaudeBot, PerplexityBot and the rest The content only: no UI, no scripts, the same facts at the same URL.

Measured, not estimated. A real product page of 386.8 KB is served to AI agents as 4.6 KB — 98.8% smaller, with the same facts. Every figure on this site comes from a fetch anybody can repeat.

One list. The 81 AI agents named on these pages are the list the product itself runs on. A page here cannot describe a crawler Citable does not serve.

Search is never touched. Googlebot and Bingbot always get your real page. The clean copy is for AI agents only — at the same address, with the same facts.

Checked, not trusted. A user agent is a claim, and anyone can send one. Where an operator publishes its addresses, every read is checked against them and recorded as verified or not.

Nothing about your visitors. Citable sees which crawler read which page. It does not see who your customers are, and their details never reach us.

Honest about limits. Nobody can promise a citation or a ranking, and we do not. What we show is what an AI agent was given, and whether it sent somebody back.

How it works · The measurements · What it touches and stores · Why Citable · Pricing · Questions, answered plainly · Talk to us

Free · about 20 seconds · no signup AI picks you or a competitor. Find out why.

Type your address. We fetch a page exactly as ChatGPT's crawler does and show you what it was handed, what it missed, and what to fix.

No contract · cancel any time · Googlebot untouched