Can ChatGPT read my website? JavaScript, robots.txt, and why it sometimes cannot · Citable

View this page

Whether ChatGPT can read and crawl your website, whether it can read JavaScript, the five reasons it cannot see a page, and how to check exactly what ChatGPT and other AI crawlers are given.

Price
USD 59
Organisation
Citable

Description

written for teams no score out of 100 no ten-step checklist written for teams no score out of 100 no ten-step checklist written for teams no score out of 100 no ten-step checklist

Does ChatGPT crawl websites?

Yes. ChatGPT reads the web through three crawlers. GPTBot collects pages ahead of time. OAI-SearchBot builds the index that ChatGPT search looks things up in. ChatGPT-User opens a page at the moment somebody asks about it or pastes its address into a chat. If any of them has been to your site, its name is in your server logs.

Can ChatGPT read JavaScript?

It can read the JavaScript file as text. It does not run it. That difference is the whole problem. A browser downloads your page, runs the scripts, and the scripts draw the price, the reviews, the stock, the menu. ChatGPT’s crawlers download the page and stop there. Whatever was going to be drawn by a script never appears.

So the honest answer to “can ChatGPT see my website” is: it can see the HTML your server sends. If the page looks complete with JavaScript switched off in your browser, ChatGPT sees it. If it looks empty, so does ChatGPT.

Why can’t ChatGPT read my website? Five reasons

robots.txt refuses it. A Disallow under GPTBot, OAI-SearchBot or ChatGPT-User — or a blanket one under * — tells it to stay out. Many sites have one they never chose.

Something in front of the site blocks it. A CDN or firewall setting that blocks AI bots turns the crawler away before your site is asked. Nothing in robots.txt shows it. See AI crawlers on Cloudflare .

The content is drawn by JavaScript. The crawler got in and was handed a shell. This is the commonest reason, and the hardest to notice, because the page looks fine to you.

The page is too heavy. A typical page is hundreds of kilobytes of markup around a few hundred words of fact. A crawler with a budget reads the first part and leaves — and the first part is the navigation.

It needs a login, a cookie wall or a location. A crawler does not sign in, accept anything or sit in your country. It gets whatever a stranger with no cookies gets.

How to check what ChatGPT can read on your site

Three checks, quickest first.

Run the scan. The free AI crawler checker fetches one of your pages as an AI crawler and shows what came back: how much of it was content, which engines your robots.txt allows, what a crawler is left with.

Look at your robots.txt. Open yoursite.com/robots.txt and search for GPTBot, OAI-SearchBot and ChatGPT-User. The rule for each, and where it goes on your platform, is in the AI crawler list .

Turn JavaScript off. In your browser’s settings, disable JavaScript and reload the page you care about. What is left is roughly what ChatGPT reads.

Is it the same for Claude, Perplexity and Gemini?

Almost. Each has its own crawlers with their own names — they are listed engine by engine under AI engines — and most of them read the same way: the HTML, without running the scripts. Google is the exception for search: Googlebot does run JavaScript. That is why a site can rank well in Google and still be quoted badly, or not at all, by an AI assistant. Passing one test tells you nothing about the other.

ChatGPT is not deciding your page is unimportant. It is reading what it was handed — and on most sites, what it is handed is not the page.

How to make sure it reads all of it

You can rebuild the site so everything is in the HTML: render on the server, move prices and reviews out of widgets, trim the template. That works, and it is a project.

Or you can leave the site as it is. Citable sits in front of it and answers AI crawlers — and only AI crawlers — with the content alone: every fact present as text, at the same address, a fraction of the size. People and Google get the page you designed. Every read is then listed for you, by crawler and by page. See how it works , or what to do next in how to get cited by ChatGPT .

The chain, drawn Access, then comprehension, then the choice.

Access — an engine you refuse never reads anything.

Comprehension — the facts have to survive the markup.

Choice — which stays the model’s, whatever anyone sells you.

Read next The rest of the set.

6 min read AI search optimization

How an assistant decides what to say about your site, and which of those steps a team can actually influence.

7 min read What GEO can and cannot do

What generative engine optimisation means, what it borrowed from SEO, and which parts of it are sold as certainty but aren't.

5 min read How AI crawlers read a website

Rendering, fetch budgets, robots.txt, and why a JavaScript-drawn price is invisible to almost every engine reading you.

8 min read Generative engine optimization (GEO)

What it means, how an AI engine chooses its sources, the strategies with evidence behind them, and what GEO tools actually do.

7 min read Answer engine optimization (AEO)

What it means, three before-and-after examples, the practices that matter — and the step that comes before any of them.

6 min read How to get cited by ChatGPT

How ChatGPT picks the pages it cites, and the six things a site can actually do to show up in its answers.

5 min read llms.txt: what it is, and what it does not do

What goes in the file, a working example, best practices — and whether any AI search engine actually reads it.

5 min read utm_source=chatgpt.com, and tracking AI traffic in GA4

What the tag on your URLs means, why ChatGPT adds it, and how to see every AI engine in Google Analytics.

6 min read Should you block AI crawlers?

Training bots, search bots and live fetch are three different decisions — and the one blocked by accident costs the most.

5 min read GEO vs AEO vs AIO vs LLMO

The full form and meaning of each, where the terms came from, and the four steps underneath all of them.

Or just look at it One page, served both ways.

What survives each trip, with the token count for both.

About Citable Who is behind this page, and what we stand on.

When an AI agent — ChatGPT, Claude, Perplexity or any of the 81 that Citable recognises — opens a page, Citable removes the theme, scripts, navigation and layout and serves only the content: the same facts, at the same URL, on the site's own domain. People and Google still get the full designed page.

Who opens the page What they are given

A person in a browser The full designed page, untouched.

Googlebot, Bingbot and other search engines The full designed page, untouched.

An AI agent — GPTBot, ChatGPT-User, ClaudeBot, PerplexityBot and the rest The content only: no UI, no scripts, the same facts at the same URL.

Measured, not estimated. A real product page of 386.8 KB is served to AI agents as 4.6 KB — 98.8% smaller, with the same facts. Every figure on this site comes from a fetch anybody can repeat.

One list. The 81 AI agents named on these pages are the list the product itself runs on. A page here cannot describe a crawler Citable does not serve.

Search is never touched. Googlebot and Bingbot always get your real page. The clean copy is for AI agents only — at the same address, with the same facts.

Checked, not trusted. A user agent is a claim, and anyone can send one. Where an operator publishes its addresses, every read is checked against them and recorded as verified or not.

Nothing about your visitors. Citable sees which crawler read which page. It does not see who your customers are, and their details never reach us.

Honest about limits. Nobody can promise a citation or a ranking, and we do not. What we show is what an AI agent was given, and whether it sent somebody back.

How it works · The measurements · What it touches and stores · Why Citable · Pricing · Questions, answered plainly · Talk to us

Free · about 20 seconds · no signup Find out whether AI has ever read your site.

Type your address. We fetch a page exactly as GPTBot does, count the tokens, and score fifteen things an answer engine needs. If the log then stays empty for a month, that is worth knowing too.

No contract · cancel any time · Googlebot untouched