AI crawlers on Nginx: allow or block GPTBot, ClaudeBot, PerplexityBot and 78 more
Nginx serves robots.txt as a file from the site’s root, like any other file. The server itself can also refuse a request by its user agent. The steps on Nginx, what trips people up there, and a page for each of 81 AI crawlers.
This response
You are ClaudeBot (Anthropic). Citable recognised you as an AI agent, so this is https://getcitable.in/crawlers/on/nginx with the UI removed — the content only. A browser asking for the same address gets the full designed page.
Where Nginx keeps the rule
Nginx serves robots.txt as a file from the site’s root, like any other file. The server itself can also refuse a request by its user agent.
The steps below are shown for GPTBot. Each crawler has its own page, with its own names in the rule.
Step by step
- 1. Ask, in robots.txt. Add the group for GPTBot to robots.txt in the site’s root.
- 2. Refuse, in the configuration. Put the map in the http block and the if in the server block. The map marks the request; the if returns 403.
- 3. Test, then reload. Run nginx -t, and reload Nginx only if it passes.
What to paste
A group that names a crawler replaces the * group for that crawler; it does not add to it. Paths the * group closes are open to a crawler with its own group unless they are repeated there.
robots.txt, refusing GPTBot:
User-agent: GPTBot
Disallow: /
nginx.conf, refusing the request itself:
# http {}
map $http_user_agent $refuse_crawler {
default 0;
"~*GPTBot" 1;
}
# server {}
if ($refuse_crawler) {
return 403;
}
robots.txt, allowing GPTBot:
User-agent: GPTBot
Allow: /
What trips people up
The map belongs in the http block, not inside server; in the wrong place nginx -t fails and a reload would too. The 403 also covers /robots.txt itself, so the refused crawler cannot read the file either.
Asked, or actually refused?
Yes. The 403 is the server’s own answer.
Check it yourself
A 200 means the name is let through; a 403 means something in front of the page refuses it. This tests the name from your own address. A platform that checks a crawler’s address as well may treat the real GPTBot differently.
The request:
curl -I -A "GPTBot" https://your-site.example/
Every AI crawler, on Nginx
- OAI-SearchBot (OpenAI): https://getcitable.in/crawlers/oai-searchbot/nginx
- ChatGPT-User (OpenAI): https://getcitable.in/crawlers/chatgpt-user/nginx
- GPTBot (OpenAI): https://getcitable.in/crawlers/gptbot/nginx
- Claude-SearchBot (Anthropic): https://getcitable.in/crawlers/claude-searchbot/nginx
- Claude-User (Anthropic): https://getcitable.in/crawlers/claude-user/nginx
- Claude-Web (Anthropic): https://getcitable.in/crawlers/claude-web/nginx
- ClaudeBot (Anthropic): https://getcitable.in/crawlers/claudebot/nginx
- Perplexity-User (Perplexity): https://getcitable.in/crawlers/perplexity-user/nginx
- PerplexityBot (Perplexity): https://getcitable.in/crawlers/perplexitybot/nginx
- Google-Extended (Google): https://getcitable.in/crawlers/google-extended/nginx
- meta-externalfetcher (Meta): https://getcitable.in/crawlers/meta-externalfetcher/nginx
- meta-externalagent (Meta): https://getcitable.in/crawlers/meta-externalagent/nginx
- DuckAssistBot (DuckDuckGo): https://getcitable.in/crawlers/duckassistbot/nginx
- CCBot (Common Crawl): https://getcitable.in/crawlers/ccbot/nginx
- Applebot-Extended (Apple): https://getcitable.in/crawlers/applebot-extended/nginx
- Bytespider (ByteDance): https://getcitable.in/crawlers/bytespider/nginx
- Amazonbot (Amazon): https://getcitable.in/crawlers/amazonbot/nginx
- YouBot (You.com): https://getcitable.in/crawlers/youbot/nginx
- cohere-ai (Cohere): https://getcitable.in/crawlers/cohere-ai/nginx
- PetalBot (Huawei): https://getcitable.in/crawlers/petalbot/nginx
- Diffbot (Diffbot): https://getcitable.in/crawlers/diffbot/nginx
- AI2Bot (AI2): https://getcitable.in/crawlers/ai2bot/nginx
- Timpibot (Timpi): https://getcitable.in/crawlers/timpibot/nginx
- Kangaroo Bot (Kangaroo): https://getcitable.in/crawlers/kangaroo-bot/nginx
- omgili (Webz.io): https://getcitable.in/crawlers/omgili/nginx
- MistralAI-User (Mistral): https://getcitable.in/crawlers/mistralai-user/nginx
- ChatGPT Agent (OpenAI): https://getcitable.in/crawlers/chatgpt-agent/nginx
- Operator (OpenAI): https://getcitable.in/crawlers/operator/nginx
- Gemini-Deep-Research (Google): https://getcitable.in/crawlers/gemini-deep-research/nginx
- Google-Agent (Google): https://getcitable.in/crawlers/google-agent/nginx
- NotebookLM (Google): https://getcitable.in/crawlers/notebooklm/nginx
- Kimi-User (Moonshot AI): https://getcitable.in/crawlers/kimi-user/nginx
- TongyiBot (Alibaba): https://getcitable.in/crawlers/tongyibot/nginx
- YiyanBot (Baidu): https://getcitable.in/crawlers/yiyanbot/nginx
- DoubaoBot (ByteDance): https://getcitable.in/crawlers/doubaobot/nginx
- Amzn-User (Amazon): https://getcitable.in/crawlers/amzn-user/nginx
- AmazonBuyForMe (Amazon): https://getcitable.in/crawlers/amazonbuyforme/nginx
- NovaAct (Amazon): https://getcitable.in/crawlers/novaact/nginx
- amazon-QBusiness (Amazon): https://getcitable.in/crawlers/amazon-qbusiness/nginx
- kagi-fetcher (Kagi): https://getcitable.in/crawlers/kagi-fetcher/nginx
- PhindBot (Phind): https://getcitable.in/crawlers/phindbot/nginx
- LinerBot (Liner): https://getcitable.in/crawlers/linerbot/nginx
- Manus-User (Manus): https://getcitable.in/crawlers/manus-user/nginx
- TwinAgent (Twin): https://getcitable.in/crawlers/twinagent/nginx
- ShapBot (Parallel): https://getcitable.in/crawlers/shapbot/nginx
- bigsur.ai (Big Sur AI): https://getcitable.in/crawlers/bigsur-ai/nginx
- QualifiedBot (Qualified): https://getcitable.in/crawlers/qualifiedbot/nginx
- KlaviyoAIBot (Klaviyo): https://getcitable.in/crawlers/klaviyoaibot/nginx
- Poggio-Citations (Poggio): https://getcitable.in/crawlers/poggio-citations/nginx
- iAskBot (iAsk): https://getcitable.in/crawlers/iaskbot/nginx
- Crawl4AI (Crawl4AI): https://getcitable.in/crawlers/crawl4ai/nginx
- Crawlspace (Crawlspace): https://getcitable.in/crawlers/crawlspace/nginx
- WRTNBot (WRTN): https://getcitable.in/crawlers/wrtnbot/nginx
- UseAI (UseAI): https://getcitable.in/crawlers/useai/nginx
- AIWebIndex (Lyrenth): https://getcitable.in/crawlers/aiwebindex/nginx
- Aranet-SearchBot (Aranet): https://getcitable.in/crawlers/aranet-searchbot/nginx
- KunatoCrawler (Kunato): https://getcitable.in/crawlers/kunatocrawler/nginx
- Reflectionbot (Reflection): https://getcitable.in/crawlers/reflectionbot/nginx
- GeistHaus-PageFetcher (GeistHaus): https://getcitable.in/crawlers/geisthaus-pagefetcher/nginx
- MistralAI-Index (Mistral): https://getcitable.in/crawlers/mistralai-index/nginx
- Kimi-SearchBot (Moonshot AI): https://getcitable.in/crawlers/kimi-searchbot/nginx
- meta-webindexer (Meta): https://getcitable.in/crawlers/meta-webindexer/nginx
- Bravebot (Brave): https://getcitable.in/crawlers/bravebot/nginx
- ExaBot (Exa): https://getcitable.in/crawlers/exabot/nginx
- TavilyBot (Tavily): https://getcitable.in/crawlers/tavilybot/nginx
- AddSearchBot (AddSearch): https://getcitable.in/crawlers/addsearchbot/nginx
- LinkupBot (Linkup): https://getcitable.in/crawlers/linkupbot/nginx
- QueritBot (Querit): https://getcitable.in/crawlers/queritbot/nginx
- TerraCotta (Ceramic AI): https://getcitable.in/crawlers/terracotta/nginx
- HenkBot (Valyu): https://getcitable.in/crawlers/henkbot/nginx
- Channel3Bot (Channel3): https://getcitable.in/crawlers/channel3bot/nginx
- AzureAI-SearchBot (Microsoft): https://getcitable.in/crawlers/azureai-searchbot/nginx
- Amzn-SearchBot (Amazon): https://getcitable.in/crawlers/amzn-searchbot/nginx
- atlassian-bot (Atlassian): https://getcitable.in/crawlers/atlassian-bot/nginx
- Cloudflare-AutoRAG (Cloudflare): https://getcitable.in/crawlers/cloudflare-autorag/nginx
- Andibot (Andi): https://getcitable.in/crawlers/andibot/nginx
- Anomura (Direqt): https://getcitable.in/crawlers/anomura/nginx
- YandexAdditional (Yandex): https://getcitable.in/crawlers/yandexadditional/nginx
- GoogleOther (Google): https://getcitable.in/crawlers/googleother/nginx
- ZanistaBot (Zanista): https://getcitable.in/crawlers/zanistabot/nginx
- Poseidon Research Crawler (Poseidon Research): https://getcitable.in/crawlers/poseidon-research-crawler/nginx
Checked against Nginx's own documentation on 1 October 2026
- nginx docs: ngx_http_map_module: https://nginx.org/en/docs/http/ngx_http_map_module.html
About Citable
Citable strips the UI from a website for AI agents and serves them only the content.
When an AI agent — ChatGPT, Claude, Perplexity or any of the 81 that Citable recognises — opens a page, Citable removes the theme, scripts, navigation and layout and serves only the content: the same facts, at the same URL, on the site's own domain. People and Google still get the full designed page.
How it works: https://getcitable.in/how-it-works · Summary for AI agents: https://getcitable.in/llms.txt
Source: https://getcitable.in/crawlers/on/nginx · Citable · Run the free audit · llms.txt · getcitable@gmail.com