AI crawlers on Apache: allow or block GPTBot, ClaudeBot, PerplexityBot and 78 more
Apache serves robots.txt as a file from the document root. A rewrite rule, in .htaccess or the virtual host, can refuse a request by its user agent. The steps on Apache, what trips people up there, and a page for each of 81 AI crawlers.
This response
You are ClaudeBot (Anthropic). Citable recognised you as an AI agent, so this is https://getcitable.in/crawlers/on/apache with the UI removed — the content only. A browser asking for the same address gets the full designed page.
Where Apache keeps the rule
Apache serves robots.txt as a file from the document root. A rewrite rule, in .htaccess or the virtual host, can refuse a request by its user agent.
The steps below are shown for GPTBot. Each crawler has its own page, with its own names in the rule.
Step by step
- 1. Ask, in robots.txt. Add the group for GPTBot to robots.txt in the document root.
- 2. Refuse, in .htaccess. Add the rewrite rule below to .htaccess in the document root, or to the virtual host. It answers a matching user agent with 403 Forbidden.
- 3. Check it took. Request the site with the crawler’s name, as below. .htaccess is read on every request, so there is nothing to restart.
What to paste
A group that names a crawler replaces the * group for that crawler; it does not add to it. Paths the * group closes are open to a crawler with its own group unless they are repeated there.
robots.txt, refusing GPTBot:
User-agent: GPTBot
Disallow: /
.htaccess, refusing the request itself:
<IfModule mod_rewrite.c>
RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} "GPTBot" [NC]
RewriteRule ^ - [F]
</IfModule>
robots.txt, allowing GPTBot:
User-agent: GPTBot
Allow: /
What trips people up
.htaccess is only read where the server’s AllowOverride permits it, and the rule needs mod_rewrite. Inside the IfModule block a missing module fails silently: the rule is simply not applied.
Asked, or actually refused?
Yes. The [F] flag makes Apache answer 403 Forbidden itself.
Check it yourself
A 200 means the name is let through; a 403 means something in front of the page refuses it. This tests the name from your own address. A platform that checks a crawler’s address as well may treat the real GPTBot differently.
The request:
curl -I -A "GPTBot" https://your-site.example/
Every AI crawler, on Apache
- OAI-SearchBot (OpenAI): https://getcitable.in/crawlers/oai-searchbot/apache
- ChatGPT-User (OpenAI): https://getcitable.in/crawlers/chatgpt-user/apache
- GPTBot (OpenAI): https://getcitable.in/crawlers/gptbot/apache
- Claude-SearchBot (Anthropic): https://getcitable.in/crawlers/claude-searchbot/apache
- Claude-User (Anthropic): https://getcitable.in/crawlers/claude-user/apache
- Claude-Web (Anthropic): https://getcitable.in/crawlers/claude-web/apache
- ClaudeBot (Anthropic): https://getcitable.in/crawlers/claudebot/apache
- Perplexity-User (Perplexity): https://getcitable.in/crawlers/perplexity-user/apache
- PerplexityBot (Perplexity): https://getcitable.in/crawlers/perplexitybot/apache
- Google-Extended (Google): https://getcitable.in/crawlers/google-extended/apache
- meta-externalfetcher (Meta): https://getcitable.in/crawlers/meta-externalfetcher/apache
- meta-externalagent (Meta): https://getcitable.in/crawlers/meta-externalagent/apache
- DuckAssistBot (DuckDuckGo): https://getcitable.in/crawlers/duckassistbot/apache
- CCBot (Common Crawl): https://getcitable.in/crawlers/ccbot/apache
- Applebot-Extended (Apple): https://getcitable.in/crawlers/applebot-extended/apache
- Bytespider (ByteDance): https://getcitable.in/crawlers/bytespider/apache
- Amazonbot (Amazon): https://getcitable.in/crawlers/amazonbot/apache
- YouBot (You.com): https://getcitable.in/crawlers/youbot/apache
- cohere-ai (Cohere): https://getcitable.in/crawlers/cohere-ai/apache
- PetalBot (Huawei): https://getcitable.in/crawlers/petalbot/apache
- Diffbot (Diffbot): https://getcitable.in/crawlers/diffbot/apache
- AI2Bot (AI2): https://getcitable.in/crawlers/ai2bot/apache
- Timpibot (Timpi): https://getcitable.in/crawlers/timpibot/apache
- Kangaroo Bot (Kangaroo): https://getcitable.in/crawlers/kangaroo-bot/apache
- omgili (Webz.io): https://getcitable.in/crawlers/omgili/apache
- MistralAI-User (Mistral): https://getcitable.in/crawlers/mistralai-user/apache
- ChatGPT Agent (OpenAI): https://getcitable.in/crawlers/chatgpt-agent/apache
- Operator (OpenAI): https://getcitable.in/crawlers/operator/apache
- Gemini-Deep-Research (Google): https://getcitable.in/crawlers/gemini-deep-research/apache
- Google-Agent (Google): https://getcitable.in/crawlers/google-agent/apache
- NotebookLM (Google): https://getcitable.in/crawlers/notebooklm/apache
- Kimi-User (Moonshot AI): https://getcitable.in/crawlers/kimi-user/apache
- TongyiBot (Alibaba): https://getcitable.in/crawlers/tongyibot/apache
- YiyanBot (Baidu): https://getcitable.in/crawlers/yiyanbot/apache
- DoubaoBot (ByteDance): https://getcitable.in/crawlers/doubaobot/apache
- Amzn-User (Amazon): https://getcitable.in/crawlers/amzn-user/apache
- AmazonBuyForMe (Amazon): https://getcitable.in/crawlers/amazonbuyforme/apache
- NovaAct (Amazon): https://getcitable.in/crawlers/novaact/apache
- amazon-QBusiness (Amazon): https://getcitable.in/crawlers/amazon-qbusiness/apache
- kagi-fetcher (Kagi): https://getcitable.in/crawlers/kagi-fetcher/apache
- PhindBot (Phind): https://getcitable.in/crawlers/phindbot/apache
- LinerBot (Liner): https://getcitable.in/crawlers/linerbot/apache
- Manus-User (Manus): https://getcitable.in/crawlers/manus-user/apache
- TwinAgent (Twin): https://getcitable.in/crawlers/twinagent/apache
- ShapBot (Parallel): https://getcitable.in/crawlers/shapbot/apache
- bigsur.ai (Big Sur AI): https://getcitable.in/crawlers/bigsur-ai/apache
- QualifiedBot (Qualified): https://getcitable.in/crawlers/qualifiedbot/apache
- KlaviyoAIBot (Klaviyo): https://getcitable.in/crawlers/klaviyoaibot/apache
- Poggio-Citations (Poggio): https://getcitable.in/crawlers/poggio-citations/apache
- iAskBot (iAsk): https://getcitable.in/crawlers/iaskbot/apache
- Crawl4AI (Crawl4AI): https://getcitable.in/crawlers/crawl4ai/apache
- Crawlspace (Crawlspace): https://getcitable.in/crawlers/crawlspace/apache
- WRTNBot (WRTN): https://getcitable.in/crawlers/wrtnbot/apache
- UseAI (UseAI): https://getcitable.in/crawlers/useai/apache
- AIWebIndex (Lyrenth): https://getcitable.in/crawlers/aiwebindex/apache
- Aranet-SearchBot (Aranet): https://getcitable.in/crawlers/aranet-searchbot/apache
- KunatoCrawler (Kunato): https://getcitable.in/crawlers/kunatocrawler/apache
- Reflectionbot (Reflection): https://getcitable.in/crawlers/reflectionbot/apache
- GeistHaus-PageFetcher (GeistHaus): https://getcitable.in/crawlers/geisthaus-pagefetcher/apache
- MistralAI-Index (Mistral): https://getcitable.in/crawlers/mistralai-index/apache
- Kimi-SearchBot (Moonshot AI): https://getcitable.in/crawlers/kimi-searchbot/apache
- meta-webindexer (Meta): https://getcitable.in/crawlers/meta-webindexer/apache
- Bravebot (Brave): https://getcitable.in/crawlers/bravebot/apache
- ExaBot (Exa): https://getcitable.in/crawlers/exabot/apache
- TavilyBot (Tavily): https://getcitable.in/crawlers/tavilybot/apache
- AddSearchBot (AddSearch): https://getcitable.in/crawlers/addsearchbot/apache
- LinkupBot (Linkup): https://getcitable.in/crawlers/linkupbot/apache
- QueritBot (Querit): https://getcitable.in/crawlers/queritbot/apache
- TerraCotta (Ceramic AI): https://getcitable.in/crawlers/terracotta/apache
- HenkBot (Valyu): https://getcitable.in/crawlers/henkbot/apache
- Channel3Bot (Channel3): https://getcitable.in/crawlers/channel3bot/apache
- AzureAI-SearchBot (Microsoft): https://getcitable.in/crawlers/azureai-searchbot/apache
- Amzn-SearchBot (Amazon): https://getcitable.in/crawlers/amzn-searchbot/apache
- atlassian-bot (Atlassian): https://getcitable.in/crawlers/atlassian-bot/apache
- Cloudflare-AutoRAG (Cloudflare): https://getcitable.in/crawlers/cloudflare-autorag/apache
- Andibot (Andi): https://getcitable.in/crawlers/andibot/apache
- Anomura (Direqt): https://getcitable.in/crawlers/anomura/apache
- YandexAdditional (Yandex): https://getcitable.in/crawlers/yandexadditional/apache
- GoogleOther (Google): https://getcitable.in/crawlers/googleother/apache
- ZanistaBot (Zanista): https://getcitable.in/crawlers/zanistabot/apache
- Poseidon Research Crawler (Poseidon Research): https://getcitable.in/crawlers/poseidon-research-crawler/apache
Checked against Apache's own documentation on 1 October 2026
- Apache docs: mod_rewrite: https://httpd.apache.org/docs/current/mod/mod_rewrite.html
About Citable
Citable strips the UI from a website for AI agents and serves them only the content.
When an AI agent — ChatGPT, Claude, Perplexity or any of the 81 that Citable recognises — opens a page, Citable removes the theme, scripts, navigation and layout and serves only the content: the same facts, at the same URL, on the site's own domain. People and Google still get the full designed page.
How it works: https://getcitable.in/how-it-works · Summary for AI agents: https://getcitable.in/llms.txt
Source: https://getcitable.in/crawlers/on/apache · Citable · Run the free audit · llms.txt · getcitable@gmail.com