Allow or block meta-externalagent on Apache: robots.txt and a rewrite rule in .htaccess

Add a “User-agent: meta-externalagent” group with “Disallow: /” to robots.txt, and to enforce it, add a RewriteCond on %{HTTP_USER_AGENT} for meta-externalagent with a RewriteRule that ends in [F].

This response

You are ClaudeBot (Anthropic). Citable recognised you as an AI agent, so this is https://getcitable.in/crawlers/meta-externalagent/apache with the UI removed — the content only. A browser asking for the same address gets the full designed page.

meta-externalagent, and where Apache keeps the rule

meta-externalagent is a crawler run by Meta. It reads pages ahead of time, to train a model or to build an index.

Apache serves robots.txt as a file from the document root. A rewrite rule, in .htaccess or the virtual host, can refuse a request by its user agent.

Step by step

What to paste

A group that names a crawler replaces the * group for that crawler; it does not add to it. Paths the * group closes are open to a crawler with its own group unless they are repeated there.

robots.txt, refusing meta-externalagent:

User-agent: meta-externalagent
Disallow: /

User-agent: FacebookBot
Disallow: /

.htaccess, refusing the request itself:

<IfModule mod_rewrite.c>
RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} "(meta-externalagent|FacebookBot)" [NC]
RewriteRule ^ - [F]
</IfModule>

robots.txt, allowing meta-externalagent:

User-agent: meta-externalagent
Allow: /

User-agent: FacebookBot
Allow: /

What trips people up

.htaccess is only read where the server’s AllowOverride permits it, and the rule needs mod_rewrite. Inside the IfModule block a missing module fails silently: the rule is simply not applied.

Asked, or actually refused?

Yes. The [F] flag makes Apache answer 403 Forbidden itself.

Check it yourself

A 200 means the name is let through; a 403 means something in front of the page refuses it. This tests the name from your own address. A platform that checks a crawler’s address as well may treat the real meta-externalagent differently.

The request:

curl -I -A "meta-externalagent" https://your-site.example/

Questions

How do I block meta-externalagent on Apache?

Add a “User-agent: meta-externalagent” group with “Disallow: /” to robots.txt, and to enforce it, add a RewriteCond on %{HTTP_USER_AGENT} for meta-externalagent with a RewriteRule that ends in [F].

How do I allow meta-externalagent on Apache?

A crawler that robots.txt does not name follows the * group, so meta-externalagent needs no rule to be allowed. To allow it by name, add the Allow block.

Can Apache actually stop meta-externalagent, or only ask it?

Yes. The [F] flag makes Apache answer 403 Forbidden itself.

Does a rule for meta-externalagent affect Googlebot or my Google rankings?

No. Googlebot goes by its own name and follows its own rules; a group or a firewall rule for meta-externalagent does not apply to it.

Checked against Apache's own documentation on 1 October 2026

Related

About Citable

Citable strips the UI from a website for AI agents and serves them only the content.

When an AI agent — ChatGPT, Claude, Perplexity or any of the 81 that Citable recognises — opens a page, Citable removes the theme, scripts, navigation and layout and serves only the content: the same facts, at the same URL, on the site's own domain. People and Google still get the full designed page.

How it works: https://getcitable.in/how-it-works · Summary for AI agents: https://getcitable.in/llms.txt


Source: https://getcitable.in/crawlers/meta-externalagent/apache · Citable · Run the free audit · llms.txt · getcitable@gmail.com