Allow or block Google-Extended on WordPress: the virtual robots.txt, the file, and the filter

Put a “User-agent: Google-Extended” group with “Disallow: /” in robots.txt: in the file in the site’s root if there is one, otherwise through an SEO plugin’s editor (Yoast SEO → Tools → File Editor) or the robots_txt filter.

This response

You are ClaudeBot (Anthropic). Citable recognised you as an AI agent, so this is https://getcitable.in/crawlers/google-extended/wordpress with the UI removed — the content only. A browser asking for the same address gets the full designed page.

Google-Extended, and where WordPress keeps the rule

Google-Extended is not a crawler. It sends no requests of its own: Google reads pages with its existing crawlers, and Google-Extended is the name robots.txt uses to say whether those pages may be used to train Gemini models.

WordPress serves a virtual robots.txt when there is no robots.txt file in the site’s root. Once a file exists the web server answers with it, and WordPress is never asked.

Step by step

What to paste

A group that names a crawler replaces the * group for that crawler; it does not add to it. Paths the * group closes are open to a crawler with its own group unless they are repeated there.

Refuse Google-Extended:

User-agent: Google-Extended
Disallow: /

Allow Google-Extended:

User-agent: Google-Extended
Allow: /

The same refusal, through WordPress’s robots_txt filter:

add_filter( 'robots_txt', function ( $output, $public ) {
    $output .= "\nUser-agent: Google-Extended\nDisallow: /\n";
    return $output;
}, 10, 2 );

What trips people up

A robots.txt file in the root wins. While it exists, a plugin’s editor and the robots_txt filter change nothing a crawler will see. Yoast’s File Editor is also absent when the install has file editing disabled.

Asked, or actually refused?

No request carries the name Google-Extended, so a rule on the user agent never matches it. robots.txt is the only place this name has any effect.

Check it yourself

Read the file a crawler is given and look for the group that names Google-Extended. There is no request to test: Google-Extended never sends one.

The request:

curl -s https://your-site.example/robots.txt

Questions

How do I block Google-Extended on WordPress?

Put a “User-agent: Google-Extended” group with “Disallow: /” in robots.txt: in the file in the site’s root if there is one, otherwise through an SEO plugin’s editor (Yoast SEO → Tools → File Editor) or the robots_txt filter.

How do I allow Google-Extended on WordPress?

A crawler that robots.txt does not name follows the * group, so Google-Extended needs no rule to be allowed. To allow it by name, add the Allow block.

Can WordPress actually stop Google-Extended, or only ask it?

No request carries the name Google-Extended, so a rule on the user agent never matches it. robots.txt is the only place this name has any effect.

Does a rule for Google-Extended affect Googlebot or my Google rankings?

Refusing Google-Extended tells Google not to use your pages to train Gemini models. It does not change whether Googlebot reads them, and Google says it is not used as a ranking signal and does not affect inclusion in Google Search.

Checked against WordPress's own documentation on 1 October 2026

Related

About Citable

Citable strips the UI from a website for AI agents and serves them only the content.

When an AI agent — ChatGPT, Claude, Perplexity or any of the 81 that Citable recognises — opens a page, Citable removes the theme, scripts, navigation and layout and serves only the content: the same facts, at the same URL, on the site's own domain. People and Google still get the full designed page.

How it works: https://getcitable.in/how-it-works · Summary for AI agents: https://getcitable.in/llms.txt


Source: https://getcitable.in/crawlers/google-extended/wordpress · Citable · Run the free audit · llms.txt · getcitable@gmail.com