Applebot-Extended: what it is, who runs it, and how to allow or block it

Applebot-Extended is not a crawler. Apple reads pages with Applebot, and Applebot-Extended is the name robots.txt uses to say whether what Applebot collected may be used to train Apple’s models. Its name in robots.txt, how it shows up in server logs, whether its identity can be checked, and how to allow or refuse it on each platform.

This response

You are ClaudeBot (Anthropic). Citable recognised you as an AI agent, so this is https://getcitable.in/crawlers/applebot-extended with the UI removed — the content only. A browser asking for the same address gets the full designed page.

What Applebot-Extended is

Applebot-Extended is not a crawler. Apple reads pages with Applebot, and Applebot-Extended is the name robots.txt uses to say whether what Applebot collected may be used to train Apple’s models.

Should you refuse it?

Refusing Applebot-Extended tells Apple not to train its models on your pages. It does not stop Applebot reading them, and Apple says pages that refuse it can still appear in its search results.

robots.txt is a request. A crawler that honours it stays out; nothing in the file stops one that does not. What is enforced is what the server or the edge in front of it refuses.

Why there is nothing to verify

There is nothing to verify: no request carries the name Applebot-Extended.

Allowing or refusing Applebot-Extended in robots.txt

Add a group under its name, “User-agent: Applebot-Extended”. “Allow: /” lets it read the site and “Disallow: /” refuses it.

A group that names a crawler replaces the * group for that crawler; it does not add to it. Paths the * group closes are open to a crawler with its own group unless they are repeated there.

Refuse Applebot-Extended:

User-agent: Applebot-Extended
Disallow: /

Allow Applebot-Extended:

User-agent: Applebot-Extended
Allow: /

Where the rule for Applebot-Extended goes, platform by platform

Questions

What is Applebot-Extended?

Applebot-Extended is not a crawler. Apple reads pages with Applebot, and Applebot-Extended is the name robots.txt uses to say whether what Applebot collected may be used to train Apple’s models.

Who runs Applebot-Extended?

Apple. It is the only name Citable recognises for Apple.

How do I allow or block Applebot-Extended in robots.txt?

Add a group under its name, “User-agent: Applebot-Extended”. “Allow: /” lets it read the site and “Disallow: /” refuses it. A group that names a crawler replaces the * group for that crawler; it does not add to it. Paths the * group closes are open to a crawler with its own group unless they are repeated there.

Should I block Applebot-Extended?

Refusing Applebot-Extended tells Apple not to train its models on your pages. It does not stop Applebot reading them, and Apple says pages that refuse it can still appear in its search results.

How do I see Applebot-Extended in my server logs?

Applebot-Extended never appears in an access log: no request carries the name. The pages are fetched by Applebot.

About Citable

Citable strips the UI from a website for AI agents and serves them only the content.

When an AI agent — ChatGPT, Claude, Perplexity or any of the 81 that Citable recognises — opens a page, Citable removes the theme, scripts, navigation and layout and serves only the content: the same facts, at the same URL, on the site's own domain. People and Google still get the full designed page.

How it works: https://getcitable.in/how-it-works · Summary for AI agents: https://getcitable.in/llms.txt


Source: https://getcitable.in/crawlers/applebot-extended · Citable · Run the free audit · llms.txt · getcitable@gmail.com