What Is Applebot-Extended?
Applebot-Extended is a robots.txt user agent that Apple provides so publishers can opt out of having their content used to train Apple's generative AI models. Apple describes it as a secondary user agent that lets publishers opt out of content "being used to train Apple's general purpose foundation models powering generative AI features across Apple products, including Apple Intelligence, Services, and Developer Tools." It is a permission signal, not a crawler. Apple states plainly: "Applebot-Extended does not crawl webpages."
How It Works
Apple has one crawler, Applebot, which fetches pages for Spotlight, Siri, and Safari search features. Applebot-Extended only tells Apple what it may do with pages Applebot has already fetched. You will not see Applebot-Extended in your server logs, because nothing makes requests under that name.
| Applebot | Applebot-Extended | |
|---|---|---|
| Fetches pages | Yes | No |
| Controls | Whether content appears in Spotlight, Siri, and Safari results | Whether crawled content may be used to train Apple's generative models |
| Effect of a Disallow | Content is removed from Apple search features | Content is excluded from training. Search features are unaffected |
| Appears in logs | Yes, with Applebot/0.1 in the user agent string | No |
The two tokens are controlled separately. Apple notes that "webpages that disallow Applebot-Extended can still be included in search results."
robots.txt Examples
To keep Apple search features and opt out of training:
User-agent: Applebot
Allow: /
User-agent: Applebot-Extended
Disallow: /
To verify that a request claiming to be Applebot is real, Apple says to use a reverse DNS lookup, which should resolve to a host under applebot.apple.com, or match the address against Apple's published Applebot IP list.
Why It Matters for AI Visibility
Blocking Applebot-Extended is a trade. It keeps your content out of Apple's model training, which protects content you want to license. It also means Apple's models learn less about your brand from your own pages, and what they know comes from third parties. For most brands that want to be described accurately by Apple Intelligence and Siri, allowing both tokens is the default. Publishers whose content is the product have a real reason to block. Siri AI on iOS 27 is built on Apple Foundation Models developed with Google Gemini, so Apple's own crawl is only part of what shapes its answers: see Siri and Gemini on iOS 27 and Apple Intelligence citation patterns.
Commonly Confused With
Applebot-Extended works like Google-Extended: both are permission tokens with no crawler behind them. They differ from GPTBot and ClaudeBot, which are real crawlers that fetch pages and show up in logs.