ExplainerTechnical AEO
robots.txt for AI Crawlers: GPTBot, ClaudeBot, PerplexityBot and Friends
by the AEO GEO Labs team6 min read
Your robots.txt controls AI crawlers by user-agent name, and the key fact is that each AI company runs separate bots for training and for search. You can block GPTBot (OpenAI training) while allowing OAI-SearchBot (ChatGPT search), so your content stays out of model training but can still be cited. Blocking the wrong bot removes you from AI answers.
tl;dr
- 1Training bots and search bots have different names. Block by purpose, not by company.
- 2User-triggered fetchers such as ChatGPT-User and Perplexity-User may not follow robots.txt.
- 3Google-Extended and Applebot-Extended are opt-out tokens, not crawlers, and don't affect search ranking.
- 4Most AI crawlers don't run JavaScript, so allowing them is only half the job.
robots.txt is one of two plain-text files that shape how AI systems see your site; the other is covered in our guide to llms.txt. This page lists the bots, what each does according to its operator's own documentation, and the rules to use.
Which AI crawlers exist and what each one does
AI crawlers fall into three groups: training crawlers that collect data for models, search crawlers that index pages for AI answers, and user-triggered fetchers that load a page when someone asks a question. As of October 2026, here is what each operator says:
| User-agent | Operator | Purpose | Follows robots.txt? |
|---|---|---|---|
| GPTBot | OpenAI | Model training | Yes |
| OAI-SearchBot | OpenAI | ChatGPT search results | Yes |
| ChatGPT-User | OpenAI | User actions in ChatGPT and GPTs | "robots.txt rules may not apply" |
| ClaudeBot | Anthropic | Model training | Yes, plus Crawl-delay |
| Claude-SearchBot | Anthropic | Search result quality | Yes |
| Claude-User | Anthropic | Fetches when a user asks Claude | Yes |
| PerplexityBot | Perplexity | Perplexity search results, not training | Yes |
| Perplexity-User | Perplexity | Fetches when a user asks | "generally ignores robots.txt" |
| Google-Extended | Token: Gemini training and grounding | Token only, no separate crawler | |
| Applebot-Extended | Apple | Token: Apple model training | Token only, no separate crawler |
| Meta-ExternalAgent | Meta | Training and product improvement | Yes |
| Meta-ExternalFetcher | Meta | User-requested fetches | "may bypass robots.txt" |
| CCBot | Common Crawl | Open web archive used by many AI labs | Yes |
Sources: OpenAI's crawler docs, Anthropic's crawler help article, Perplexity's bot docs, Google's common crawlers list, Apple's Applebot page and Meta's web crawler page.
Block training, allow AI search: the rules
To block AI training while staying visible in AI search, disallow the training bots by name and leave the search bots allowed. This is the setup most businesses want:
# Block model training
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: Meta-ExternalAgent
Disallow: /
User-agent: CCBot
Disallow: /
# Allow AI search
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: *
Allow: /
Sitemap: https://example.com/sitemap.xml
A crawler follows the most specific group that matches its name, under the Robots Exclusion Protocol, standardized as RFC 9309. So GPTBot reads its own group and ignores the * group. The explicit Allow lines are not strictly needed, but they make your intent clear to the next person who edits the file.
Our view: unless you sell content itself (news, data, courses), blocking training buys little. Models already learn about your brand from the rest of the web. Blocking search bots, on the other hand, costs citations immediately.
Allow all AI crawlers
To allow all AI crawlers, you don't need any AI-specific lines. A robots.txt with User-agent: * and Allow: / (or an empty Disallow:) admits every bot that follows the protocol. Check for leftovers: many sites added blanket GPTBot blocks in 2023 and forgot about them.
Google AI Overviews and Google-Extended
Google-Extended does not control AI Overviews. Google's crawler documentation says Google-Extended manages whether content may be used for training future Gemini models and for grounding in Gemini Apps and Vertex AI, and that it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal." As of October 2026, AI Overviews and AI Mode use normal Googlebot crawling. Google's AI features guidance says to limit what appears there with nosnippet, data-nosnippet, max-snippet or noindex, not with robots.txt tokens.
Applebot-Extended works the same way for Apple: Apple says pages that disallow it can still appear in search results.
What robots.txt cannot do
robots.txt is a request, not a lock. Three limits matter:
- User-triggered fetchers may ignore it. OpenAI and Perplexity both say their user-initiated agents may not apply robots.txt rules, because a person asked for the page.
- Not every crawler declares itself. A bot can send any user-agent string it likes, and robots.txt only binds bots that choose to obey it. If you need enforcement, use firewall rules and verify bots against the IP lists OpenAI, Anthropic and Common Crawl publish.
- Changes take time. OpenAI says its systems take about 24 hours to adjust after a robots.txt update, and Meta says up to 24 hours.
There is a fourth trap that isn't a robots.txt problem at all. Vercel's December 2024 analysis of AI crawler traffic found that none of the major AI crawlers rendered JavaScript, so a page that builds its content in the browser can be allowed and still look empty to them. Check yours with What AI Sees.
How to test AI crawlers against your robots.txt
The fastest test is to run your URL through the AI crawler checker. It checks 21 crawlers from OpenAI, Anthropic, Perplexity, Google and others against your robots.txt, shows the exact line that decides each one, and requests the page as several of them to catch firewall blocks that robots.txt can't show. After crawlers can get in, give them clean signals about what your pages are with schema markup.
Frequently asked questions
Does blocking GPTBot remove my site from ChatGPT search?
No. As of October 2026, OpenAI uses GPTBot for model training and OAI-SearchBot for ChatGPT's search features. Blocking GPTBot tells OpenAI not to use your content for training. To stay eligible for citations in ChatGPT search, keep OAI-SearchBot allowed. Blocking both removes you from training and from ChatGPT search results.
Should I block AI crawlers in robots.txt?
Block training crawlers only if your content is your product, such as news, research or paid courses. Most businesses gain more from being cited than they lose from being trained on. Never block search crawlers like OAI-SearchBot, Claude-SearchBot or PerplexityBot unless you want to disappear from those AI answers. Decide bot by bot, not company by company.
Does Google-Extended affect AI Overviews?
No. Google's documentation says Google-Extended controls whether content is used for training future Gemini models and for grounding, and that it does not affect inclusion or ranking in Google Search. As of October 2026, AI Overviews draw on pages crawled by Googlebot. To limit how your content appears in them, use nosnippet, max-snippet or noindex.
Do AI crawlers respect robots.txt?
The major declared crawlers say they do: GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Meta-ExternalAgent and CCBot all document robots.txt support. User-triggered fetchers are different. OpenAI says robots.txt "may not apply" to ChatGPT-User, and Perplexity says Perplexity-User "generally ignores" it. For hard blocking, use firewall rules with published IP ranges.
How long do robots.txt changes take to apply to AI crawlers?
Expect about a day. OpenAI says it can take roughly 24 hours for its systems to adjust after you update robots.txt, and Meta says its crawlers cache the file for up to 24 hours. Other operators don't publish a figure. Test the live file with a crawler checker right after the change, then again the next day.
If you want this done for your site, AEO GEO Labs audits crawler access, robots.txt and rendering as part of its GEO, AEO and SEO programs for B2B and SaaS teams. See our services.