Do AI Chatbots Respect Robots.txt and Crawling Rules
Patrick Bushe
July 1, 2024 · 5 min read
Robots.txt is a plain text file where website owners tell automated crawlers which pages they may visit. It was standardized as RFC 9309 in 2022, but it has always been voluntary. Most major AI companies say their crawlers follow it; the details matter.
The main AI crawlers
- OpenAI:
GPTBotcollects training data,OAI-SearchBotindexes for ChatGPT search, andChatGPT-Uservisits pages when a user asks. - Google:
Google-Extendedis a robots.txt token controlling whether content is used for Gemini training. It doesn't affect Google Search or AI Overviews, which use regular Googlebot. - Anthropic:
ClaudeBotfor training, plus separate agents for search and user requests. - Apple:
Applebot-Extendedcontrols use for Apple's AI training. - Perplexity:
PerplexityBotfor its index. - Common Crawl:
CCBot, whose public dataset has been used to train many AI models.
Check each company's documentation for current names; they add and rename crawlers.
Do they actually comply?
The big companies' declared crawlers generally follow robots.txt. Problems come from crawlers that don't identify themselves. In August 2025, Cloudflare reported that Perplexity was using undeclared crawlers to reach sites that had blocked its declared ones, which Perplexity disputed. Requests a user triggers directly can also be treated differently from automatic crawling, depending on the company.
Block training, allow search
Many sites want to appear in AI search results but not be used for training. In robots.txt:
User-agent: GPTBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: OAI-SearchBot
Allow: /
Blocking AI search crawlers means AI tools can't cite your pages. See getting cited by AI.
Stronger controls
- Firewall rules: services such as Cloudflare can block AI crawlers by behavior, not just name. Since July 2025, Cloudflare has blocked known AI crawlers by default on new sites.
- Login walls and rate limits for content you don't want collected.
- Server logs: check which crawlers actually visit and how often.
What robots.txt can't do
- Remove content already collected.
- Stop crawlers that ignore it.
- Protect anything legally; it's a request, not a lock.
Decide what you want
For most businesses, being cited by AI search is worth more than blocking it. For publishers whose content is the product, limiting training use may matter more. See also how ChatGPT browses the web.