Back to Blog

Do AI Chatbots Respect Robots.txt and Crawling Rules

PB

Patrick Bushe

July 1, 2024 · 5 min read

Robots.txt is a plain text file where website owners tell automated crawlers which pages they may visit. It was standardized as RFC 9309 in 2022, but it has always been voluntary. Most major AI companies say their crawlers follow it; the details matter.

The main AI crawlers

  • OpenAI: GPTBot collects training data, OAI-SearchBot indexes for ChatGPT search, and ChatGPT-User visits pages when a user asks.
  • Google: Google-Extended is a robots.txt token controlling whether content is used for Gemini training. It doesn't affect Google Search or AI Overviews, which use regular Googlebot.
  • Anthropic: ClaudeBot for training, plus separate agents for search and user requests.
  • Apple: Applebot-Extended controls use for Apple's AI training.
  • Perplexity: PerplexityBot for its index.
  • Common Crawl: CCBot, whose public dataset has been used to train many AI models.

Check each company's documentation for current names; they add and rename crawlers.

Do they actually comply?

The big companies' declared crawlers generally follow robots.txt. Problems come from crawlers that don't identify themselves. In August 2025, Cloudflare reported that Perplexity was using undeclared crawlers to reach sites that had blocked its declared ones, which Perplexity disputed. Requests a user triggers directly can also be treated differently from automatic crawling, depending on the company.

Block training, allow search

Many sites want to appear in AI search results but not be used for training. In robots.txt:

User-agent: GPTBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: OAI-SearchBot
Allow: /

Blocking AI search crawlers means AI tools can't cite your pages. See getting cited by AI.

Stronger controls

  • Firewall rules: services such as Cloudflare can block AI crawlers by behavior, not just name. Since July 2025, Cloudflare has blocked known AI crawlers by default on new sites.
  • Login walls and rate limits for content you don't want collected.
  • Server logs: check which crawlers actually visit and how often.

What robots.txt can't do

  • Remove content already collected.
  • Stop crawlers that ignore it.
  • Protect anything legally; it's a request, not a lock.

Decide what you want

For most businesses, being cited by AI search is worth more than blocking it. For publishers whose content is the product, limiting training use may matter more. See also how ChatGPT browses the web.

More Tools by Patrick Bushe

Free Chrome extensions to boost your productivity and privacy