How to Block AI Crawlers From Scraping Your Website
Patrick Bushe
June 30, 2024 · 5 min read
AI companies use crawlers to collect web pages, both to train models and to answer questions in AI search. You can ask them to stay away with robots.txt, and block them more firmly at your server or CDN. First decide what you're blocking, because blocking everything can also remove you from AI answers.
Know which crawler does what
- Training: OpenAI's GPTBot, Anthropic's ClaudeBot, Common Crawl's CCBot, Applebot-Extended and Bytespider collect pages that may be used to train models.
- AI search: OAI-SearchBot (ChatGPT search) and PerplexityBot fetch pages so they can show and link to them in answers.
- User requests: ChatGPT-User fetches a page when someone asks ChatGPT to read it.
- Google-Extended is a control, not a separate crawler: it tells Google not to use your pages for Gemini. It doesn't affect Google Search, and AI Overviews are part of Search.
Block with robots.txt
Add rules to the robots.txt file at the root of your site. To block training but allow AI search:
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: Google-Extended
Disallow: /
The AI crawler robots.txt builder writes these rules for each bot you pick, and the robots.txt tester checks that a rule blocks what you meant.
The limits of robots.txt
- It's a request, not a lock. Major AI companies say they follow it, but some crawlers have been reported ignoring it or hiding their identity.
- It doesn't remove content already collected.
- New crawlers appear often, so review your list a few times a year. See do AI chatbots respect robots.txt?
Firmer blocking
- CDN settings: Cloudflare and similar services can block known AI crawlers with a setting, without editing files.
- Server rules: block by user agent in your server or firewall, and check IP addresses against the ranges some AI companies publish, since user agents can be faked.
- Logins and paywalls: content behind a login isn't reachable by crawlers at all.
Should you block them?
It depends on what your site does. Publishers whose work is the product often block training crawlers. Businesses that want customers to find them usually allow AI search crawlers, because being cited in ChatGPT or Perplexity answers sends visitors. Note that llms.txt is a guide for AI tools, not a way to block them.
If you want your business to appear in AI answers rather than disappear from them, see AI search optimization.