If AI crawlers can't access your content, it can't be cited in AI-generated answers from ChatGPT, Claude, or Perplexity. Allowing them is a business decision, similar to allowing search engine indexing.
See exactly which AI crawlers — GPTBot, ClaudeBot, Google-Extended, PerplexityBot, and more — are allowed or blocked on your site, in plain language, by reading your live robots.txt.
robots.txt allows their crawlers in
the first place. Many sites unknowingly block every AI crawler by copying an overly broad
"block all bots" rule years ago — this tool reads your live robots.txt and tells you exactly
which AI crawlers can currently access your site, and which ones are shut out.
Any domain — no login or verification needed.
Read live, directly from your server.
Allowed, Blocked, or Partially Blocked — with the exact rule that determined it.
Update your robots.txt with our generator if anything needs changing.
| Crawler | Company | Purpose |
|---|---|---|
| GPTBot | OpenAI | Trains ChatGPT / GPT models |
| ChatGPT-User | OpenAI | Live browsing during a ChatGPT conversation |
| OAI-SearchBot | OpenAI | Powers ChatGPT search results |
| Google-Extended | AI training (Gemini, AI Overviews) — separate from Googlebot search indexing | |
| ClaudeBot | Anthropic | Trains Claude models |
| Claude-Web | Anthropic | Live browsing during a Claude conversation |
| anthropic-ai | Anthropic | Legacy Anthropic crawler identifier |
| PerplexityBot | Perplexity | Crawls for Perplexity AI search answers |
| Perplexity-User | Perplexity | Live browsing during a Perplexity query |
| CCBot | Common Crawl | Open dataset used to train many AI models |
| Applebot-Extended | Apple | AI training for Apple Intelligence |
| Bytespider | ByteDance / TikTok | Crawls for ByteDance AI products |
| Amazonbot | Amazon | Crawls for Amazon AI products and Alexa |
| Meta-ExternalAgent | Meta | Trains Meta AI / Llama models |
| cohere-ai | Cohere | Trains Cohere language models |
| Bingbot | Microsoft | Search indexing — also powers Copilot answers |
If AI crawlers can't access your content, it can't be cited in AI-generated answers from ChatGPT, Claude, or Perplexity. Allowing them is a business decision, similar to allowing search engine indexing.
No. robots.txt is voluntary. Reputable companies say they respect it, but it has no technical enforcement.
Googlebot handles search indexing. Google-Extended separately controls whether your content trains Google's AI and powers AI Overviews — you can block one without blocking the other.