AI Crawler Access Checker

AI companies crawl the web to train models and to answer questions with live citations. This checker reads your robots.txt and shows, crawler by crawler, whether each major AI bot is allowed or blocked - so you can decide who gets access to your content.

10
Allowed
2
Blocked
GPTBotTrainingOpenAI
Trains ChatGPT models
Blocked
OAI-SearchBotSearchOpenAI
Cites pages in ChatGPT search
Allowed
ChatGPT-UserLive fetchOpenAI
Fetches when a user asks ChatGPT to browse
Allowed
ClaudeBotTrainingAnthropic
Trains Claude models
Allowed
Claude-UserLive fetchAnthropic
Fetches for Claude user requests
Allowed
PerplexityBotSearchPerplexity
Indexes pages for Perplexity answers
Allowed
Perplexity-UserLive fetchPerplexity
Fetches when a user asks Perplexity
Allowed
Google-ExtendedTrainingGoogle
Gemini training (not Google Search)
Allowed
CCBotTrainingCommon Crawl
Open dataset used to train many models
Blocked
BytespiderTrainingByteDance
Trains ByteDance / Doubao models
Allowed
Meta-ExternalAgentTrainingMeta
Trains Meta AI / Llama
Allowed
Applebot-ExtendedTrainingApple
Trains Apple Intelligence
Allowed

Want to block AI training crawlers?

Copy this snippet into your robots.txt, or build a full file with our robots.txt generator.

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: CCBot
User-agent: Bytespider
User-agent: Meta-ExternalAgent
User-agent: Applebot-Extended
Disallow: /

How to use the ai crawler check tool

  1. 1Paste the contents of your current robots.txt (from yourdomain.com/robots.txt).
  2. 2Review the status of each AI crawler - allowed or blocked.
  3. 3Decide which to allow (for AI answer visibility) or block (to protect training data).
  4. 4Use our robots.txt generator to create the exact directives you want.

Frequently asked questions

Should I block or allow AI crawlers?

It's a trade-off. Allowing crawlers like OAI-SearchBot and PerplexityBot can get your content cited in AI answers, driving referral traffic. Blocking bots like GPTBot and CCBot keeps your content out of model training. Many sites allow answer and search bots while blocking training bots.

What's the difference between GPTBot and OAI-SearchBot?

GPTBot crawls content to train OpenAI's models, while OAI-SearchBot fetches pages to cite in ChatGPT's live search answers. If you want visibility in ChatGPT's answers but not in training data, allow OAI-SearchBot and block GPTBot.

Does blocking an AI crawler remove me from its answers?

It stops that company from crawling your site directly, but models may still reference content they learned earlier or find through third parties. Blocking is about future access, not erasing existing knowledge.

Where do I add AI crawler rules?

In your robots.txt at the root of your domain, as separate User-agent groups. Our robots.txt generator can build the exact directives for you.

Do all AI companies respect robots.txt?

Most major ones - OpenAI, Anthropic, Google, and Perplexity - honor robots.txt, but compliance isn't universally guaranteed. For content you must protect, use server-side access controls, not robots.txt alone.

More free SEO tools

Skip the manual work - let Spook rank you

These tools help with the pieces. Spook does the whole job: it finds winnable queries, writes the articles, and publishes them on autopilot.