Web tool / 05

AI Crawler Checker

Translate a site’s robots.txt rules into a readable policy report for well-known AI crawlers.

AI Crawler Checker: TOEA evaluates robots.txt rules for a curated group of named AI, search, and dataset crawlers and shows the most specific matching rule. Processed securely on demand.

Category
Web tools
Runs
On TOEA's server
Cost
Free · no sign-up
Availability
Ready to use
AI crawler policy readerSafe public URL check

Explains robots.txt rules; it cannot guarantee crawler behavior or AI inclusion.

The AI crawlers this checker reads

Each row is a User-agent token you can name in robots.txt. Blocking a training crawler does not remove a site from ordinary search, and blocking a search crawler can keep pages out of that assistant's cited answers.

User-agent tokenOperatorWhat it is for
GPTBotOpenAICollects public pages that may be used to train OpenAI's models.
ChatGPT-UserOpenAIFetches a page when a ChatGPT user asks for it, rather than crawling on its own.
OAI-SearchBotOpenAIIndexes pages so they can appear, with links, in ChatGPT search results.
ClaudeBotAnthropicCollects public pages that may be used to train Anthropic's Claude models.
Claude-SearchBotAnthropicIndexes pages to improve search results in Claude.
PerplexityBotPerplexityIndexes pages for Perplexity's cited answers.
Google-ExtendedGoogleNot a separate crawler: a token that controls whether content Googlebot fetches may be used for Gemini. It does not affect Google Search.
CCBotCommon CrawlBuilds the open Common Crawl archive, widely used as AI training data.
BytespiderByteDanceCrawls for ByteDance products, including its AI models.
AmazonbotAmazonCrawls for Amazon services, including answers given by Alexa.

Crawler names and policies change; check each operator's own documentation before relying on a rule.

Block AI training, keep search

Several User-agent lines can share one group. This stops the training crawlers while leaving Googlebot, Bingbot, and the AI search crawlers alone:

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: CCBot
User-agent: Bytespider
Disallow: /

Block every listed AI crawler

Add the search and user-request agents to the same group. Ordinary search engines are unaffected because they are not named:

User-agent: GPTBot
User-agent: ChatGPT-User
User-agent: OAI-SearchBot
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: Google-Extended
User-agent: CCBot
User-agent: Bytespider
User-agent: Amazonbot
Disallow: /

Then run this checker again, and the robots.txt checker to confirm the file parses as intended.

How to use it

  1. Enter a page on the site you want to inspect.
  2. Choose “Check AI crawlers.”
  3. Review the matched policy and source rule for each crawler.

Privacy & limitations

The website URL is checked through TOEA’s API and is not stored. Robots.txt is voluntary and crawler names or behavior can change; an allowed result does not mean a site is included in any AI product.

Related tools

Frequently asked questions

Does “allowed” mean my content is used for AI?

No. It only means the current robots.txt rules do not block that named crawler on the root path.

Can robots.txt enforce a legal policy?

Robots.txt expresses crawler preferences; it is not access control or legal advice.

Will blocking GPTBot keep my pages out of ChatGPT?

Not entirely. GPTBot collects training data; ChatGPT search results come from OAI-SearchBot, and pages a user asks ChatGPT to open are fetched by ChatGPT-User. Each has its own token.

Does blocking Google-Extended hurt my Google Search rankings?

No. Google-Extended only controls whether content Googlebot fetches may be used for Gemini; Google says it does not affect Search inclusion or ranking.

Which crawlers does this check?

Ten tokens: GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, Claude-SearchBot, PerplexityBot, Google-Extended, CCBot, Bytespider, and Amazonbot. The table below says what each is for.

Free tool · runs on toea's server · no account required