Web tool / 02

Robots.txt Checker

Turn a site’s robots.txt file into a quick, readable crawler-policy report.

Robots.txt Checker: Read a site's robots.txt crawler rules, agent groups, sitemap declarations and raw source. TOEA fetches the file from the site root and summarises the wildcard policy for the root path. This helps spot broad blocks and misplaced rules; it does not test whether every page can be indexed. Processed securely on demand.

Category
Web tools
Runs
On TOEA's server
Cost
Free · no sign-up
Availability
Ready to use
Robots.txt readerSafe public URL check

TOEA checks the site's root robots.txt file.

A short robots.txt example

Place the file at the root of the host, as /robots.txt. User-agent selects a crawler group, Disallow asks that group not to fetch matching paths, and Allow can make an exception. This example blocks a private directory except its public press folder and declares a sitemap.

User-agent: *
Disallow: /private/
Allow: /private/press/

Sitemap: https://example.com/sitemap.xml

The longest matching path wins; Allow wins a tie. Paths are case-sensitive. Rules from groups matching the same agent are combined; a named crawler does not inherit the wildcard group's rules. Protocol reference: https://www.rfc-editor.org/rfc/rfc9309.html

Common mistakes to check

Disallow: / blocks every path, while an empty Disallow value blocks nothing. A file under /blog/robots.txt does not control the site root, and rules on one host do not cover a subdomain. Avoid blocking assets a search engine needs to render a page.

Robots.txt controls crawling, not access or guaranteed removal from search results. Use authentication for private content. For a public page that must be excluded from search, allow crawling and serve noindex. This checker summarises groups and the wildcard root policy; it does not test every URL. Search guidance: https://developers.google.com/search/docs/crawling-indexing/robots/intro

How to use it

  1. Enter any page on the site you want to check.
  2. Choose “Check robots.txt.”
  3. Review its groups, rule totals, sitemaps, and full source.

Privacy & limitations

The site URL is sent to TOEA’s bounded fetch service and is not stored. A robots.txt file is advisory; this report does not prove that every crawler will follow it.

Related tools

Frequently asked questions

Where should robots.txt go?

At /robots.txt on the host you want to control, such as https://example.com/robots.txt. A file under /blog/ does not control the whole site, and a subdomain needs its own file.

Does Disallow remove a page from search results?

No. Robots.txt controls crawling, not indexing or access. A blocked URL may still appear in search. To exclude a public page, allow crawling and serve a noindex directive; protect private content with authentication.

What mistakes should I check first?

Disallow: / blocks the whole site, while an empty Disallow value blocks nothing. Check case-sensitive paths, group boundaries, and whether essential assets are blocked. A named crawler group does not inherit rules from the wildcard group.

Free tool · runs on toea's server · no account required