Introducing Highlights: the context that matters
Free Tools

Robots.txt Checker and Tester

Fetch any site's robots.txt, test whether a URL is allowed for Googlebot, GPTBot, ClaudeBot, or any user agent, and see its sitemaps and AI crawler rules.

How do I test a robots.txt file?

  1. Enter a domain, such as example.com, and click Fetch robots.txt.
  2. Choose a user agent from the list or type your own, such as Googlebot or GPTBot.
  3. Type a path or a full URL to test. The verdict updates as you type and names the rule that decided it.
  4. Read the AI crawler table, the listed sitemaps, and any ignored lines to see the rest of the file at a glance.

How does robots.txt decide which URLs a crawler can fetch?

The file is split into groups, each starting with one or more User-agent lines followed by Allow and Disallow rules. A crawler picks the group that names it, or the * group if none does, then compares each rule path against the URL path. The longest matching rule wins, and an empty Disallow allows everything. See the robots.txt glossary entry for the full syntax.

Should I block AI crawlers in robots.txt?

It depends on what you want. Blocking a training crawler such as GPTBot or ClaudeBot keeps your content out of future model training, while blocking a search crawler such as OAI-SearchBot or Claude-SearchBot can remove you from AI search answers. The table above shows which of them your file already blocks. Rules for a crawler that is not listed fall through to the * group.

How do I check robots.txt before scraping at scale?

Fetch the file once per domain, cache it, and test each URL against it before you request the page. The web scraping API returns the pages you choose to fetch, and the free sitemap checker shows which pages a site wants crawled.

What can you use a robots.txt tester for?

  • Confirming that a new Disallow rule blocks what you expect and nothing else
  • Finding out why Googlebot or another crawler cannot reach a page
  • Auditing which AI crawlers a site allows before you publish or scrape
  • Checking that a robots.txt lists your sitemap
  • Reading a competitor's crawl rules before you build a crawler

Frequently asked questions

Is the robots.txt tester free?
Yes. You can check up to 50 sites a day, with no account needed. Testing different paths and user agents against a file you already fetched does not count again.
Which rules does the tester follow?
It follows RFC 9309, the Robots Exclusion Protocol standard. A crawler uses the most specific group for its name, falls back to the * group, and inside that group the longest matching path wins. When an Allow and a Disallow rule match with the same length, Allow wins. The * wildcard and the $ end anchor are supported.
What happens when a site has no robots.txt?
A 404 or other 4xx response means there are no restrictions, so every path is allowed. A 5xx response means the file is unreachable, and the standard tells crawlers to assume the whole site is disallowed until it comes back.
Does robots.txt keep a page out of Google or AI answers?
Not reliably. Robots.txt controls crawling, not indexing. A blocked URL can still appear in search results if other pages link to it. To keep a page out of an index, use a noindex directive on a page crawlers are allowed to fetch.
Which AI crawlers does the tool check?
GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Google-Extended, Applebot-Extended, and CCBot. The table shows whether each is allowed, partly restricted, or blocked from the whole site.
Should I respect robots.txt when I scrape a site?
Yes, check it before you crawl, and treat Disallow rules as the site owner stating its preference. This tool shows what a site asks of your crawler. Whether a given use is allowed also depends on the site terms and on the law where you operate.

Ship an agent that actually knows things.

Free tier, 10-minute integration, and the same API powering agents at Mintlify, daily.dev, and Propane. No credit card to start.