How the tester decides

This tester implements RFC 9309, the Robots Exclusion Protocol standard published in 2022 and used by Google, Bing and most well-behaved crawlers:

  1. Pick the group. It reduces your user-agent to a product token (Mozilla/5.0 (compatible; Googlebot/2.1) becomes Googlebot) and looks for User-agent lines that match it, ignoring case. All matching groups are merged. If none match, the User-agent: * group is used. If there is no * group either, everything is allowed.
  2. Find matching rules. Each Allow and Disallow path is compared with the start of the URL path and query string. * matches any characters and a trailing $ means "ends here".
  3. Longest rule wins. The rule with the most characters is the most specific and decides the answer. If an Allow and a Disallow tie, Allow wins. Line order does not matter.

Reading the result

  • Allowed / Blocked is the answer for that exact path and crawler.
  • Matched rule shows the winning line, so you can see which line to edit. "None" means no rule matched and the path is allowed by default.
  • All rules that match appears when several rules compete, with their lengths, so you can see why one won.
  • When fetching, the HTTP status matters: a 404 means no rules (allow all), while a 5xx means crawlers should treat the whole site as disallowed until the file can be read.

Common mistakes

  • Disallow: with nothing after it blocks nothing. To block everything use Disallow: /.
  • Expecting the * group to add to a named group. A crawler with its own group ignores the * rules completely, so repeat shared rules in each group.
  • Case. Paths are case-sensitive: Disallow: /Admin does not block /admin. User-agent names are not.
  • Blocking to hide. robots.txt is public and stops only polite crawlers. Use authentication or noindex for pages that must stay out of search results; a disallowed page can still be indexed from links.
  • Wrong place. The file must be at the root of each host and protocol: https://www.example.com/robots.txt does not cover https://shop.example.com.

FAQ

Which rule wins when Allow and Disallow both match?

The most specific one: the rule with the longest path pattern. If an Allow and a Disallow rule are exactly as long, Allow wins. The order of lines in the file does not matter.

What do * and $ mean in robots.txt?

* matches any sequence of characters (including none), and $ at the end of a rule means the URL must end there. Disallow: /*.pdf$ blocks /files/report.pdf but not /files/report.pdf?download=1.

Which group applies to my crawler?

The group whose User-agent line matches your crawler's product token (case-insensitive), such as Googlebot. If several groups name it they are merged. Only if none does, the User-agent: * group applies; the * rules are not added on top of a specific group.

What if robots.txt returns 404 or 500?

Under RFC 9309 a 4xx response (like 404) means there are no restrictions, so everything is allowed. A 5xx error or an unreachable server means the crawler should assume everything is disallowed until it can read the file.

Is robots.txt legally binding?

It is a voluntary standard, not access control, and it does not hide anything: blocked URLs are still public. Respecting it is part of polite, defensible scraping; see the lessons below.