How the tester decides
This tester implements RFC 9309, the Robots Exclusion Protocol standard published in 2022 and used by Google, Bing and most well-behaved crawlers:
- Pick the group. It reduces your user-agent to a product token (
Mozilla/5.0 (compatible; Googlebot/2.1)becomesGooglebot) and looks forUser-agentlines that match it, ignoring case. All matching groups are merged. If none match, theUser-agent: *group is used. If there is no*group either, everything is allowed. - Find matching rules. Each
AllowandDisallowpath is compared with the start of the URL path and query string.*matches any characters and a trailing$means "ends here". - Longest rule wins. The rule with the most characters is the most specific and decides the answer. If an Allow and a Disallow tie, Allow wins. Line order does not matter.
Reading the result
- Allowed / Blocked is the answer for that exact path and crawler.
- Matched rule shows the winning line, so you can see which line to edit. "None" means no rule matched and the path is allowed by default.
- All rules that match appears when several rules compete, with their lengths, so you can see why one won.
- When fetching, the HTTP status matters: a 404 means no rules (allow all), while a 5xx means crawlers should treat the whole site as disallowed until the file can be read.
Common mistakes
Disallow:with nothing after it blocks nothing. To block everything useDisallow: /.- Expecting the
*group to add to a named group. A crawler with its own group ignores the*rules completely, so repeat shared rules in each group. - Case. Paths are case-sensitive:
Disallow: /Admindoes not block/admin. User-agent names are not. - Blocking to hide. robots.txt is public and stops only polite crawlers. Use authentication or
noindexfor pages that must stay out of search results; a disallowed page can still be indexed from links. - Wrong place. The file must be at the root of each host and protocol:
https://www.example.com/robots.txtdoes not coverhttps://shop.example.com.
FAQ
Which rule wins when Allow and Disallow both match?
The most specific one: the rule with the longest path pattern. If an Allow and a Disallow rule are exactly as long, Allow wins. The order of lines in the file does not matter.
What do * and $ mean in robots.txt?
* matches any sequence of characters (including none), and $ at the end of a rule means the URL must end there. Disallow: /*.pdf$ blocks /files/report.pdf but not /files/report.pdf?download=1.
Which group applies to my crawler?
The group whose User-agent line matches your crawler's product token (case-insensitive), such as Googlebot. If several groups name it they are merged. Only if none does, the User-agent: * group applies; the * rules are not added on top of a specific group.
What if robots.txt returns 404 or 500?
Under RFC 9309 a 4xx response (like 404) means there are no restrictions, so everything is allowed. A 5xx error or an unreachable server means the crawler should assume everything is disallowed until it can read the file.
Is robots.txt legally binding?
It is a voluntary standard, not access control, and it does not hide anything: blocked URLs are still public. Respecting it is part of polite, defensible scraping; see the lessons below.