freeRobots.txt Tester & Generator

Robots.txt tester: which line decides?

Fetch a site's robots.txt or paste your own, pick any crawler, and test as many paths as you like. You see allowed or blocked and the exact rule that decided it. Switch to the generator to write a new file.

Or paste or edit a file below. Testing runs in your browser as you type.

0.2 KB

Follows its own group: User-agent: gptbot (line 7).

Results for GPTBot

4 paths
  • Blocked/line 9 · Disallow: /
  • Blocked/admin/help/faqline 9 · Disallow: /
  • Blocked/blog/post?sessionid=42line 9 · Disallow: /
  • Blocked/searchline 9 · Disallow: /

The file, with the deciding rules

16 lines
  1. 1User-agent: *
  2. 2Disallow: /admin/
  3. 3Disallow: /cart
  4. 4Disallow: /*?sessionid=
  5. 5Allow: /admin/help/
  6. 6
  7. 7User-agent: GPTBot
  8. 8User-agent: CCBot
  9. 9✕Disallow: /
  10. 10
  11. 11User-agent: Googlebot
  12. 12Disallow: /search
  13. 13Crawl-delay: 5
  14. 14
  15. 15Sitemap: https://example.com/sitemap.xml
  16. 16

Things to know

1
  • Crawl-delay on line 13. Google ignores it; Bing and some others honour it.

Groups

3
  • L1 *3 disallow · 1 allow
  • L7 GPTBot, CCBot1 disallow · 0 allow
  • L11 Googlebot1 disallow · 0 allow · crawl-delay 5

Sitemaps

  • L15 https://example.com/sitemap.xml

How matching works

Three rules decide every URL.

robots.txt looks simple, and most mistakes come from assuming it's read top to bottom. It isn't. Since RFC 9309 (2022), and in Google's parser for years before that, a crawler does this:

  1. 1Find its group

    The group whose User-agent names the crawler wins. Only if there's none does it fall back to User-agent: *. It never uses both.

  2. 2Collect matching rules

    Every Allow and Disallow in that group whose path matches the start of the URL, with * and $ wildcards.

  3. 3Longest wins

    The most specific (longest) matching rule decides. If an Allow and a Disallow tie, Allow wins.

  4. 4No match means allowed

    An empty Disallow, or no matching rule, allows the URL. /robots.txt itself is always allowed.

The tester implements exactly that, keeps line numbers for every rule, and shows the group a crawler ends up in. It also flags lines crawlers will skip: unknown directives, rules before any User-agent, paths that don't start with a slash, and Noindex, which Google stopped supporting in robots.txt in 2019.

Common mistakes

What we see break most often.

A named group that forgets the basics. Adding User-agent: Googlebot with one rule means Googlebot no longer sees your User-agent: * exclusions. Repeat the rules you still want in the named group.

Blocking CSS and JavaScript. Disallowing /assets/ or /wp-includes/ stops Google from rendering your pages properly. The tester warns when a rule looks like it covers asset files.

Blanket AI blocks. A Disallow: / under a list of AI user agents often catches search crawlers like OAI-SearchBot and PerplexityBot along with the training ones. Check what each AI crawler is allowed with the AI crawler checker.

A robots.txt that errors. If /robots.txt returns a 5xx, Google treats the whole site as disallowed until it can read the file again. A 404 is fine; it means “no restrictions”.

The generator

A sensible file in thirty seconds.

Pick a starting point, adjust individual crawlers, list the paths you want kept out of every crawler, and add your sitemap. The presets reflect the real choices: allow everything, let AI search cite you while opting out of model training, or block AI entirely and keep classic search.

Your draft is saved in this browser. Paste the output into the tester tab to check it against your real URLs before you upload it, and point a sitemap validator at the sitemap you list.

FAQ

Good questions.

Keep going

Related free tools

All 24 tools

Found problems? We fix them, then get you named in AI answers.

These tools show what's broken. In 20 minutes we'll show you who ChatGPT recommends instead of you, and the three fixes we'd make first.

Book a call

Or get the free AI visibility report by email.

no strings, really