freeAI Crawler Access Checker
Can AI crawlers reach your pages?
Enter a URL. We check 21 crawlers from OpenAI, Anthropic, Perplexity, Google and others against your robots.txt, show the exact line that decides each one, then request the page as six of them to catch firewall blocks.
Try
What it checks
Two ways to lock a crawler out.
A crawler can be stopped politely or at the door. robots.txt is the polite way: a text file that says which bots may fetch which paths, and well-behaved crawlers read it first. A firewall is the door: your CDN or server refuses the request outright, often because someone switched on a “block AI bots” setting without thinking about AI search.
1robots.txt, rule by rule
We parse your file the way RFC 9309 and Google describe: the most specific user-agent group wins, the longest matching rule wins, Allow wins ties.
221 crawlers, grouped by purpose
AI search, assistant fetchers, classic search engines and AI training. Blocking each group costs you something different.
3Live requests
We fetch your page as GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Googlebot and Bingbot and compare each response with a normal browser's.
4Firewall fingerprints
403s, 429s, Cloudflare challenge pages, and responses that are suspiciously smaller than what a browser gets.
5Page-level directives
noindex, nosnippet and noai in meta robots tags and the X-Robots-Tag header.
6A fix you can paste
A robots.txt group that lets AI search in without undoing the rest of your file.
Search vs. training
Not every AI bot is the same bot.
Most AI companies now run separate crawlers for separate jobs. OpenAI's GPTBot gathers training data; OAI-SearchBot builds the index ChatGPT search draws on; ChatGPT-User fetches a page when someone asks about it. Anthropic splits the same way with ClaudeBot, Claude-SearchBot and Claude-User. Perplexity runs PerplexityBot for its index and Perplexity-User for live requests.
That split is what makes a sensible policy possible. If you want to be cited in AI answers, the search crawlers have to get in. Whether training crawlers get in is a separate decision, and for most businesses that want to be recommended, letting models learn about you is worth more than keeping them out.
Two tokens never fetch anything. Google-Extended and Applebot-Extended exist only in robots.txt, as a way to tell Google and Apple not to use what Googlebot and Applebot already crawled for model training. Blocking them doesn't affect Google Search, AI Overviews or Siri.
Fixing it
Unblocking AI search without opening everything.
If a search crawler is blocked by robots.txt, the fix is to give it its own group. A crawler that matches a named group ignores the User-agent: * rules entirely, so the snippet the tool generates copies your existing exclusions (admin, cart, internal search) into the new group rather than opening them up.
If the block is at the firewall, robots.txt can't help. On Cloudflare, look at the AI crawler and bot settings under Security, and check whether verified AI search bots are allowed. On other CDNs and WAFs, search the logs for the crawler's user agent and see which rule fired. Then check again here, and test your edited file in the robots.txt tester before you deploy it.
Access is only the first step. A crawler that gets in but finds an empty JavaScript shell still has nothing to quote. See what it actually reads with What AI Sees.
FAQ
Good questions.
Keep going
Related free tools
Found problems? We fix them, then get you named in AI answers.
These tools show what's broken. In 20 minutes we'll show you who ChatGPT recommends instead of you, and the three fixes we'd make first.
Book a callOr get the free AI visibility report by email.
no strings, really