Shruwd
Shruwd

Can AI crawlers
read your site?

Enter a domain. Shruwd reads its robots.txt and its homepage the way a crawler does, and tells you what it found.

What the scan checks

Two requests to your site, and nothing is stored.

robots.txt, for every AI crawler token

Shruwd fetches your robots.txt and works out, token by token, whether the crawlers of OpenAI, Anthropic, Perplexity and Google’s AI products may fetch your homepage. It applies the precedence the standard sets, so rules naming one crawler override the rules for all of them. It also tells a block on a training crawler, which is often deliberate, from a block on a crawler that fetches pages to answer questions, which costs you citations.

Which crawlers Shruwd recognises

Your homepage, as raw HTML

AI crawlers largely do not run JavaScript. Shruwd fetches your homepage without running any and counts the readable text in the response. A page with almost none, and an empty app container where the content should be, is a page those crawlers receive blank, even while it looks complete in a browser and ranks on Google.

The findings these checks become

Questions and answers

Do I need an account?

No. Enter a domain and read the result. An account adds what a scan cannot do: measuring how often AI answers name you, and reading your server logs to see which AI crawlers are turned away.

What does the scan request from my site?

Two things: your homepage and your robots.txt, once each, from Shruwd’s server, with an ordinary browser user agent and a header that says it is a Shruwd scan. It does not crawl the rest of the site.

Why not fetch my site as GPTBot and see what happens?

Because the answer would mislead you. A request that claims to be GPTBot from an address OpenAI does not own is a spoofed crawler, and sites that verify crawlers by address refuse those on purpose. Whether the real crawlers are refused shows in your server logs, which the free plan reads, or in your CDN’s security analytics if the CDN refused them.

Is my result stored?

No. It is held in memory for ten minutes, so that repeating a scan does not fetch your site again, and then it is gone.

Can I scan a site that is not mine?

Yes. It reads two public files that any crawler reads.

It says my homepage could not be measured. Why?

Usually because a firewall refused the request. The scan comes from a data-centre address, and some firewalls turn those away. That is about our request, not about AI crawlers, so it tells you little either way.

Find out what the answers say about you.

One measurement, your own crawler data, and the three access findings that matter most. Free, no card, about two minutes.

Start free