Every one of those four answers Googlebot perfectly happily. Their robots.txt files are clean. No standard SEO audit would ever surface this, and the companies almost certainly do not know it is happening.
55 mid-market B2B domains, drawn from two unrelated verticals: industrial and environmental services, and fractional marketing leadership. Each homepage was requested six times over HTTPS, once as a desktop browser, once as Googlebot, and once as each of four answer-engine crawlers. Nothing else. No robots.txt parsing, no assumptions, just what the server actually returns to each requester.
The interesting number is not how many sites block crawlers. It is how many block the answer engines specifically while leaving classic search untouched. That asymmetry is the signature of a bot-protection default that predates AI crawlers, not a policy anyone sat down and chose.
| Crawler | Blocked | Rate |
|---|---|---|
| ClaudeBot (Claude) | 8 | 16% |
| GPTBot (ChatGPT) | 5 | 10% |
| PerplexityBot | 4 | 8% |
| Google-Extended (Gemini) | 3 | 6% |
ClaudeBot is turned away most often, which tracks: it is the newest of the four and the least likely to be on anyone's allowlist.
An answer engine cannot quote a page it is not allowed to fetch. Not "ranks it lower." Cannot quote it at all. Every content investment, every schema fix, every thought-leadership piece is invisible to ChatGPT for as long as the 403 stands.
One of the blocked sites in this scan is cited 11 times across the AI answer engines. Its closest competitor, a company of comparable size in the same vertical, is cited 77 times. The blocked site has a clean robots.txt, a healthy domain rating, and no idea.
The fix is usually a single allowlist rule at the CDN. The hard part is finding out.
The free AI Visibility Check runs exactly this test against your site, live, along with Core Web Vitals and structured data. It reports what it found and it names what it could not measure instead of quietly leaving it out.
This test sends a user-agent string. It does not come from Google's or OpenAI's IP ranges. A site that verifies crawlers properly, by reverse DNS rather than by user-agent, will refuse this test no matter which bot it claims to be, and it is right to do so.
So a site that refuses everything tells you nothing. That case is inconclusive and is not what the 12% figure counts.
The number above counts only the asymmetric case: the site accepted a Googlebot user-agent from the same IP, in the same second, and refused the answer engines. That cannot be IP verification, because the same unverified requester was let through moments earlier. It is user-agent filtering with the AI crawlers left off the list.
Crawler access is one layer. Score your whole marketing machine free, eleven questions, one tap each.