Every one of those four answers Googlebot perfectly happily. Their robots.txt files are clean. No standard SEO audit would ever surface this, and the companies almost certainly do not know it is happening.
28 mid-market B2B domains, drawn from two unrelated verticals: industrial and environmental services, and fractional marketing leadership. Each homepage was requested six times over HTTPS, once as a desktop browser, once as Googlebot, and once as each of four answer-engine crawlers. Nothing else. No robots.txt parsing, no assumptions, just what the server actually returns to each requester.
The interesting number is not how many sites block crawlers. It is how many block the answer engines specifically while leaving classic search untouched. That asymmetry is the signature of a bot-protection default that predates AI crawlers, not a policy anyone sat down and chose.
| Crawler | Blocked | Rate |
|---|---|---|
| ClaudeBot (Claude) | 5 | 19% |
| GPTBot (ChatGPT) | 3 | 11% |
| PerplexityBot | 2 | 7% |
| Google-Extended (Gemini) | 2 | 7% |
ClaudeBot is turned away most often, which tracks: it is the newest of the four and the least likely to be on anyone's allowlist.
An answer engine cannot quote a page it is not allowed to fetch. Not "ranks it lower." Cannot quote it at all. Every content investment, every schema fix, every thought-leadership piece is invisible to ChatGPT for as long as the 403 stands.
One of the blocked sites in this scan is cited 11 times across the AI answer engines. Its closest competitor, a company of comparable size in the same vertical, is cited 77 times. The blocked site has a clean robots.txt, a healthy domain rating, and no idea.
The fix is usually a single allowlist rule at the CDN. The hard part is finding out.
The free AI Visibility Check runs exactly this test against your site, live, along with Core Web Vitals and structured data. It reports what it found and it names what it could not measure instead of quietly leaving it out.