Whether AI assistants can read your website is decided by a handful of user-agents and your robots.txt. In 2023 many sites blocked them all on principle; in 2026, with AI answers driving discovery, that same block is quietly deleting brands from the channel. Here is who the crawlers are and how to set policy deliberately.
The crawlers that matter
- GPTBot (OpenAI) — collects content that can inform OpenAI models. Blocking it keeps you out of that pipeline.
- ChatGPT-User (OpenAI) — fetches pages when a ChatGPT user's request needs live browsing; blocking it breaks live citations of your site.
- ClaudeBot (Anthropic) — crawls for Claude; same trade-offs as GPTBot.
- PerplexityBot — powers Perplexity's web-grounded, citation-heavy answers; blocking it removes you from the most citation-friendly engine.
- Google-Extended — controls whether Google may use your content for its AI models; separate from normal Googlebot indexing.
- CCBot (Common Crawl) — feeds the open dataset many models train on.
Blocking vs allowing: the honest trade-off
Allowing AI crawlers means your content can be read, cited, and represented in answers — and may inform model training. Blocking protects content from reuse but makes you unciteable and, over time, underrepresented in model knowledge. For content businesses selling the content itself, blocking can be rational. For every business that wants to be found and recommended, allowing is the default that matches the goal.
The most expensive misconfiguration we see: a blanket User-agent: * Disallow: / or a 2023-era AI block that nobody remembers, silently keeping every engine from ever citing the site.
The robots.txt for maximum AI visibility
User-agent: GPTBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: ClaudeBot Allow: / User-agent: PerplexityBot Allow: / User-agent: Google-Extended Allow: / User-agent: CCBot Allow: / Sitemap: https://yoursite.com/sitemap.xml
Selective policies are legitimate too: many sites allow the browsing agents (ChatGPT-User, PerplexityBot) for citations while blocking training-oriented ones (GPTBot, CCBot). Just make it a decision, not an accident — and re-verify after site migrations, when robots.txt files most often regress.
Verify in 60 seconds
- Open
yoursite.com/robots.txtand read every Disallow block. - Check that no wildcard block covers the AI agents you want to allow.
- Confirm the Sitemap line exists and resolves.
- Run an AI-readiness scan to confirm crawl access plus the rest of the machine-readability stack.
Key takeaways
- Six user-agents decide your AI readability: GPTBot, ChatGPT-User, ClaudeBot, PerplexityBot, Google-Extended, CCBot.
- Blocking protects content but deletes you from citations and, over time, model knowledge.
- Browsing agents vs training agents can be handled with different policies — deliberately.
- Re-check robots.txt after every migration; regressions are common and silent.
Includes crawler check
Scan your site's AI crawler access now
The free audit flags exactly which AI bots your robots.txt blocks — plus 10 other AI-readiness checks.
Frequently asked questions
Should I block GPTBot?
Only if protecting content from model training outweighs being cited and recommended by ChatGPT. For most businesses seeking customers, allowing GPTBot (and especially ChatGPT-User, the live-browsing agent) is the right default.
What is the difference between GPTBot and ChatGPT-User?
GPTBot crawls broadly to gather content that can inform OpenAI models. ChatGPT-User fetches specific pages in real time when a user's ChatGPT request needs browsing — blocking it prevents ChatGPT from citing your live pages.
Does allowing AI crawlers hurt my Google SEO?
No. AI-crawler directives are separate user-agent rules; normal Googlebot indexing is unaffected. Google-Extended specifically controls AI-model use of your content, not search indexing.