Robots.txt and AI Bots: What to Allow, What to Block

Quick answer: For almost every business that wants customers, the right robots.txt policy is: allow retrieval bots (OAI-SearchBot, PerplexityBot, ChatGPT-User) unconditionally – they create citations that send you buyers – and allow training bots (GPTBot, ClaudeBot, Google-Extended) unless you sell content itself. Blocking AI bots to “protect content” is, for a visibility-seeking business, the modern equivalent of blocking Google in 1999.

Know the two bot species

Retrieval/search bots fetch pages to answer live user questions with citations: OAI-SearchBot, PerplexityBot, plus user-triggered fetchers ChatGPT-User and Perplexity-User. Blocking these removes you from AI answers directly – pure visibility loss, no upside for a normal business.

Training bots collect data for future models: GPTBot, ClaudeBot, Google-Extended, Meta-ExternalAgent. Blocking them keeps your content out of next-generation model memory – defensible for publishers monetising content, self-sabotage for businesses that want AI engines to know and recommend them.

Recommended configuration

For a visibility-first business, simply ensure no disallow rules target: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, Google-Extended, or Bingbot (Bing feeds multiple AI engines). Explicit allow blocks are optional – absence of disallow means allowed. Then verify behaviour in logs, because robots.txt is only one gate: firewalls override it silently (full crawlability guide).

When blocking is rational

Paywalled publishers, sites whose content is the product, and businesses with contractual confidentiality obligations may reasonably block training bots while keeping retrieval bots open – a split policy robots.txt handles cleanly with per-user-agent groups. Block retrieval bots only if you genuinely never want to appear in AI answers.

Common mistakes

Blanket disallows inherited from staging, security plugins silently adding AI-bot blocks, blocking Bingbot while expecting Perplexity citations, and treating robots.txt as enforcement (bad actors ignore it – it is a signal for compliant crawlers, which all the majors are). Audit quarterly: rules rot.

FAQ

Does allowing AI bots hurt rankings or bandwidth?

Ranking impact: none. Bandwidth: negligible for normal sites – AI crawlers are modest compared with scrapers you already tolerate.

How do I confirm my policy is working?

Check logs for each bot’s visits and status codes (detection guide), or run IndexGraph.ai‘s crawler audit for a bot-by-bot access report.

Unsure what your robots.txt is actually telling AI engines? Adexorb Technologies will audit it free as part of any GEO consultation.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top