Should You Block or Allow GPTBot, PerplexityBot, and Google Extended? A robots.txt Decision Guide

blog-main
  • user1
    admin
  • time-and-date
    18 Sep, 2026
  • clock-2
    Blog

Your CTO wants to protect the company’s content. Your SEO team wants customers to find it in AI-generated answers. Both goals belong in the same conversation, but “block AI bots” is not a precise enough policy.

Before changing robots.txt, ask a more useful question: which content should be available, to whom, and for what purpose?

That is where this guide starts. Instead of choosing between blanket access and blanket restrictions, you will build a crawler policy around your business model, content ownership, and discovery goals.

For businesses that want AI search visibility without approving OpenAI’s training-oriented crawling, a practical starting point is to allow OAI-SearchBot and PerplexityBot while blocking GPTBot. Evaluate Google Extended separately because it controls specified Gemini training and grounding uses, not Google Search rankings or AI Overview eligibility. 

The question is not simply whether your website should “allow AI.” It is which agents should access which content, for what purpose, and under whose approval.

A SaaS company wants buyers to find its integration documentation. A subscription publisher wants to protect reporting that supports paid memberships. An online retailer wants accurate product recommendations without exposing customer accounts.

These businesses should not automatically use the same crawler policy.

An effective AI crawler robots.txt audit connects access permissions with commercial priorities. It separates search discovery from training preferences, identifies content that requires stronger protection, and verifies whether production settings match the approved policy.

For CTOs and SEO leaders, the objective is straightforward: make useful public content discoverable without granting access or approving content uses that the business does not intend.

Understand the AI Crawler Controls

Start by identifying what each agent or control actually does.

An AI company may operate a training crawler, a search crawler, and a user-triggered page fetcher. Those functions should not be treated as interchangeable.

Google-Extended adds another distinction. It is a robots.txt control token, not a separately identifiable HTTP crawler with that name in its user-agent string.

Agent or token Documented purpose Main policy question
GPTBot Crawls content that may be used to train OpenAI’s generative AI foundation models Do you approve training-oriented collection?
OAI-SearchBot Helps surface websites in ChatGPT search features Should approved public pages be discoverable through ChatGPT search?
ChatGPT Use Visits pages for certain user-initiated actions What enforceable restrictions apply to user-requested retrieval?
PerplexityBot Surfaces and links websites in Perplexity search results Should approved public pages support Perplexity search discovery?
Perplexity-User Retrieves pages in response to user actions How do you protect content when robots.txt generally does not apply?
Google-Extended Controls specified Gemini training and grounding uses of content. Google crawls. Do those content uses align with your policy?
Googlebot Crawls for Google Search and its search features Can Google access pages intended for Search?

These distinctions follow the platforms’ official documentation. Use them as the basis for decisions rather than assuming that every AI-related agent collects training data. 

GPT-Bot and ChatGPT Search Are Separate

GPT-Bot and OAI-SearchBot have independent controls.

OpenAI supports allowing OAI-SearchBot while disallowing GPTBot. This lets a website restrict training-oriented crawling without making the same choice about ChatGPT search discovery.

Blocking GPTBot therefore does not, by itself, block OAI-SearchBot.

Likewise, allowing GPTBot is not a documented requirement for appearing in ChatGPT searches. Its stated purpose concerns potential model-training collection, not search inclusion.

For an SEO team, this distinction prevents an expensive misunderstanding: sacrificing search access when the intended decision was only to restrict training.

For a CTO, it creates a clearer approval process. Search access and training-oriented collection can be documented and reviewed separately.

PerplexityBot Is Search-Oriented

PerplexityBot is designed to surface and link websites in Perplexity search results. Perplexity states that it is not used to crawl content for AI foundation models.

For a public lead-generation website, allowing it may align with the purpose of publishing service pages, implementation guides, product comparisons, and educational articles.

However, blocking PerplexityBot does not guarantee that every reference to the website disappears. Perplexity says it may still index a blocked site’s domain, headline, and a brief factual summary while not indexing its full or partial text content. 

Treat the decision as a content-access preference, not a guaranteed removal mechanism.

Google Extended Includes a Grounding Tradeoff

Google-Extended is often described too narrowly as a training bot.

Google documents it as a standalone product token controlling whether crawled content may be used for training future Gemini models and for grounding in Gemini apps and grounding with Google Search on Vertex AI. 

Grounding supplies source information to a model when it generates a response.

Blocking Google-Extended therefore restricts its specified grounding uses as well as training uses. That makes its scope different from GPTBot’s documented role.

Google also states that Google-Extended does not affect inclusion in Google Search and is not a Search ranking signal. AI Overviews and AI Mode use search-related controls rather than Google Extended as an inclusion switch. 

Skyram’s guide to the technical SEO foundation for AEO provides context for the broader crawlability review. Within that review, keep Google Search access, Gemini content-use preferences, and private-content protection as separate decisions.

Choose by Business Model

The strongest AI crawler robots.txt policy starts with how your business earns revenue.

Does public content attract customers to a separate product? Is the information itself the product? Do you publish licensed material? Which pages contain sensitive information?

The matrix below provides recommended starting positions based on those questions and the agents’ documented purposes. It is a planning framework, not a universal rule or legal conclusion.

Business model GPTBot starting position OAI-SearchBot and PerplexityBot Google-extended decision Main consideration
B2B SaaS and technology services Block unless a training-oriented collection is approved Allow public product, documentation, and solution pages. Evaluate specified Gemini grounding alongside training preferences. Separate acquisition content from proprietary material
E-commerce retailer Block unless approved. Allow public products, categories, and buying guides. Evaluate public discovery value and approved content uses. Protect accounts, checkout, and customer information
Local service business Block if the training-oriented collection has no approved purpose. Allow public services, locations, and FAQs. Consider allowing approved public marketing content. Make services discoverable without exposing private workflows.
Advertising supported the publisher. Block pending a commercial and rights review. Selectively allow content with measurable discovery value. Evaluate training and grounding together. Compare referral value with answer substitution risk.
Subscription publisher Block unless licensed or approved. Allow previews and selected public articles. Evaluate previews separately from premium content. Protect the paid information product.
Proprietary data provider Block pending explicit approval. Allow summaries, methodology, and marketing pages. Restrict valuable proprietary material. Protect the underlying commercial asset
Healthcare or financial services Block unless approved. Allow reviewed public information. Apply a documented governance decision. Protect sensitive records and maintain accurate public information
Open educational resource or nonprofit Allow if rights and mission support it. Allow approved public resources. Allow if consistent with rights and mission. Expand access while respecting licensing restrictions.

The essential qualifier is “public.”

A retailer’s product specifications and a customer’s order history require different treatment. Public API documentation and a private incident report should not share the same access policy.

SaaS: Expose Answers, Protect Workspaces

Public documentation can help buyers understand compatibility, setup requirements, pricing, and product capabilities.

A reasonable starting policy is to permit search-oriented crawling of approved product pages, integration guides, and use-case content. Decide on GPTBot separately according to the company’s training preferences.

Google-Extended requires its own review because it covers specified grounding uses as well as training.

Consider a fictional observability platform. Its public integration guides could be available to search crawlers, while customer dashboards and incident data remain behind authentication.

That policy supports product discovery without treating customer information as marketing content.

Retail and Local Services: Separate Public Routes

Retailers and service businesses should identify the pages customers need to make a decision. These may include product specifications, service areas, pricing explanations, return policies, and public FAQs.

Account details, order history, private appointments, and internal systems belong in a different category.

Google explicitly warns that robots.txt should not be used to hide private information. Use enforceable access controls for those routes, regardless of your crawler preferences. 

For teams connecting crawler decisions with a broader acquisition plan, Skyram’s SEO services guide for US businesses covers the surrounding technical, content, and measurement work.

Publishers: Test the Commercial Exchange

Publishers face a different question because the information may be the product.

An AI answer can introduce the brand to new readers, but it may also satisfy a question without generating a visit. Avoid assuming that every mention creates equal value.

Divide the content library into public previews, evergreen explainers, original reporting, and premium research. Decide which categories warrant search access and which require tighter restrictions.

For example, a subscription publisher could allow public summaries while keeping full reports behind a login.

Write the hypothesis before changing access: “Public summaries should generate qualified subscription visits without exposing the premium report.”

Then measure that outcome instead of treating crawler activity as success.

Regulated Organizations: Govern Before Allowing

Healthcare and financial services should distinguish reviewed public education from sensitive records and individualized information.

Public service descriptions, locations, professional credentials, and approved educational resources may have a clear discovery purpose.

Patient records, financial accounts, private documents, and internal recommendations require security controls, not simply disallow rules.

Nonprofits and educational organizations may favor broader access when it supports their mission. They still need to check whether licensed or third-party material carries different restrictions.

The practical rule is consistent: decide by content category, ownership, and approved purpose, not only by company type.

Build the Robots.txt Policy

Before editing production rules, save the existing file and identify its owner.

Document the intended behavior for representative URLs. Review restrictions added by plugins, developers, security tools, and previous migrations.

The examples below are starting patterns. Their directory names are illustrative and must match your website. Merge them into a reviewed policy rather than replacing an existing file without testing.

Search Access With Collection Restrictions

This pattern allows the two search-oriented crawlers while blocking GPTBot and Google Extended:

text

User-agent: GPTBot

Disallow: /

 

User-agent: Google-Extended

Disallow: /

 

User-agent: OAI-SearchBot

Allow: /

Disallow: /account/

Disallow: /checkout/

Disallow: /private/

 

User-agent: PerplexityBot

Allow: /

Disallow: /account/

Disallow: /checkout/

Disallow: /private/

 

OpenAI supports independent GPTBot and OAI-SearchBot preferences. PerplexityBot is search-oriented, while Google Extended covers its documented Gemini training and grounding uses. 

Do not describe this as “block training while preserving every AI retrieval use.” Blocking Google-Extended also restricts its specified Gemini grounding uses.

Record that tradeoff in the approval documentation. The private-route exclusions express crawl preferences. They do not secure those routes. Authentication and authorization remain necessary.

Approved Public Content Collection

Some organizations approve training-oriented collection of their public marketing content. An illustrative GPTBot group could look like this:

text

User-agent: GPTBot

Allow: /

Disallow: /account/

Disallow: /checkout/

Disallow: /private/

Disallow: /research/premium/

 

This expresses a different policy from a sitewide block: approved public content is accessible, while named routes are excluded.

Create separately reviewed groups for other agents or tokens where needed.

Do not assume that approving one agent approves every possible AI use. Contracts, licensing terms, ownership rights, and other collection methods remain separate from robots.txt configuration.

Public Previews Only

A subscription business may prefer narrow allowances for search crawlers:

text

User-agent: OAI-SearchBot

Disallow: /

Allow: /public/

Allow: /previews/

 

User-agent: PerplexityBot

Disallow: /

Allow: /public/

Allow: /previews/

 

User-agent: GPTBot

Disallow: /

 

User-agent: Google-Extended

Disallow: /

 

The intended policy is straightforward: public resources and previews may be crawled, while other routes are excluded for the named search agents.

Test representative URLs against each crawler’s applicable behavior.

If previews use a different directory or URL structure, adjust the patterns before deployment. Premium content must remain protected at the application layer.

Avoid Group-Inheritance Assumptions

A subtle configuration error appears when teams add a specific agent group and assume it inherits every restriction under User-agent: *.

Google documents that a specific Googlebot group takes precedence over the global group rather than automatically inheriting all its rules. Do not assume every crawler combines groups identically. 

Whenever you add a dedicated group, review whether it includes the exclusions that agent needs.

Also inspect separately served properties. Robots.txt scope is tied to the relevant host and protocol, so a main-site policy should not be assumed to cover every documentation subdomain or other property. 

The objective is a policy your team can understand and maintain, not a growing file of exceptions with no documented owner.

Audit Access and Business Results

A robots.txt file can be correct while the crawler experience remains broken.

A CDN rule may reject an allowed agent. A security challenge may replace the page. An accidental noindex directive may prevent indexing. A canonical tag may point to a different destination.

Follow the request from policy to delivery to business outcome.

Inventory the Content

Create a URL-category inventory before selecting rules.

For each category, record:

  • Whether the content is public, restricted, or confidential.
  • Whether your organization owns or licenses it.
  • Which discovery channels matter.
  • Whether training-oriented collection is approved.
  • Whether specified Gemini grounding is approved.
  • Who owns the final decision.

This turns vague disagreement into a concrete review.

Marketing may want broader discovery while legal or product leadership prefers narrower content use. Resolve that difference before deployment.

Test Representative URLs

Build an acceptance test set covering the homepage, a service page, a product page, a blog article, an account route, and any premium material.

For every named agent, record the expected result.

Inspect the complete robots.txt file for broad disallows, duplicate groups, obsolete paths, and patterns that conflict with the approved policy.

A useful test record includes the URL, agent, intended permission, observed response, and unresolved issue. Save it with the change record.

Verify the Delivery Layer

OpenAI recommends allowing its published IP ranges for OAI-SearchBot. Perplexity recommends using its current published IP ranges alongside user-agent identification when configuring WAF access. 

Do not rely on the user-agent string alone. Google notes that HTTP user-agent strings can be spoofed. 

Verify that:

  • Requests match the platform’s documented verification method.
  • Approved public pages return usable content.
  • Security challenges do not replace the intended response.
  • Redirects resolve to the correct destination.
  • Private routes remain protected.
  • Crawler exceptions are narrowly scoped.

Allowing a verified search crawler should not mean disabling security protections across the website.

Check Google Search Eligibility

For Google AI Overviews and AI Mode, supporting pages must be indexed and eligible to appear in Google Search with a snippet. Google states that no additional technical requirements or special AI schema are required. 

Crawler access alone is therefore insufficient.

Review indexing status, canonical destinations, visible text, and applicable preview controls.

If you use noindex, Google must be able to crawl the page to read the directive. Blocking the URL in robots.txt can prevent it from discovering that instruction. 

Skyram’s Google AI Overview optimization guide offers context for the broader technical and content review. Treat Google’s official documentation as the authority on eligibility and access controls.

Review Sitemaps and Internal Links

The sitemap review and the crawler policy review should support each other.

Check sitemap entries for exact duplicate URLs, redirected destinations, noncanonical variants, unavailable pages, and pages intentionally excluded from indexing.

Separately, review editorial duplication. Two different URLs can target the same question even when neither is an exact duplicate entry.

For this topic, keep the article focused on crawler permissions and business-model decisions. A general AI visibility audit should cover query testing, competitor citations, and content gaps rather than repeat this guide.

Internal links should deliver what their anchor text promises. A link labeled “technical SEO foundation” should open that resource, not an unrelated service.

Google recommends making content discoverable through internal links as part of the foundational practices relevant to its AI search features. 

Separate Access From Performance

Use three measurement layers:

Layer What to record What it establishes
Technical access Verified requests, response codes, routes, and security challenges Whether permitted agents can retrieve approved content
Search visibility Query, platform, date, cited URL, brand mention, and factual accuracy Whether content appears in relevant answers
Business performance Identifiable referral sessions, engaged visits, qualified leads, purchases, or subscriptions Whether discovery contributes commercial value

A crawler visit is not a citation. A citation is not automatically a qualified lead.

Run the same buyer questions before and after the policy change. Record platform and geographic context, including the US market where relevant. Repeat observations instead of treating one generated answer as a stable benchmark.

Skyram’s step-by-step AI search visibility audit provides a framework for query selection, platform testing, and competitor benchmarking.

Google includes traffic from its AI features in the overall Search Console Performance report under the Web search type. Do not assume you have a complete, separate AI Overview reporting view.

Assign Owners and Review Triggers

The CTO or engineering lead should own implementation and verification. The SEO lead should own discovery requirements and measurement. Content, legal, and security stakeholders should approve relevant restrictions.

Review the policy after migrations, CDN changes, paywall launches, documentation changes, and material updates to platform documentation.

OpenAI says search-related robots.txt changes can take approximately 24 hours to reach its systems. Perplexity says changes may take up to 24 hours. These are processing expectations, not promises that citations or revenue will change within a day.

Skyram Technologies’ AI visibility audit provides a starting point for assessing how a brand appears across AI platforms. Its answer engine optimization services connect that assessment with content structure, entity consistency,  and visibility measurement.

The useful deliverable is an approved policy, verified implementation, and measurable discovery plan. The goal is not maximum bot access. It is useful discovery without unnecessary exposure.

Frequently Asked Questions

Should I Block GPTBot?

Block GPTBot if your organization does not approve OpenAI’s training-oriented collection of your content. You can separately allow OAI-SearchBot for ChatGPT search discovery. Allowing GPTBot is not a documented requirement for appearing in ChatGPT searches.

Does Blocking GPTBot Affect Google Rankings?

GPTBot controls OpenAI’s training-oriented crawling, not Google Search crawling. Keep Googlebot access and indexing controls separate. Google also states that Google Extended does not affect search inclusion and is not a search ranking signal.

Should I Allow PerplexityBot?

Allow PerplexityBot on approved public pages when Perplexity search discovery supports your business goals. Perplexity identifies it as a search crawler, not a foundation-model training crawler. Publishers and proprietary-data businesses may prefer selective access instead of a sitewide allowance.

Does Google-Extended Control AI Overviews?

No. Google Extended Controls specified Gemini training and grounding uses. Google AI Overviews and AI Mode are Search features governed by Googlebot access and applicable indexing or preview controls. Blocking Google-Extended does not remove ordinary Google Search eligibility.

Can I Allow Search but Block Training?

Yes. For OpenAI, you can allow OAI-SearchBot while blocking GPTBot. PerplexityBot is also documented as search-oriented rather than a foundation-model training crawler. Evaluate Google-Extended separately because its scope includes specified grounding uses as well as training.

Do AI Agents Always Follow Robots.txt?

No single rule covers every agent. OpenAI says robots.txt may not apply to user-initiated ChatGPT-User actions. Perplexity says Perplexity-User generally ignores robots.txt for user-requested fetches. Use enforceable access controls for information that must remain private.

Will Allowing Crawlers Guarantee Citations?

No. Allowing the relevant crawler removes one potential access obstacle, but it does not guarantee selection. Google explicitly states that meeting its requirements does not guarantee crawling, indexing, or serving. Measure relevant queries, citation accuracy, and business outcomes rather than treating permission as performance.

Can Robots.txt Remove Previously Collected Content?

Do not treat a new disallow rule as a deletion request or retroactive removal guarantee. The documented controls express crawler or content-use preferences within each platform’s stated scope. Questions about previously collected content require a separate review of applicable platform processes and policies.

How Do I Protect Private Content?

Use authentication, authorization, and appropriate application or infrastructure controls. Robots.txt is a crawl-preference mechanism, not a security boundary. Google warns against using it to hide private information, and user-triggered AI fetchers may not follow its directives.

What Is the Best Default Policy?

For a public lead-generation site that rejects training-oriented collection, start by allowing OAI-SearchBot and PerplexityBot on approved public pages while blocking GPTBot. Decide Google-Extended separately, document its Gemini grounding tradeoff, and verify that private routes remain protected by enforceable controls.

Do you want more traffic?

Our team at Skyram Technologies is ready to make a business grow. Our only question is, do you want it too?