Skip to main content

How to Block AI Crawlers with robots.txt

Protect your website from AI crawlers: Complete guide to blocking ChatGPT, Claude & Co. with robots.txt. Simply explained with examples.

FHFinn Hillebrandt
AI Application
How to Block AI Crawlers with robots.txt
Links marked with * are affiliate links. If a purchase is made through such links, we receive a commission.

As an online entrepreneur or blogger, you're facing a new challenge:

Web crawlers from OpenAI, Anthropic, or Google are searching the web and collecting training data for LLMs and other AI models.

Your valuable blog posts that you created with great effort could be used without your knowledge or consent to generate AI-generated texts in ChatGPT & Co.

This can not only violate your copyrights but also jeopardize your competitive position. Somewhat unsettling, isn't it?

Perhaps you're already asking yourself: How can I protect my work? How do I prevent my content from being used for AI training without my consent?

No problem!

In this article, I'll show you simply and step by step how to configure your robots.txt to protect your content.

TL;DRKey Takeaways
  • Create or edit your robots.txt to block specific AI crawlers like GPTBot, ClaudeBot, Claude-User, Claude-SearchBot, Google-Extended, and Applebot-Extended
  • Use selective blocking to protect only certain areas while keeping others accessible
  • Test your configuration with Google Search Console and pair robots.txt with a WAF (e.g., Cloudflare AI bot blocking), since Grok and Perplexity-User partly ignore robots.txt

1. Preparation

Before we get started protecting your website from curious AI crawlers, you need to make a few preparations. Don't worry, it's easier than you might think!

Access to the web server

First, you need access to your web server. This sounds technical but is often just a login to your hosting account.

If you're using WordPress, you can access your files directly via FTP or the File Manager Plugin.

Backup of existing robots.txt

Safety first! If you already have a robots.txt file, make sure to create a copy. This way you can always revert to the old version in case of emergency:

  • Find the robots.txt file in your website's root directory
  • Download it to your computer or copy the content into a text document
  • Store this backup in a safe place

2. Creating/Editing robots.txt

You don't need to be a programming genius to create or edit your robots.txt file.

Only a few steps are required:

2.1 Opening or Creating the File

First, you need to check if a robots.txt already exists on your website. There's a simple trick for this:

  1. Open your browser
  2. Enter your domain followed by "/robots.txt" (e.g., www.yourwebsite.com/robots.txt)
  3. Do you see text? Great, the file already exists. If not, we'll create a new one.

If you need to create a new file:

  • Open a simple text editor (Notepad, TextEdit, etc.)
  • Create a new, empty document
  • Save it as "robots.txt" (Note: don't add a file extension like .txt!)

2.2 Setting Up the Basic Structure

The robots.txt follows a specific syntax (structure). Here are the basics:

User-agent: [Name of the crawler]
Disallow: [Path to be blocked]

For starters, you could write something like this:

User-agent: *
Disallow:

This means: All crawlers (*) may crawl everything (empty "Disallow"). This is our starting point from which we'll further customize the file.

3. Blocking Specific AI Crawlers

To block common AI crawlers, you need to add the following blocks to your robots.txt:

OpenAI (ChatGPT)

OpenAI has a total of three different crawlers that serve different functions. To prevent content theft as effectively as possible, you should exclude all of them:

User-agent: OAI-SearchBot
Disallow: /

User-agent: ChatGPT-User
Disallow: /

User-agent: GPTBot
Disallow: /

Anthropic (Claude)

Since early 2026, Anthropic officially documents four bots: ClaudeBot (model training), Claude-User (live fetches when a Claude user asks a question), Claude-SearchBot (indexing for in-Claude web search), and claude-code (Claude Code CLI). The legacy user agents Claude-Web and anthropic-ai are officially deprecated, but many crawlers still send them. Block all of them to be safe:

User-agent: ClaudeBot
Disallow: /

User-agent: Claude-User
Disallow: /

User-agent: Claude-SearchBot
Disallow: /

User-agent: claude-code
Disallow: /

User-agent: anthropic-ai
Disallow: /

User-agent: Claude-Web
Disallow: /

Google (Gemini)

User-agent: Google-Extended
Disallow: /

Common Crawl

User-agent: CCBot
Disallow: /

Perplexity

Perplexity ships two bots: PerplexityBot for indexing and Perplexity-User for user-triggered fetches. Note that Perplexity explicitly says Perplexity-User ignores robots.txt because it counts as "user-initiated." An additional WAF or firewall rule is often a good idea here.

User-agent: PerplexityBot
Disallow: /

User-agent: Perplexity-User
Disallow: /

Meta AI / Facebook

User-agent: FacebookBot
Disallow: /

User-agent: Meta-ExternalAgent
Disallow: /

User-agent: Meta-ExternalFetcher
Disallow: /

Webz.io

User-agent: OmgiliBot
Disallow: /

Cohere

User-agent: cohere-ai
Disallow: /

xAI/SpaceXAI (Grok)

xAI has been part of SpaceX since February 2026 and now operates as SpaceXAI, but it still officially documents three user agents (GrokBot/1.0, xAI-Grok/1.0, and Grok-DeepSearch/1.0). In practice, however, Grok often masquerades as a regular browser or as Go-http-client/1.1, so robots.txt rules often miss it. Add the official strings anyway, but back them up with a firewall rule or Cloudflare AI bot blocking:

User-agent: GrokBot
Disallow: /

User-agent: xAI-Grok
Disallow: /

User-agent: Grok-DeepSearch
Disallow: /

Apple (Apple Intelligence)

Applebot-Extended controls exclusively whether Apple may use your content to train Apple Intelligence. The regular Applebot for Spotlight and Siri Suggestions is unaffected. With the Apple Intelligence rollout, Applebot has risen to one of the most active AI crawlers in 2026.

User-agent: Applebot-Extended
Disallow: /

Amazon

Amazonbot collects content for Alexa, Echo, and Amazon's internal AI shopping assistants. Amazon documents IP ranges cleanly and respects robots.txt reliably.

User-agent: Amazonbot
Disallow: /

4. Selective Blocking

Sometimes you don't want to completely lock out AI crawlers, but only protect certain areas of your website.

No problem!

Blocking specific directories/pages for AI crawlers

If you have an area with exclusive content, you can exclude this from crawlers with the following code:

User-agent: GPTBot
Disallow: /exclusive/

User-agent: anthropic-ai
Disallow: /premium-content/

In this example, you're blocking GPTBot from your "/exclusive/" directory and Anthropic's crawler from "/premium-content/".

Defining exceptions

Sometimes you might want to block most of your site but make certain areas accessible to AI crawlers. Here's an example:

User-agent: GPTBot
Disallow: /
Allow: /blog/

User-agent: anthropic-ai
Disallow: /
Allow: /public/

In this case, you first block everything with Disallow: /, then allow specific areas with Allow.

So GPTBot is allowed to crawl your blog, while Anthropic's crawler can only access the public area.

5. Verification and Testing

Everything set up? Great!

But before you sit back, you should make sure your robots.txt is actually doing what it's supposed to.

Google provides you with a great tool for this: The robots.txt Tester in Google Search Console.

robots.txt Tester in Google Search Console

Here you can see if your robots.txt can be properly fetched by Google and if it contains any errors.

Frequently Asked Questions About Blocking AI Crawlers

Effectiveness varies significantly: Large companies like OpenAI and Anthropic generally respect robots.txt, but there's no legal obligation to do so. Studies show only 7.8% of top websites block GPTBot. The crawl-to-referral ratio is problematic: While Google crawls at 14:1 and brings traffic back, OpenAI's ratio is 1,700:1 and Anthropic's is even 73,000:1. Many smaller or less transparent AI crawlers completely ignore robots.txt. For effective protection, a combination of robots.txt, firewall rules, and legal measures is recommended.

The legal landscape continues to change rapidly:

  • EU AI Act: Explicitly requires AI providers to respect robots.txt (Sub-Measure 4.1)
  • Copyright: Your content is generally protected, but enforcement is complex
  • TDM Directive: EU law permits text and data mining, but you can opt out
  • Terms of Service: Explicit prohibition of AI crawling in your terms

Precedent cases like the NYT lawsuit against OpenAI focus on copyright violations, not robots.txt violations.

This is a strategic decision. Common Crawl itself doesn't use the data for AI, but provides monthly snapshots that anyone can download - including AI companies. CCBot is the second most blocked crawler after GPTBot. Advantage of blocking: Your content won't end up in public datasets. Disadvantage: Legitimate research projects and smaller developers also use Common Crawl. If you want maximum protection, block CCBot. If you want to support open-source projects, allow it.

Robots.txt alone is often not enough. Consider these additional measures:

  • WAF/Firewall: Cloudflare offers special AI bot blocking, but beware of false positives (3.31% of legitimate users affected)
  • Noindex meta tags: Prevents indexing, but note: Don't combine with robots.txt!
  • IP blocking: Many AI companies publish their IP ranges
  • Password protection: For particularly sensitive content
  • License agreements: Like Associated Press with OpenAI

The most effective strategy is a multi-layered approach.

There are several standardization efforts, but adoption is slow. The W3C is working on the 'Text and Data Mining Reservation Protocol' for fine-grained control. Google is experimenting with extensions to robots.txt. The ai.txt and llms.txt approaches try to distinguish media and usage types and map to EU directives. Problem: None of these standards are widely adopted in 2026 - neither by web servers nor crawlers. The rapid development of new AI crawlers (Anthropic's Claude-SearchBot and claude-code, Applebot-Extended, xAI/SpaceXAI Grok) makes it hard to keep up. In the medium term, standards could help; in the short term, robots.txt remains the most important tool, ideally paired with a WAF.

Yes, there are potential disadvantages. ChatGPT and other AI assistants are increasingly being used as information sources and could partially replace traditional search engines. If you block AI crawlers, your content won't appear in AI-generated answers - this could cost you traffic. On the other hand: The current crawl-to-referral ratio is extremely unbalanced (1,700:1 for OpenAI). Recommendation: Block selectively - allow access to marketing content or product descriptions, but protect exclusive content, guides, or creative works. This way you balance visibility and protection.
FH

Finn Hillebrandt

AI Expert & Blogger

Finn Hillebrandt is the founder of Gradually AI, an SEO and AI expert. He helps online entrepreneurs simplify and automate their processes and marketing with AI. Finn shares his knowledge here on the blog in 50+ articles as well as through the AI Business Club.

Learn more about Finn and the team, follow Finn on LinkedIn, join his Facebook group for ChatGPT, OpenAI & AI Tools or do like 17,500+ others and subscribe to his AI Newsletter with tips, news and offers about AI tools and online business. Also visit his other blog, Blogmojo, which is about WordPress, blogging and SEO.