# How to prevent ChatGPT from crawling your website

**URL:** <https://forums.realmacsoftware.com/t/how-to-prevent-chatgpt-from-crawling-your-website/42591>\
**Category:** RapidWeaver Classic\
**Tags:** ai, chatgpt\
**Created:** [August 15, 2023, 1:13pm UTC](https://forums.realmacsoftware.com/t/how-to-prevent-chatgpt-from-crawling-your-website/42591 "2023-08-15T13:13:00Z")\
**Posts on this page:** 17\
**Page:** 1

<div class="post-metadata">

**Author:** ![differentdan](https://dub1.discourse-cdn.com/flex005/user_avatar/forums.realmacsoftware.com/differentdan/32/33530_2.png) [@differentdan](https://forums.realmacsoftware.com/u/differentdan)\
**Post date:** [August 15, 2023, 1:13pm UTC](https://forums.realmacsoftware.com/t/how-to-prevent-chatgpt-from-crawling-your-website/42591/1 "2023-08-15T13:13:00Z")

</div>

Came across an interesting article and thought I’d share it. ChatGPT is pretty neat and useful for a lot of folks, but there are also those who don’t take too kindly to it scraping their data to train their LLMs (large language models). For example those who publish copyrighted works online, or those that have to pay the server costs for the bandwidth that ChatGPT consumes when it’s crawling their site. There are even those that just don’t want to contribute to the potential future AI uprising. 🤖

If you fall into the **“ChatGPT = Bad”** camp, below are two methods you can use to prevent our future AI overlord from assimilating your website’s data.

![Season 2 Borg GIF by Paramount+](https://europe1.discourse-cdn.com/flex005/uploads/realmacsoftware1/original/3X/3/5/35c1b665d854060213406521a41c7c547d0e5783.gif)

* * *

## Block the ChatGPT Bot via robots.txt

1. At your web host, locate your website’s root directory (usually **public\_html** if your web host is using cPanel) and create a new file in that directory called **robots.txt**. Open that **robots.txt** file by clicking **Edit**.

 ![no-chatgpt-robots-1](https://europe1.discourse-cdn.com/flex005/uploads/realmacsoftware1/original/3X/8/0/803633a2b1c2339c6988d7b2d26fdf97898c5923.jpeg)

1. Enter the below rules to block ChatGPT from accessing **all areas** of your website, then click the “ **Save Changes** ” button, then the “ **Close** ” button.

```auto
User-agent: GPTBot
Disallow: /

```

 ![no-chatgpt-robots-2](https://europe1.discourse-cdn.com/flex005/uploads/realmacsoftware1/original/3X/d/9/d907e757266f2de3ab5ac2ee1fc972af801c94f4.jpeg)

1. If you would like to prevent ChatGPT from accessing only certain parts of your site, you can selectively list what directories/folders it can and cannot crawl by entering the below rules (replacing **directory-\*/** with the **actual path to your directory** ), then click the “ **Save Changes** ” button, then the “ **Close** ” button.

```auto
User-agent: GPTBot
Allow: /directory-1/
Disallow: /directory-2/

```

 ![no-chatgpt-robots-3](https://europe1.discourse-cdn.com/flex005/uploads/realmacsoftware1/original/3X/9/4/94dacee79437e310db0c7819d8eb1b13f5a6e592.jpeg)

* * *

## Block the ChatGPT bot via .htaccess

1. In RapidWeaver, go to your Publishing settings, then click on the “ **Edit .htaccess File** ” button.

 ![no-chatgpt-1](https://europe1.discourse-cdn.com/flex005/uploads/realmacsoftware1/original/3X/5/1/5114609a5dc20738f211e45a70684a7313169c95.png)

1. On a new line, enter the below .htaccess rules to block ChatGPT from accessing **all areas** of your website, then click the “ **Save and Upload** ” button.

```auto
# Apache 2.2
<IfModule !authz_core_module>
    Order Allow,Deny
    Allow from all
    Deny from 52.230.152.0/24
    Deny from 52.233.106.0/24
</IfModule>

# Apache 2.4+
<IfModule authz_core_module>
    <RequireAll>
        Require all granted
        Require not ip 52.230.152.0/24
        Require not ip 52.233.106.0/24
    </RequireAll>
</IfModule>

```

 ![htaccess-rules-to-block-chatgpt](https://europe1.discourse-cdn.com/flex005/uploads/realmacsoftware1/original/3X/f/9/f9f84b3cd9c307ace4cbd4d6e02c75b144dfd7c8.jpeg)

1. Another possible way to block the ChatGPT Bot via .htaccess is provided below.

```auto
# Apache 2.2
<IfModule !authz_core_module>
    SetEnvIf User-Agent GPTBot NoChatGPT=1
    Order Allow,Deny
    Allow from all
    Deny from env=NoChatGPT
</IfModule>

# Apache 2.4+
<IfModule authz_core_module>
    <If "%{HTTP_USER_AGENT} == 'GPTBot'">
        Require all denied
    </If>
</IfModule>

```

 ![htaccess-rules-to-block-chatgpt-alt](https://europe1.discourse-cdn.com/flex005/uploads/realmacsoftware1/original/3X/0/7/07ce678efec79ffa56f439b4b03d09bf95942469.jpeg)

* * *

**That’s it!**

You’ve now blocked the ChatGPT bot from crawling your website. More information can be found on OpenAI’s website [here](https://platform.openai.com/docs/gptbot).

Hope that helps…humanity. 😰

---

<div class="post-metadata">

**Author:** ![jacksona](https://dub1.discourse-cdn.com/flex005/user_avatar/forums.realmacsoftware.com/jacksona/32/16622_2.png) [@jacksona](https://forums.realmacsoftware.com/u/jacksona)\
**Post date:** [August 16, 2023, 10:35am UTC](https://forums.realmacsoftware.com/t/how-to-prevent-chatgpt-from-crawling-your-website/42591/2 "2023-08-16T10:35:14Z")

</div>

Can’t remember where I picked this up from, but this _robots.txt_ prevents many of the other bots as well.

```auto
User-agent: CCBot
Disallow: /

User-agent: ChatGPT-User
Disallow: /

User-agent: GPTBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: anthropic-ai
Disallow: /

User-agent: Omgilibot
Disallow: /

User-agent: Omgili
Disallow: /

User-agent: FacebookBot
Disallow: /

User-agent: Bytespider
Disallow: /

```

Might be helpful to someone!

---

<div class="post-metadata">

**Author:** ![differentdan](https://dub1.discourse-cdn.com/flex005/user_avatar/forums.realmacsoftware.com/differentdan/32/33530_2.png) [@differentdan](https://forums.realmacsoftware.com/u/differentdan)\
**Post date:** [August 16, 2023, 10:40am UTC](https://forums.realmacsoftware.com/t/how-to-prevent-chatgpt-from-crawling-your-website/42591/3 "2023-08-16T10:40:59Z")

</div>

Yeah I saw that over on Adam’s forum yesterday, someone posted it.

Thanks for adding it here, it’s a great resource for excluding more AI bots from crawling a website. 🙂 🙌

One thing in that above robots.txt, GPTBot covers ChatGPT-User as well, so only the GPTBot needs to be added to the robots.txt file and it would block both.

---

<div class="post-metadata">

**Author:** ![zamboknee](https://dub1.discourse-cdn.com/flex005/user_avatar/forums.realmacsoftware.com/zamboknee/32/25532_2.png) [@zamboknee](https://forums.realmacsoftware.com/u/zamboknee)\
**Post date:** [August 16, 2023, 6:00pm UTC](https://forums.realmacsoftware.com/t/how-to-prevent-chatgpt-from-crawling-your-website/42591/4 "2023-08-16T18:00:22Z")

</div>

Awesome post, dan. Thanks so much.

---

<div class="post-metadata">

**Author:** ![differentdan](https://dub1.discourse-cdn.com/flex005/user_avatar/forums.realmacsoftware.com/differentdan/32/33530_2.png) [@differentdan](https://forums.realmacsoftware.com/u/differentdan)\
**Post date:** [August 17, 2023, 11:48am UTC](https://forums.realmacsoftware.com/t/how-to-prevent-chatgpt-from-crawling-your-website/42591/5 "2023-08-17T11:48:18Z")

</div>

No problem.

If anybody knows of other AI bots out there and ways to block them, please add them in the comments here.

OpenAI has changed their IP egress ranges three times since I originally wrote this article, so it’s a constant battle of staying up to date to make sure these AI bots are properly blocked from crawling people’s websites. 😮‍💨

---

<div class="post-metadata">

**Author:** ![Flash](https://avatars.discourse-cdn.com/v4/letter/f/f475e1/32.png) [@Flash](https://forums.realmacsoftware.com/u/Flash)\
**Post date:** [August 17, 2023, 3:13pm UTC](https://forums.realmacsoftware.com/t/how-to-prevent-chatgpt-from-crawling-your-website/42591/6 "2023-08-17T15:13:14Z")

</div>

I was looking up this data yesterday. What I have found hundreds of bots. And this is growing everyday it seems.

---

<div class="post-metadata">

**Author:** ![differentdan](https://dub1.discourse-cdn.com/flex005/user_avatar/forums.realmacsoftware.com/differentdan/32/33530_2.png) [@differentdan](https://forums.realmacsoftware.com/u/differentdan)\
**Post date:** [April 12, 2024, 1:22pm UTC](https://forums.realmacsoftware.com/t/how-to-prevent-chatgpt-from-crawling-your-website/42591/7 "2024-04-12T13:22:12Z")

</div>

(April 12th, 2024) - Updated list of AI bots to add to your robots.txt file to direct them not to crawl/scrape your website.

* * *

```auto
# Amazon Bot - enabling Alexa to answer even more questions for customers.
User-agent: Amazonbot
Disallow: /

# Anthropic AI Bot
User-agent: anthropic-ai
Disallow: /

# Apple Bot - collects website data for its Siri and Spotlight services.
User-agent: Applebot
Disallow: /

# Claude Bot run by Anthropic
User-agent: Claude-Web
Disallow: /

# Cohere AI Bot - unconfirmed bot believed to be associated with Cohere’s chatbot.
User-agent: cohere-ai
Disallow: /

# Common Crawl's bot - Common Crawl is one of the largest public datasets used by AI for training, with ChatGPT, Bard and other large language models.
User-agent: CCBot
Disallow: /

# Diffbot - somewhat dishonest scraping bot used to collect data to train LLMs.
User-agent: Diffbot
Disallow: /

# Google Bard and VertexAI. This will not have an impact on Google Search indexing. This will not affect GoogleBot crawling.
User-agent: Google-Extended
Disallow: /

# ImagesiftBot is billed as a reverse image search tool, but it's associated with The Hive, a company that produces models for image generation.
User-agent: ImagesiftBot 
Disallow: /

# KUKA's youBot
User-agent: YouBot
Disallow: /

# OMGilibot - They sell data for training LLMs (large language models)
User-agent: omgilibot
Disallow: /

# Omgili (Oh My God I Love It)
User-agent: omgili
Disallow: /

# OpenAI API - bot that OpenAI specifically uses to collect bulk training data from your website for ChatGPT.
User-agent: GPTBot
Disallow: /

# Perplexity AI
User-agent: PerplexityBot
Disallow: /

## Social Media Bots

# Bytespider is a web crawler operated by ByteDance, the Chinese owner of TikTok
User-agent: Bytespider
Disallow: /

# Meta’s bot that crawls public web pages to improve language models for their speech recognition technology
User-agent: FacebookBot
Disallow: /

#Twitter's bot used to index the content of any given URL
User-agent: Twitterbot
Disallow: /

```

---

<div class="post-metadata">

**Author:** ![dan](https://dub1.discourse-cdn.com/flex005/user_avatar/forums.realmacsoftware.com/dan/32/28516_2.png) [@dan](https://forums.realmacsoftware.com/u/dan)\
**Post date:** [June 11, 2024, 8:25am UTC](https://forums.realmacsoftware.com/t/how-to-prevent-chatgpt-from-crawling-your-website/42591/8 "2024-06-11T08:25:06Z")

</div>

There’s a good GitHub project that keeps an updated listed of AI crawlers, worth keeping an eye on if you want to keep away those pesky robots!

> <https://github.com/ai-robots-txt/ai.robots.txt/blob/main/table-of-bot-metrics.md>

---

<div class="post-metadata">

**Author:** ![Heroic\_Nonsense](https://dub1.discourse-cdn.com/flex005/user_avatar/forums.realmacsoftware.com/heroic_nonsense/32/20666_2.png) [@Heroic\_Nonsense](https://forums.realmacsoftware.com/u/Heroic_Nonsense)\
**Post date:** [June 20, 2024, 2:37pm UTC](https://forums.realmacsoftware.com/t/how-to-prevent-chatgpt-from-crawling-your-website/42591/9 "2024-06-20T14:37:52Z")

</div>

Ugh, I hope openAI is more sincere than Perplexity is! Perplexity (an AI search engine) outright ignores your robots.txt, Robb Knight (a researcher) and WIRED Magazine have found.

> **[Perplexity AI Is Lying about Their User Agent](https://rknight.me/blog/perplexity-ai-is-lying-about-its-user-agent/)**
>
> Perplexity AI claims it sends a user agent and respects robots.txt but it absolutely does not

> **[Perplexity Is a Bullshit Machine](https://www.wired.com/story/perplexity-is-a-bullshit-machine/)**
>
> A WIRED investigation shows that the AI-powered search startup Forbes has accused of stealing its content is surreptitiously scraping—and making things up out of thin air.

Cheers,  
Erwin

---

<div class="post-metadata">

**Author:** ![Heroic\_Nonsense](https://dub1.discourse-cdn.com/flex005/user_avatar/forums.realmacsoftware.com/heroic_nonsense/32/20666_2.png) [@Heroic\_Nonsense](https://forums.realmacsoftware.com/u/Heroic_Nonsense)\
**Post date:** [June 25, 2024, 2:09pm UTC](https://forums.realmacsoftware.com/t/how-to-prevent-chatgpt-from-crawling-your-website/42591/10 "2024-06-25T14:09:27Z")

</div>

Well, apparently OpenAI _also_ ignores robots.txt:

> **[OpenAI and Anthropic are ignoring an established rule that prevents bots...](https://www.businessinsider.com/openai-anthropic-ai-ignore-rule-scraping-web-contect-robotstxt?international=true&r=US&IR=T)**
>
> OpenAI and Anthropic have said publicly they respect robots.txt. But they are among the biggest tech companies ignoring the rule, BI has learned.

---

<div class="post-metadata">

**Author:** ![differentdan](https://dub1.discourse-cdn.com/flex005/user_avatar/forums.realmacsoftware.com/differentdan/32/33530_2.png) [@differentdan](https://forums.realmacsoftware.com/u/differentdan)\
**Post date:** [June 25, 2024, 2:31pm UTC](https://forums.realmacsoftware.com/t/how-to-prevent-chatgpt-from-crawling-your-website/42591/11 "2024-06-25T14:31:57Z")

</div>

I don’t doubt that there is some shady scraping of data going on among the different AI companies, however that BI article isn’t listing its sources yet.

> OpenAI and Anthropic have been found to be either ignoring or circumventing an established web rule, called robots.txt, that prevents automated scraping of websites, **according to a person with knowledge of the analytics of TollBit, as well as another person familiar with the matter.**

The article does reference an [earlier article from Reuters here](https://www.reuters.com/technology/artificial-intelligence/multiple-ai-companies-bypassing-web-standard-scrape-publisher-sites-licensing-2024-06-21/), in which TollBit states:

> According to the TollBit letter, Perplexity is not the only offender that appears to be ignoring robots.txt. TollBit said its analytics indicate “numerous” AI agents are bypassing the protocol, a standard tool used by publishers to indicate which parts of its site can be crawled.

It doesn’t specifically mention which AI companies they are referring to in the letter.

With that said, it wouldn’t surprise me if it turns out OpenAI is ignoring the robots.txt file, but I’d need a bit more to go on aside from that BI article. Let’s see if OpenAI publicly addresses it in the coming days.

In the meantime, perhaps blocking via the .htaccess method would be more ironclad.

---

<div class="post-metadata">

**Author:** ![differentdan](https://dub1.discourse-cdn.com/flex005/user_avatar/forums.realmacsoftware.com/differentdan/32/33530_2.png) [@differentdan](https://forums.realmacsoftware.com/u/differentdan)\
**Post date:** [June 25, 2024, 2:42pm UTC](https://forums.realmacsoftware.com/t/how-to-prevent-chatgpt-from-crawling-your-website/42591/12 "2024-06-25T14:42:26Z")

</div>

Another blocking alternative, for any website that runs behind CloudFlare (which is recommended), they offer AI bot management on all of their plans, including the free one according to their below article.

> **[Easily manage AI crawlers with our new bot categories](https://blog.cloudflare.com/ai-bots/)**
>
> Manage AI crawlers, out of the box with Cloudflare

---

<div class="post-metadata">

**Author:** ![Bruno](https://dub1.discourse-cdn.com/flex005/user_avatar/forums.realmacsoftware.com/bruno/32/25457_2.png) [@Bruno](https://forums.realmacsoftware.com/u/Bruno)\
**Post date:** [June 25, 2024, 3:02pm UTC](https://forums.realmacsoftware.com/t/how-to-prevent-chatgpt-from-crawling-your-website/42591/13 "2024-06-25T15:02:06Z")

</div>

Hi, I continue to be surprised by this desire for confidentiality when the online posting is obviously public. Remember back in the 90s all those programs to suck up entire websites under the guise of saving the prohibitive connection fees at the time. From forum to forum I read posts asking how to protect your photos (which you put online voluntarily), how to protect your documents (which you put online voluntarily), how to protect your music (which you… . yes, you understand). It’s like movie stars, as long as they’re not yet, they absolutely want to be known and recognized, when they finally are, they put on dark glasses… I know my psychological side… I can’t do anything 🥱. I believe that the real desire behind these posts is: “how can we only give to those to whom we want to give while showing everyone?” Put differently: “how can we ensure that it only benefits those for whom we accept that it will benefit.” It reminds me of a joke about perverts: who, the sadist and the masochist, will win if they play together? The masochist says “go ahead and hurt me.” The sadist replies: “if I want.” 😬

---

<div class="post-metadata">

**Author:** ![differentdan](https://dub1.discourse-cdn.com/flex005/user_avatar/forums.realmacsoftware.com/differentdan/32/33530_2.png) [@differentdan](https://forums.realmacsoftware.com/u/differentdan)\
**Post date:** [June 25, 2024, 3:07pm UTC](https://forums.realmacsoftware.com/t/how-to-prevent-chatgpt-from-crawling-your-website/42591/14 "2024-06-25T15:07:24Z")

</div>

I think it has more to do with people wanting to protect their intellectual property and copyrighted works, not concerns about privacy or confidentiality.

---

<div class="post-metadata">

**Author:** ![Flash](https://avatars.discourse-cdn.com/v4/letter/f/f475e1/32.png) [@Flash](https://forums.realmacsoftware.com/u/Flash)\
**Post date:** [June 25, 2024, 3:10pm UTC](https://forums.realmacsoftware.com/t/how-to-prevent-chatgpt-from-crawling-your-website/42591/15 "2024-06-25T15:10:39Z")

</div>

> [@differentdan](#):
>
> I think it has more to do with people wanting to protect their intellectual property and copyrighted works, not concerns about privacy or confidentiality.

Yeah that, exactly!

---

<div class="post-metadata">

**Author:** ![Bruno](https://dub1.discourse-cdn.com/flex005/user_avatar/forums.realmacsoftware.com/bruno/32/25457_2.png) [@Bruno](https://forums.realmacsoftware.com/u/Bruno)\
**Post date:** [June 25, 2024, 3:15pm UTC](https://forums.realmacsoftware.com/t/how-to-prevent-chatgpt-from-crawling-your-website/42591/16 "2024-06-25T15:15:01Z")

</div>

In first intention I agree, but how to design a thought lock, graphic representations or others? When one has indicated “all rights reserved” and the other does not care, all that remains is justice… and its cost.

---

<div class="post-metadata">

**Author:** ![Heroic\_Nonsense](https://dub1.discourse-cdn.com/flex005/user_avatar/forums.realmacsoftware.com/heroic_nonsense/32/20666_2.png) [@Heroic\_Nonsense](https://forums.realmacsoftware.com/u/Heroic_Nonsense)\
**Post date:** [July 4, 2024, 12:35pm UTC](https://forums.realmacsoftware.com/t/how-to-prevent-chatgpt-from-crawling-your-website/42591/17 "2024-07-04T12:35:03Z")

</div>

Cloudflare have just announced a tool that helps you prevent AI-bots from scraping your website.

> **[Declare your AIndependence: block AI bots, scrapers and crawlers with a...](https://blog.cloudflare.com/declaring-your-aindependence-block-ai-bots-scrapers-and-crawlers-with-a-single-click)**
>
> To help preserve a safe Internet for content creators, we’ve just launched a brand new “easy button” to block all AI bots. It’s available for all customers, including those on our free tier.

Cheers,  
Erwin
