# Google internal docs on search engine algorithm have leaked

**URL:** <https://forums.realmacsoftware.com/t/google-internal-docs-on-search-engine-algorithm-have-leaked/43359>\
**Category:** Lounge\
**Created:** [May 30, 2024, 7:30am UTC](https://forums.realmacsoftware.com/t/google-internal-docs-on-search-engine-algorithm-have-leaked/43359 "2024-05-30T07:30:26Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![Heroic\_Nonsense](https://dub1.discourse-cdn.com/flex005/user_avatar/forums.realmacsoftware.com/heroic_nonsense/32/20666_2.png) [@Heroic\_Nonsense](https://forums.realmacsoftware.com/u/Heroic_Nonsense)\
**Post date:** [May 30, 2024, 7:30am UTC](https://forums.realmacsoftware.com/t/google-internal-docs-on-search-engine-algorithm-have-leaked/43359/1 "2024-05-30T07:30:26Z")

</div>

Google’s internal documentation on how their search engine works, ave been leaked online. The documents are causing quite a stir, as they contradict a few statements Google has made in the past about the way they “rank” sites in the search results.

For example: subdomains _are_ treated as separate websites by the crawler and _do_ impact your site’s ranking. Also, newly registered domains _are_ penalised and ranked lower than more established ones. Domains that mimic search queries (like [mens-luxury-watches.com](http://mens-luxury-watches.com) or [big-purple-car.com](http://big-purple-car.com)) are also penalised.

One particularly troubling reveal is that Google ranks results by user centric click signals. In other words: heavily advertised websites and websites that generate a lot of traffic _are_ ranked higher, contrary to what Google has stated numerous times in the past. This means that sites from big brands _always_ show up higher in the list, no matter how hard a smaller brand’s tries to optimise their website or has authority. As **[today’s episode of Techlinked](https://www.youtube.com/watch?v=mlKax2El2Cg)** puts it: _“It’s like Mr. Beast showing up higher in the search results than your own doctor, just because Mr. Beast has more Twitter followers.”_

Google has in the past also stated that data from Chrome (their browser) is in no way used to determin a page’s ranking. Guess what? They _do_ use that data - including from people who have opted out from allowing Google to use browsing data.

The documents were accidentally leaked by a Google employee, who included them in his Github repro. He tried to delete the docs soon after, but they’d already been backed up by third part document management services, who still have them available.

The leak is quite big news - SEO specialists and website owners need to read them, as you can read exactly how to get traffic to your site by reading between the lines of the documents.

I’m not linking to the docs themselves, as I’m not certin of the legal implications of doing so for both me and Realmac (as the owner of this platform), but you can ( **oh I love irony** ) find them using Google Search just fine…

Google has confirmed the authenticity of the documents to **[The Verge](https://www.theverge.com/2024/5/29/24167407/google-search-algorithm-documents-leak-confirmation)**, but followed that with a statement that one should not jump to conclusions as “_the pages are out of context_”. Just how 2,500 pages of text can lack context was not explained though. Google also states that the documents are for an older version of their systems.

Cheers,  
Erwin

Sources:

> **[Google confirms the leaked Search documents are real](https://www.theverge.com/2024/5/29/24167407/google-search-algorithm-documents-leak-confirmation)**
>
> The confirmation comes after Google refused to comment.

> **[What’s going on with the Google Search leak?](https://www.siliconrepublic.com/business/what-is-google-search-leak-algorithm-api-seo)**
>
> A massive batch of documents were leaked online and appear to contain internal details and secrets about how Google Search operates.

> **[An Anonymous Source Shared Thousands of Leaked Google Search API Documents...](https://sparktoro.com/blog/an-anonymous-source-shared-thousands-of-leaked-google-search-api-documents-with-me-everyone-in-seo-should-see-them/)**
>
> On Sunday, May 5th, I received an email from a person claiming to have access to a massive leak of API documentation from inside Google’s Search division.

---

<div class="post-metadata">

**Author:** ![scottsteven](https://dub1.discourse-cdn.com/flex005/user_avatar/forums.realmacsoftware.com/scottsteven/32/3418_2.png) [@scottsteven](https://forums.realmacsoftware.com/u/scottsteven)\
**Post date:** [May 30, 2024, 10:00am UTC](https://forums.realmacsoftware.com/t/google-internal-docs-on-search-engine-algorithm-have-leaked/43359/2 "2024-05-30T10:00:25Z")

</div>

Google is the real threat in IT and the future of the internet I have no doubt and lots of my suspicions about them over the years seem to validated now.

---

<div class="post-metadata">

**Author:** ![Bruno](https://dub1.discourse-cdn.com/flex005/user_avatar/forums.realmacsoftware.com/bruno/32/25457_2.png) [@Bruno](https://forums.realmacsoftware.com/u/Bruno)\
**Post date:** [May 30, 2024, 10:30am UTC](https://forums.realmacsoftware.com/t/google-internal-docs-on-search-engine-algorithm-have-leaked/43359/3 "2024-05-30T10:30:48Z")

</div>

In History it has always been very bad and wrong to let just one make the lists… in fact in History it’s always very bad and wrong to let someone make a list

---

<div class="post-metadata">

**Author:** ![Heroic\_Nonsense](https://dub1.discourse-cdn.com/flex005/user_avatar/forums.realmacsoftware.com/heroic_nonsense/32/20666_2.png) [@Heroic\_Nonsense](https://forums.realmacsoftware.com/u/Heroic_Nonsense)\
**Post date:** [May 30, 2024, 11:51am UTC](https://forums.realmacsoftware.com/t/google-internal-docs-on-search-engine-algorithm-have-leaked/43359/4 "2024-05-30T11:51:46Z")

</div>

The first hint this could happen was when they dropped “Don’t be evil.” as their credo.

---

<div class="post-metadata">

**Author:** ![Flash](https://avatars.discourse-cdn.com/v4/letter/f/f475e1/32.png) [@Flash](https://forums.realmacsoftware.com/u/Flash)\
**Post date:** [May 30, 2024, 12:26pm UTC](https://forums.realmacsoftware.com/t/google-internal-docs-on-search-engine-algorithm-have-leaked/43359/5 "2024-05-30T12:26:27Z")

</div>

If you have been doing real scientific A-B research on SEO for any length of time. Absolutely NONE of this is news. This is why my original Rapidweaver-SEO dot com website closed a long time ago.

Trying to win the search engine race for riches and fame is a fools game. You must find a very narrow niche and carve out your market. There is one tool out there that does it. And it works with Rapidweaver products as well as all other website builders if you take the time to learn the system.

---

<div class="post-metadata">

**Author:** ![Flash](https://avatars.discourse-cdn.com/v4/letter/f/f475e1/32.png) [@Flash](https://forums.realmacsoftware.com/u/Flash)\
**Post date:** [May 30, 2024, 12:28pm UTC](https://forums.realmacsoftware.com/t/google-internal-docs-on-search-engine-algorithm-have-leaked/43359/6 "2024-05-30T12:28:05Z")

</div>

> [@Heroic\_Nonsense](#):
>
> The first hint this could happen was when they dropped “Don’t be evil.” as their credo.

It happened several years before that. They dropped it because of other mistakes and legal hearings they got called out on and lost.

---

<div class="post-metadata">

**Author:** ![scottsteven](https://dub1.discourse-cdn.com/flex005/user_avatar/forums.realmacsoftware.com/scottsteven/32/3418_2.png) [@scottsteven](https://forums.realmacsoftware.com/u/scottsteven)\
**Post date:** [May 30, 2024, 1:12pm UTC](https://forums.realmacsoftware.com/t/google-internal-docs-on-search-engine-algorithm-have-leaked/43359/7 "2024-05-30T13:12:07Z")

</div>

What I think we all want is copies of these documents. This could help us get an edge up on managing Evil Google.

---

<div class="post-metadata">

**Author:** ![Bruno](https://dub1.discourse-cdn.com/flex005/user_avatar/forums.realmacsoftware.com/bruno/32/25457_2.png) [@Bruno](https://forums.realmacsoftware.com/u/Bruno)\
**Post date:** [May 30, 2024, 2:41pm UTC](https://forums.realmacsoftware.com/t/google-internal-docs-on-search-engine-algorithm-have-leaked/43359/8 "2024-05-30T14:41:39Z")

</div>

@Flash Hi, I think about it since a quite long time (sorry for grammatical injuries 😁). The result of my intense thoughts 🤯 is that SEO can’t help alone tiny business among megaone , other sources of advertising are required (social media, YouTube, books, social and associative activities, teams, friends of a friend of a friend, newspapers if still read by anyone….). Sad but realistic it seems 😔

---

<div class="post-metadata">

**Author:** ![differentdan](https://dub1.discourse-cdn.com/flex005/user_avatar/forums.realmacsoftware.com/differentdan/32/33530_2.png) [@differentdan](https://forums.realmacsoftware.com/u/differentdan)\
**Post date:** [May 30, 2024, 2:52pm UTC](https://forums.realmacsoftware.com/t/google-internal-docs-on-search-engine-algorithm-have-leaked/43359/9 "2024-05-30T14:52:29Z")

</div>

> [@Heroic\_Nonsense](#):
>
> Just how 2,500 pages of text can lack context was not explained though.

I doubt many people will go through 2,500 pages, but seems like you could plug those documents into ChatGPT and ask it to summarize _“Based on the provided docs, what is the best way to achieve high page rankings in Google’s search engine results pages?”_

Just don’t use Google Gemini 🤣:

_Google’s AI Overviews feature suggested a user put glue on pizza so the cheese would stay put._

_Touts the health benefits of tobacco for kids._

_How many rocks should I eat? and the AI Overview suggested the user take at least one rock per day for vitamins and minerals quoting UC Berkeley geologists._

---

<div class="post-metadata">

**Author:** ![Flash](https://avatars.discourse-cdn.com/v4/letter/f/f475e1/32.png) [@Flash](https://forums.realmacsoftware.com/u/Flash)\
**Post date:** [May 30, 2024, 3:19pm UTC](https://forums.realmacsoftware.com/t/google-internal-docs-on-search-engine-algorithm-have-leaked/43359/10 "2024-05-30T15:19:17Z")

</div>

> [@Bruno](#):
>
> The result of my intense thoughts 🤯 is that SEO can’t help alone tiny business among megaone

This is true. The trick or magic is to compete in a sub market or a small niche. Then once you bring your website visitors in using the niche you can slowly and respectfully introduce them to the larger market you want to win.

---

<div class="post-metadata">

**Author:** ![Flash](https://avatars.discourse-cdn.com/v4/letter/f/f475e1/32.png) [@Flash](https://forums.realmacsoftware.com/u/Flash)\
**Post date:** [May 30, 2024, 3:21pm UTC](https://forums.realmacsoftware.com/t/google-internal-docs-on-search-engine-algorithm-have-leaked/43359/11 "2024-05-30T15:21:19Z")

</div>

> [@differentdan](#):
>
> could plug those documents into ChatGPT

I assume this was a joke. But if not, keep in mind it can only read about 6000 words at a time. That’s a lot of uploading that it cannot keep in context.

---

<div class="post-metadata">

**Author:** ![differentdan](https://dub1.discourse-cdn.com/flex005/user_avatar/forums.realmacsoftware.com/differentdan/32/33530_2.png) [@differentdan](https://forums.realmacsoftware.com/u/differentdan)\
**Post date:** [May 30, 2024, 3:24pm UTC](https://forums.realmacsoftware.com/t/google-internal-docs-on-search-engine-algorithm-have-leaked/43359/12 "2024-05-30T15:24:48Z")

</div>

> [@Flash](#):
>
> I assume this was a joke.

Nope wasn’t a joke, I wasn’t aware there were any character limits in regards to documents it can scan through. Are you sure that 6000 limit is across the board? For example do ChatGPT Plus members get higher character limits, or would using a different model give higher limits (like GPT 3.5 vs 4o), or would using the API give higher limits?

I seem to recall reading about people plugging in **a lot** of information into ChatGPT and it ingesting it fine. Maybe this was in regards to people making their own GPTs in the GPT store though..

I also seem to recall reading about people plugging in legal docs which can be extremely word heavy to get a summarization.

---

<div class="post-metadata">

**Author:** ![Flash](https://avatars.discourse-cdn.com/v4/letter/f/f475e1/32.png) [@Flash](https://forums.realmacsoftware.com/u/Flash)\
**Post date:** [May 30, 2024, 3:44pm UTC](https://forums.realmacsoftware.com/t/google-internal-docs-on-search-engine-algorithm-have-leaked/43359/13 "2024-05-30T15:44:48Z")

</div>

> [@differentdan](#):
>
> I seem to recall reading about people plugging in **a lot** of information into ChatGPT and it ingesting it fine.

It can take in everything, but it will not be an accurate summary.

GPT 4 Turbo can read 128k if you pay for it. That’s about $11US for input on just that amount alone. However output is still limited to about 4K. Will it be a more accurate summary? Yes. However in this case, a technical document, details matter, and since this is more than likely a white paper, there probably is not a lot of fluff. Also if you pay for and develop an API there is obviously a lot more that can be done. But I am assuming we are discussing consumer level products here.

Honestly the players are too big to game for a consistent income anymore. They have been for years.

See my comment above about [winning at niches](https://forums.realmacsoftware.com/t/google-internal-docs-on-search-engine-algorithm-have-leaked/43359/10).

---

<div class="post-metadata">

**Author:** ![Bruno](https://dub1.discourse-cdn.com/flex005/user_avatar/forums.realmacsoftware.com/bruno/32/25457_2.png) [@Bruno](https://forums.realmacsoftware.com/u/Bruno)\
**Post date:** [May 30, 2024, 3:45pm UTC](https://forums.realmacsoftware.com/t/google-internal-docs-on-search-engine-algorithm-have-leaked/43359/14 "2024-05-30T15:45:03Z")

</div>

Yes you’re right. I have created my own private gpt for psychological research. I put them thousands pages (btw many thanks to pdf) and then I ask. The important thing is to be very careful about the constant way you build your asks because it also train itself to learn about that. Of course the statistics lead the game, an example : impossible for ChatGPT to tell that there still isn’t any proof for Psychotherapeutic effectiveness, too much papers about Psychotherapeutic effectiveness to tell the contrary. That’s a drastic and very problematic limitation I think for these tools.

---

<div class="post-metadata">

**Author:** ![scottsteven](https://dub1.discourse-cdn.com/flex005/user_avatar/forums.realmacsoftware.com/scottsteven/32/3418_2.png) [@scottsteven](https://forums.realmacsoftware.com/u/scottsteven)\
**Post date:** [May 30, 2024, 8:39pm UTC](https://forums.realmacsoftware.com/t/google-internal-docs-on-search-engine-algorithm-have-leaked/43359/15 "2024-05-30T20:39:17Z")

</div>

I use jennieAI upload research papers and the like into a folder and ask questions based on those documents normally PDF
