Google Search API documents leak

SpaceYams

Registered Member
Joined
Jan 17, 2023
Messages
64
Reaction score
329
it seems recently G exposed some of their API info on Github:

During our call, this contact showed me the leak itself: more than 2,500 pages of API documentation containing 14,014 attributes (API features) that appear to come from Google’s internal “Content API Warehouse.” Based on the document’s commit history, this code was uploaded to GitHub on Mar 27, 2024th and not removed until May 7, 2024th.

article posted yesterday: An Anonymous Source Shared Thousands of Leaked Google Search API Documents with Me; Everyone in SEO Should See Them

Naturally, I was skeptical. The claims made by this source (who asked to remain anonymous) seemed extraordinary–claims like:
  • In their early years, Google’s search team recognized a need for full clickstream data (every URL visited by a browser) for a large percent of web users to improve their search engine’s result quality.
  • A system called “NavBoost” (cited by VP of Search, Pandu Nayak, in his DOJ case testimony) initially gathered data from Google’s Toolbar PageRank, and desire for more clickstream data served as the key motivation for creation of the Chrome browser (launched in 2008).
  • NavBoost uses the number of searches for a given keyword to identify trending search demand, the number of clicks on a search result (I ran several experiments on this from 2013-2015), and long clicks versus short clicks (which I presented theories about in this 2015 video).
  • Google utilizes cookie history, logged-in Chrome data, and pattern detection (referred to in the leak as “unsquashed” clicks versus “squashed” clicks) as effective means for fighting manual & automated click spam.
  • NavBoost also scores queries for user intent. For example, certain thresholds of attention and clicks on videos or images will trigger video or image features for that query and related, NavBoost-associated queries.
  • Google examines clicks and engagement on searches both during and after the main query (referred to as a “NavBoost query”). For instance, if many users search for “Rand Fishkin,” don’t find SparkToro, and immediately change their query to “SparkToro” and click SparkToro.com in the search result, SparkToro.com (and websites mentioning “SparkToro”) will receive a boost in the search results for the “Rand Fishkin” keyword.
  • NavBoost’s data is used at the host level for evaluating a site’s overall quality (my anonymous source speculated that this could be what Google and SEOs called “Panda”). This evaluation can result in a boost or a demotion.
  • Other minor factors such as penalties for domain names that exactly match unbranded search queries (e.g. mens-luxury-watches.com or milwaukee-homes-for-sale.net), a newer “BabyPanda” score, and spam signals are also considered during the quality evaluation process.
  • NavBoost geo-fences click data, taking into account country and state/province levels, as well as mobile versus desktop usage. However, if Google lacks data for certain regions or user-agents, they may apply the process universally to the query results.
  • During the Covid-19 pandemic, Google employed whitelists for websites that could appear high in the results for Covid-related searches
  • Similarly, during democratic elections, Google employed whitelists for sites that should be shown (or demoted) for election-related information

another relevant post with info: Secrets from the Algorithm: Google Search’s Internal Engineering Documentation Has Leaked

i believe this may be one of the indexed repos but i have not confirmed: https://hexdocs.pm/google_api_content_warehouse/api-reference.html

happy reading/digging nerds :D
 
The recent leak of over 2,500 pages of Google’s internal API documentation, including 14,014 attributes from their "Content API Warehouse," exposes fascinating details about Google’s search mechanisms. Key revelations include the NavBoost system’s use of clickstream data from Chrome to enhance search quality, and how Google’s algorithms leverage user clicks, query changes, and engagement to rank results. It also discusses Google's use of whitelists for critical searches during events like Covid-19 and elections. For more insights, explore the full article on SparkToro and the indexed repo on hexdocs
it seems recently G exposed some of their API info on Github:


article posted yesterday: https://sparktoro.com/blog/an-anonymous-source-shared-thousands-of-leaked-google-search-api-documents-with-me-everyone-in-seo-should-see-them/


another relevant post with info: https://ipullrank.com/google-algo-leak

i believe this may be one of the indexed repos but i have not confirmed: https://hexdocs.pm/google_api_content_warehouse/api-reference.html

happy reading/digging nerds :D
 
The recent leak of over 2,500 pages of Google’s internal API documentation, including 14,014 attributes from their "Content API Warehouse," exposes fascinating details about Google’s search mechanisms. Key revelations include the NavBoost system’s use of clickstream data from Chrome to enhance search quality, and how Google’s algorithms leverage user clicks, query changes, and engagement to rank results. It also discusses Google's use of whitelists for critical searches during events like Covid-19 and elections. For more insights, explore the full article on SparkToro and the indexed repo on hexdocs
thanks for your shitty GPT rehash
 
Do you have leaked document?

i linked one of them. view v0.4.0 https://hexdocs.pm/google_api_content_warehouse/0.4.0/api-reference.html

for example here are the docs for NavBoost which the article talks about w.r.t clicks:

In reality, Navboost has a specific module entirely focused on click signals.
The summary of that module defines it as “click and impression signals for Craps,” one of the ranking systems. As we see below, bad clicks, good clicks, last longest clicks, unsquashed clicks, and unsquashed last longest clicks are all considered as metrics. According to Google’s “Scoring local search results based on location prominence” patent, “Squashing is a function that prevents one large signal from dominating the others.” In other words, the systems are normalizing the click data to ensure there is no runaway manipulation based on the click signal. Googlers argue that systems in patents and whitepapers are not necessarily what are in production, but NavBoost would be a nonsensical thing to build and include were it not a critical part of Google’s information retrieval systems.

https://hexdocs.pm/google_api_conte...NavboostCrapsCrapsData.html#module-attributes

https://hexdocs.pm/google_api_conte...CrapsCrapsClickSignals.html#module-attributes
 
Has anyone here made any changes in how they're approaching SEO based off of these docs or is it business as usual?
 
From what I've read, most of it was pretty much expected. I will say that even if they collect the information it doesn't mean they're actually using it.
 
Exactly my opinion. So G might have used domain age. But not using it nowadays.

And: you never know, if features (e.g. domain age, #words in title tags, #clicks on porn sites after site visit, whatsoever) might be useful in the future, when extending the canon of features. Still possible to figure out new, latent insights, where the feature for itself might not seem to be useful right now.

Only completely surprising takeaway: Matt Cutts was lying to us for many years :)
 
From what I skimmed through, not much that we don't already know to be honest.

Most factors have been discussed here for years.
 
Until I see that leak for myself this is just fake-news
 
Back
Top