[Reverse Engineer Google] Journey to understand Google TOGETHER. Help me.

splishsplash

Elite Member
Executive VIP
Jr. VIP
Joined
Oct 9, 2013
Messages
3,471
Reaction score
14,453
As some of you know I'm working on new guides, but before I write them I'm doing some really deep case studies and analyses.

I remember 4 years ago, back in May 2021. I wanted to reverse engineer the Google update.

I set about to gather all the data I could, then I signed up to this cool statistical analysis site that listed and had guides for almost every stat method there is.

I spent a month or 2 total on this. Learned a lot of math, but in the end I had to give up.

The reality was, to gain meaningful insights I would have had to spend a year studying 8 hours a day, and then another year working as a mathematician.

Fast forward to today.

In the space of an hour, using agents, I was able to produce something that would have taken a mathematician with 10 years industry experience a week to do himself.

We truly are entering an era of unbelievable opportunity.

Now..

Let's not get too crazy here. We aren't going to "reverse engineer Google". That's just a convenient shorthand way to write.

Our goal is to gain useful insights that we can apply to our campaigns.

So far, the first thing i'm working on is my theory of a legitimacy filter. This is a very hard and complex one, but I'm making progress.

I'm also going deep on the Helpful Content Update and the most recent spam update of August 2025.

What makes this particularly difficult is that everything is connected.

For example, here's one thing I've observed.

March this year was a core update. In this update they let go of the reigns a bit.

Sites that had no(what I call legitimacy) started to rank again for YMYL and commercial terms.

In Aug 2025, we had a spam update.

A lot of sites that came from nowhere(like dead sites, 0-20 traffic then shot up to the 1000's in march) tanked again in Aug 2025.

Some COMPLETELY tanked, while others only partially.

There was also a June core update, and I noticed a lot of those sites dropped in July too, so there was re-adjustments there, but Aug was the one they tackled it, aiming for more 'balance'.

Ie, it was too strict before, and many legit small publishers were harmed. Now it's more balanced.

Just to give you an idea of how deep I'm going here and how much work is going into this. I asked this to sonnet in the research folder.

1762170760539.png


Here's the response:

1762170796339.png


It's an absolute beast. It's pushing my working memory to the maximum just to keep track of everything.

Each insight leads to 5 different research ideas, which lead to more insights and it spirals off and needs to be summarized and connected into meaningful predictive theories.

It's exciting, and I don't exaggerate when i say this..

What one person can do with these agents is on par with what only billionaires could do just 10 years ago.

Not even millionaires, but billionaires. The power at your fingertips is out of this world. People have no fucking idea yet.

Especially if you combine them with your human intelligence and guide them properly, it's like having teams of 100's of devs, mathematicians and research assistants.

How You Can Help
I have pretty good resources for this. I can afford multiple copies of the $200/mo plans for openai/anthropic and burn a few thousand above that on APIs and data. Although to do real damage you'd need about $100k-$200k for data, and another $200k for inference costs. That would be something. The insights from that would be incredible.

So while I do have the resources to make progress, I still need more data.

I really need people to come together on this. It won't be enough if I just get 10-15 people sending me their site.

I need everyone to pool their resources here.

I specifically need this :-

Small to medium sized sites ONLY for everything.

  • Sites that were hit with HCU in 2023 and never recovered
  • Sites that were hit with HCU in 2023, recovered in march 2025, and were hit again aug 2025
  • Sites that were hit with HCU in 2023, recovered in march 2025, and SURVIVED aug 2025
  • Sites that have the "anonymous blog" feel, but are ranking. Ie, they are transparent entities.
  • Sites that are less than 1 year old and are ranking really well.
  • Sites that are less than 1 year old and cannot rank despite strong link building and effort(Please, no beginners sending me sites they can't rank. For this, I need experienced guys who have ranked in the past, but are struggling with newer sites)
  • Sites that are older than 2 years, and ranked well for a long time, but have since tanked. Can be clean or spammy sites. Both have value. Clean is more valuable.


Even better are people that have mass data they would be willing to share. Any data can be useful.

I'm also open to the sharing of ideas and critique of any approaches in the analyses.

This is one of the hardest times in SEO with Google stealing so much of people's traffic with generative AI.

We have an opportunity to pool our resources together to achieve insights that can help us collectively.

If the thread is successful and people are contributing, then I will share insights here in the journey. If not.. No problem, I will just continue with my own resources and publish my guides when they're ready.
 
Just to get the ball rolling I'll share a few insights here from what I'm working on.

I analysed 6 low quality info sites that rose in march 2025 and dropped june-aug-2025 (very small dataset to start, I'm aware. I will need to code something to hunt for 100's of them to do a bigger analysis, but results are statistically significant here.

- Overall non‑new pages: 2,656; Dropped: 2,292; Drop rate: 86.3%.
- High‑volume vs lower‑volume:
- Drop: 36.9% (n=160) vs 89.5% (n=2,496); diff −52.6 pp; z=−18.75; p≈0.0 [SIGNIFICANT].
- YMYL vs non‑YMYL:
- Drop: 92.4% (n=171) vs 85.9% (n=2,485); diff +6.5 pp; p≈0.016 [SIGNIFICANT].
- Commercial vs non‑commercial:
- Drop: 91.9% (n=358) vs 85.4% (n=2,298); diff +6.5 pp; p≈9.2e‑4 [SIGNIFICANT].
- Parasitic vs non‑parasitic:
- Drop: 68.8% (n=32) vs 86.5% (n=2,624); diff −17.8 pp; p≈0.0037 [SIGNIFICANT].
- Navigational/About vs other:
- Drop: 85.9% (n=384) vs 86.4% (n=2,272); diff −0.4 pp; p≈0.826 [not significant].

So YMYL and commercial dropped significantly. Almost completely wiped out.

Parasitic(brand jacking) pages dropped significantly less. Beware when you buy guest posts from sites that have a lot of parasitic/brand jacking keywords and nothing else.

navigational/about style - not significant.

From that we can see YMYL and commercial are targeted when legitimacy score is low (All 6 sites have almost 0 legitimacy score)

The weird one is high vol vs low vol. More low vol dropped

But, that one is a perfect example of correlation does not imply causation.

This highlights the problem.

I asked codex: › Analyse and report the average volume for YMYL, commercial, parasitic and nav

We get

- YMYL
- Mean volume: 6,353
- Median: 45
- n: 24
- Commercial
- Mean volume: 9,290
- Median: 80
- n: 39
- Parasitic
- Mean volume: 1,253
- Median: 1,100
- n: 19
- Nav/About
- Mean volume: 1,539
- Median: 450
- n: 105

Mean, can be misrepresentative of the cohort.

If I give you this data

3,8,1,5,15,4,3492

The mean = 504
median = 5

So notice the median for YMYL, only 45.
Commercial, only 80

Yet parasitic.. 1100!

and we know from the parasitic cohort that parasitic survive a LOT more than non-parasitic.

p = 0.0037. If you calculate the z-score for that, it satisfies 5%, 1% and strict 0.1% thresholds.


Next, what I'm doing is this :-


› Here's what i want you to do. Create a new folder in site-drops called mild-vs-severe where we will compare the content of sites who experienced mild drops vs those that experienced severe. I want you to make a csv where you note the site, page, keyword and whether that site was mild or
severe, and whether the page was mild or severe. So we end up with site mild, page mild, site mild, page severe, site severe, page severe. 3 groups. With enough in each group to do a statistically significant study. Only create the csv for this, don't start the study



Then I created a sub-agent inside claude code to do this


------

Content Quality Assessor Agent

What It Does

The content-quality-assessor is a specialized SEO agent that performs competitive content gap analysis. It evaluates how well a page's content quality and user intent alignment compares against the top-ranking competitors for a specific keyword.

Core Function

This agent conducts a rigorous side-by-side comparison between:
- Your target page (the page you want to improve)
- Top 5 ranking competitors for your target keyword

It identifies exactly why top pages rank and reveals what your page needs to compete effectively.

Analysis Methodology

Phase 1: Benchmark Analysis (Top 5 Competitors)

- User Intent Determination - Identifies what users actually want (informational, commercial, transactional, navigational)
- Content Quality Baseline - Analyzes topical depth, entities mentioned, facts/data provided, content formats used
- Pattern Recognition - Identifies "table stakes" content elements common across all top performers
- Differentiator Identification - Notes unique approaches individual pages use to stand out

Phase 2: Primary Page Assessment

Evaluates your page across 7 dimensions:
1. User Intent Alignment - Does it match what users searching this keyword want?
2. Topical Depth & Coverage - Are you covering all the topics top pages cover?
3. Entity Coverage - Do you mention the brands, products, concepts that matter?
4. Factual Content & Evidence - How does your data, citations, and proof compare?
5. Usefulness & User Value - Would searchers find your page more or less helpful?
6. E-E-A-T Signals - How does expertise, experience, authority, and trust compare?
7. Content Format & Structure - Does your format meet user expectations?

Output Deliverables

The agent provides:
- User Intent Analysis - What users really want when searching this keyword
- Top 5 Benchmark Summary - What makes competitors successful
- Primary Page Assessment - Overall verdict with detailed gap analysis
- Strategic Recommendations - Prioritized action plan (Critical → High Impact → Optimization)
- Competitive Positioning - Realistic path to outranking competitors

When to Use This Agent

Perfect for:
- Understanding why a new page isn't ranking
- Pre-publishing validation of content refreshes
- Site-wide content audits
- Capitalizing on competitor ranking drops
- Identifying content opportunities in your niche

Example Scenarios:
- "Why isn't my infrared sauna guide ranking on page 1?"
- "I rewrote our buying guide - is it competitive now?"
- "Our competitor just dropped from #2 to #8 - what changed?"
- "Do our product pages match user intent for target keywords?"

Technical Details

- Model: Haiku (fast and cost-efficient)
- Color: Purple
- Agent Type: Task-based subagent
- Focus: Content gap analysis, not technical SEO or link building


-----


The new topic studies we have :-

How We Label Things

- Site class: “mild” = the site had a smaller overall drop; “severe” = the site had a big drop.
- Page class: “mild” = the page still gets traffic (survived); “severe” = the page lost traffic (dropped).

Main File (All Examples)

- File: analyses/site-drops/mild-vs-severe/dataset.csv
- What’s inside: site, page URL, top keyword, plus the site class and page class.
- Only the three combos you asked for:
- mild site + mild page: 25 examples
- mild site + severe page: 27 examples
- severe site + severe page: 2,265 examples

Four Topic Study Files

- Each file focuses on a topic type and includes up to 10 examples from each combo (mild/mild, mild/severe, severe/severe). Some have fewer than 10 where the pool is small after removing New pages.
- YMYL (money/health/legal): analyses/site-drops/mild-vs-severe/dataset_study-1.csv
- mild/mild: 1; mild/severe: 5; severe/severe: 10
- Commercial (reviews/best/prices/etc.): analyses/site-drops/mild-vs-severe/dataset_study-2.csv
- mild/mild: 1; mild/severe: 3; severe/severe: 10
- Parasitic/manufactured keywords: analyses/site-drops/mild-vs-severe/dataset_study-3.csv
- mild/mild: 8; mild/severe: 3; severe/severe: 10
- Misc/About/Nav (the rest + navigational/“about”): analyses/site-drops/mild-vs-severe/dataset_study-4.csv
- mild/mild: 10; mild/severe: 10; severe/severe: 10

So we have..

YMYL topic study.

Broken up into 3 cohorts

mild site (site had mild drop. < 70% total traffic)

with mild page (page had mild traffic drop. < 70% total traffic)

mild site + severe page (severe is 80% or more traffic lost)

severe site + severe page (both the site and the page lost 80% or more)

Then the same for commercial, parasitic and misc/about/nav.

Then sonnet will go through each, use the above sub-agent to check the current top 5, and do a quality/user-intent analysis on the page in the study RELATIVE TO, the top 5.

Google is using AI these days rather than simple keyword counts to determine quality of pages, so this gives us something closer to what they have since we are using the top 5(their own choices) as reference points.

This then lets us do a study to see if there's things in common that the different groups have within their CONTENT relative to what google wants to see now in the top 5.

I'm about to run this. It'll take a little while, it's a lot of checks, but I'll report back when I have insights.

Any questions feel free to ask
 
Back
Top