- Oct 9, 2013
- 3,471
- 14,453
As some of you know I'm working on new guides, but before I write them I'm doing some really deep case studies and analyses.
I remember 4 years ago, back in May 2021. I wanted to reverse engineer the Google update.
I set about to gather all the data I could, then I signed up to this cool statistical analysis site that listed and had guides for almost every stat method there is.
I spent a month or 2 total on this. Learned a lot of math, but in the end I had to give up.
The reality was, to gain meaningful insights I would have had to spend a year studying 8 hours a day, and then another year working as a mathematician.
Fast forward to today.
In the space of an hour, using agents, I was able to produce something that would have taken a mathematician with 10 years industry experience a week to do himself.
We truly are entering an era of unbelievable opportunity.
Now..
Let's not get too crazy here. We aren't going to "reverse engineer Google". That's just a convenient shorthand way to write.
Our goal is to gain useful insights that we can apply to our campaigns.
So far, the first thing i'm working on is my theory of a legitimacy filter. This is a very hard and complex one, but I'm making progress.
I'm also going deep on the Helpful Content Update and the most recent spam update of August 2025.
What makes this particularly difficult is that everything is connected.
For example, here's one thing I've observed.
March this year was a core update. In this update they let go of the reigns a bit.
Sites that had no(what I call legitimacy) started to rank again for YMYL and commercial terms.
In Aug 2025, we had a spam update.
A lot of sites that came from nowhere(like dead sites, 0-20 traffic then shot up to the 1000's in march) tanked again in Aug 2025.
Some COMPLETELY tanked, while others only partially.
There was also a June core update, and I noticed a lot of those sites dropped in July too, so there was re-adjustments there, but Aug was the one they tackled it, aiming for more 'balance'.
Ie, it was too strict before, and many legit small publishers were harmed. Now it's more balanced.
Just to give you an idea of how deep I'm going here and how much work is going into this. I asked this to sonnet in the research folder.

Here's the response:

It's an absolute beast. It's pushing my working memory to the maximum just to keep track of everything.
Each insight leads to 5 different research ideas, which lead to more insights and it spirals off and needs to be summarized and connected into meaningful predictive theories.
It's exciting, and I don't exaggerate when i say this..
What one person can do with these agents is on par with what only billionaires could do just 10 years ago.
Not even millionaires, but billionaires. The power at your fingertips is out of this world. People have no fucking idea yet.
Especially if you combine them with your human intelligence and guide them properly, it's like having teams of 100's of devs, mathematicians and research assistants.
How You Can Help
I have pretty good resources for this. I can afford multiple copies of the $200/mo plans for openai/anthropic and burn a few thousand above that on APIs and data. Although to do real damage you'd need about $100k-$200k for data, and another $200k for inference costs. That would be something. The insights from that would be incredible.
So while I do have the resources to make progress, I still need more data.
I really need people to come together on this. It won't be enough if I just get 10-15 people sending me their site.
I need everyone to pool their resources here.
I specifically need this :-
Small to medium sized sites ONLY for everything.
Even better are people that have mass data they would be willing to share. Any data can be useful.
I'm also open to the sharing of ideas and critique of any approaches in the analyses.
This is one of the hardest times in SEO with Google stealing so much of people's traffic with generative AI.
We have an opportunity to pool our resources together to achieve insights that can help us collectively.
If the thread is successful and people are contributing, then I will share insights here in the journey. If not.. No problem, I will just continue with my own resources and publish my guides when they're ready.
I remember 4 years ago, back in May 2021. I wanted to reverse engineer the Google update.
I set about to gather all the data I could, then I signed up to this cool statistical analysis site that listed and had guides for almost every stat method there is.
I spent a month or 2 total on this. Learned a lot of math, but in the end I had to give up.
The reality was, to gain meaningful insights I would have had to spend a year studying 8 hours a day, and then another year working as a mathematician.
Fast forward to today.
In the space of an hour, using agents, I was able to produce something that would have taken a mathematician with 10 years industry experience a week to do himself.
We truly are entering an era of unbelievable opportunity.
Now..
Let's not get too crazy here. We aren't going to "reverse engineer Google". That's just a convenient shorthand way to write.
Our goal is to gain useful insights that we can apply to our campaigns.
So far, the first thing i'm working on is my theory of a legitimacy filter. This is a very hard and complex one, but I'm making progress.
I'm also going deep on the Helpful Content Update and the most recent spam update of August 2025.
What makes this particularly difficult is that everything is connected.
For example, here's one thing I've observed.
March this year was a core update. In this update they let go of the reigns a bit.
Sites that had no(what I call legitimacy) started to rank again for YMYL and commercial terms.
In Aug 2025, we had a spam update.
A lot of sites that came from nowhere(like dead sites, 0-20 traffic then shot up to the 1000's in march) tanked again in Aug 2025.
Some COMPLETELY tanked, while others only partially.
There was also a June core update, and I noticed a lot of those sites dropped in July too, so there was re-adjustments there, but Aug was the one they tackled it, aiming for more 'balance'.
Ie, it was too strict before, and many legit small publishers were harmed. Now it's more balanced.
Just to give you an idea of how deep I'm going here and how much work is going into this. I asked this to sonnet in the research folder.

Here's the response:

It's an absolute beast. It's pushing my working memory to the maximum just to keep track of everything.
Each insight leads to 5 different research ideas, which lead to more insights and it spirals off and needs to be summarized and connected into meaningful predictive theories.
It's exciting, and I don't exaggerate when i say this..
What one person can do with these agents is on par with what only billionaires could do just 10 years ago.
Not even millionaires, but billionaires. The power at your fingertips is out of this world. People have no fucking idea yet.
Especially if you combine them with your human intelligence and guide them properly, it's like having teams of 100's of devs, mathematicians and research assistants.
How You Can Help
I have pretty good resources for this. I can afford multiple copies of the $200/mo plans for openai/anthropic and burn a few thousand above that on APIs and data. Although to do real damage you'd need about $100k-$200k for data, and another $200k for inference costs. That would be something. The insights from that would be incredible.
So while I do have the resources to make progress, I still need more data.
I really need people to come together on this. It won't be enough if I just get 10-15 people sending me their site.
I need everyone to pool their resources here.
I specifically need this :-
Small to medium sized sites ONLY for everything.
- Sites that were hit with HCU in 2023 and never recovered
- Sites that were hit with HCU in 2023, recovered in march 2025, and were hit again aug 2025
- Sites that were hit with HCU in 2023, recovered in march 2025, and SURVIVED aug 2025
- Sites that have the "anonymous blog" feel, but are ranking. Ie, they are transparent entities.
- Sites that are less than 1 year old and are ranking really well.
- Sites that are less than 1 year old and cannot rank despite strong link building and effort(Please, no beginners sending me sites they can't rank. For this, I need experienced guys who have ranked in the past, but are struggling with newer sites)
- Sites that are older than 2 years, and ranked well for a long time, but have since tanked. Can be clean or spammy sites. Both have value. Clean is more valuable.
Even better are people that have mass data they would be willing to share. Any data can be useful.
I'm also open to the sharing of ideas and critique of any approaches in the analyses.
This is one of the hardest times in SEO with Google stealing so much of people's traffic with generative AI.
We have an opportunity to pool our resources together to achieve insights that can help us collectively.
If the thread is successful and people are contributing, then I will share insights here in the journey. If not.. No problem, I will just continue with my own resources and publish my guides when they're ready.