Big changes in Google content scoring

Some fine info in this thread. You'd be hard pressed to find this type of discourse on some premier white hat seo forums. Empirical data used to break down google's algo is a great way to test/enable your S.E placement, and I've done so on more than one occasion in the past..but nothing on the scale of O.P.

IPOPBB- you running your seo lab from a vm?
 
Very intelligent observations ipopbb. This thread contains some valuable information.

I also think Matt Cutts is providing the seo world with false information.

I own a price comparison website and would rather have you as my seo consultant than Matt Cutts.
 
How do you compile all results, by "hand" ??!?

Sometimes... usually not, but google modifies their result HTML fairly frequently so to insure the parser is accurate does require periodic manual review and adjustment... bots thats true for most web based data scraping bots.
 
MaTT said there are over 200 factors for ranking.
But value of those vary from time to time. I feel they do some changes daily. What more i sould say when one of my friend told they do minor changes at every 5 minutes... OMG

I don't entirely trust matt cutts... I find him purposefully misleading.

In my opinion I think M.A.D. is correct on the kinds of math and methods being used to determine if a page is strongly referring to a keyword based on node topography, but I think there is an AI Baysian Categorizer that creates some kind of spammyFeeling metric. I say this because the key to displaying quality results comes down to sorts and filters on factors largely outside of document search. I do believe there are 200 weighted factors that funnel down to a hypothetical spammyFeeling metric between 0 and 1 or 0 and 100 or what ever scale. I think M.A.D. is correct that the document search part of Google doesn't care about tags. But something is sorting my experiment pages and I think it is an effect of the Baysian Categorizer that is acting as Google Secret Sauce and is the Reason Matt Cutts job exists. Matt oversees the list of spammyFeeling factors and there weights in making a determination and then they throw a large number of known abuses and known guinuine pages as a training set for the AI. The slow changes I see over time feeling remarkably like the changes I see when my own Baysian application have a large training set that slowly changes over time. And then occasionally when I pull a Matt Cutts and re-weigh factors in making fuzzy classifications.

I believe(a.k.a.assume without a way of testing) currently (as of this morning) that tag order from Google is Baysian in nature. It explains why tag agnostic search innovation would consistently and reliably sort my experiment pages that only differ by a single tag in most cases. It also explains the nature and frequency of the changes.

When matt cutts, cancer man, says the factors change all the time... He's playing with the notion that the machine is constantly learning and updating the statistics and classifications... the deterministic factors probably change quarterly at best (every 2-4 months, and usually the factors don't change but their weights in making a determination) and most people have little if any likelihood of seeing changes that frequently. (New sites excluded, they see a lot of change as they progress from 0-10 months... I'm only talking about established sites)
 
A huge thank you to all the above contributors. Such a valuable thread with some actual meat in it. Real reading matter. More please!

Bugs

I have Charts of common SEO factors in my other thread! ;)
I can't believe the crickets in that thread... Its the meatiest data I've ever seen.

Its called "More Data from the SEO Lab" or something
 
Some fine info in this thread. You'd be hard pressed to find this type of discourse on some premier white hat seo forums. Empirical data used to break down google's algo is a great way to test/enable your S.E placement, and I've done so on more than one occasion in the past..but nothing on the scale of O.P.

IPOPBB- you running your seo lab from a vm?

Yup... a VPS host... I have more info about my lab setup in my other thread... With Charts of the influence of common SEO factors!

Its called More data from the SEO Lab or something.
 
Very intelligent observations ipopbb. This thread contains some valuable information.

I also think Matt Cutts is providing the seo world with false information.

I own a price comparison website and would rather have you as my seo consultant than Matt Cutts.

Thanks! I think! ;) I set the bar a lot higher than beating Matt Cutts at SEO. A dead possum on the highway would give you better SEO advice than Matt... Of course Matt calls himself Google's Web Spam guy on his "About Me" page of his blog so he's really not out to help anyone but the searcher. He is of no help to us or anyone else running a website as a business.

Black Hack SEO Science can own Google if we all work together like the Scientific Community's love child with the Open Source Community.

Empirical data, Peer Review, Reproducible results,... now that's a force to be reckoned with!

Cheers
 
somehow my post go deleted...
why is caption in there twice at #10 and #28?
and not sure what #8 is?

Thanks
 
great thread! please keep discussing and elaborating more about it! thanks to the OP and to M.A.D for your valuable information :)
 
I agree that there's a lot of Bayesian type-math going on behind the scenes. The thing that I think people under-estimate is how often the training data is updated. I know there's a HUGE amount of human review on the index. A lot of that is simply finding footprints and removing bad pages from the index, but that team also directly contributes on a daily basis to the index of "things that aren't spammy enough to be deindexed by the rules, but shouldn't rank well".
 
somehow my post go deleted...
why is caption in there twice at #10 and #28?
and not sure what #8 is?

Thanks

#8 means

Code:
<html>
<head>
</head>
<body>
<textarea>keyword</textarea>
</body>
</html>

I notice from time to time that pages with caption tags appear in the results twice. It tends to happen when google is making big changes. It usually only lasts a month or two... but its already gone now, which makes me think Google has less deployment overhead now than in recent years.

I used to use captions on all pages just to have it there and primed for when google changes things and it pops back. But only got about 4-5 days this go around.
 
Thanks for sharing. These findings are just too much to get my head around.

I guess Market Samurai's SEO Competition module needs an update or two :)


ffranko
 
Thanks for the explanation and definitely for sharing
It is good to see this and I know there are other code factors going on It just helps to make it all up to date and not 1995 basic compuserve html type stuff with just <p> everywhere.
 
Wow, fantastic information, I'm blown away by your SEO knowledge, your clearly a veteran and have certainly done your testing thoroughly!
 
Wow, fantastic information, I'm blown away by your SEO knowledge, your clearly a veteran and have certainly done your testing thoroughly!

Thanks! I don't know about my SEO knowledge being that impressive... its mostly unproven theories. My real accomplishment was realizing most White Hat advice is often total BS and frequently not helpful and sometimes even damaging. My brand of SEO works for me. Things improved quickly when I stopped believing in the "experts" and started measuring and testing Winners and SERPs for myself. That's what people should really take away from this thread... that and sharing findings for peer review.

I don't think I've ever gotten so many thanks and so much rep at a faster rate.

;)
 
I wish I had a lab like that. ipopbb, thanks for sharing all these stuff.


But hey, what the hell is samp..

Definition:
<samp> = Defines sample computer code
 
Last edited:
Thanks a lot, this is pure gold.

A few issues I would like your opinion on:
What is the best way of using multiple types of tags in a page?
How their combined density and individual density have an effect, because I understand that between h1 and blockquote I should only pick one but how about the strong, em, samp, kbd or i, b, u, big, small tags, should I aim of using each of them at least once ?
 
hmm....very interesting. Some things I have to think about here. Thanks for the share.
 
One thing I am seeing them do is linguistic ranking shifts due to how they are deciding what people are really looking for. I know an ecommerce site that held #1 for 7+ years for one of its primary keywords and all of a sudden it dropped to #3. The two sites that suddenly appeared in #1 and #2 were not competition. They were old sites that had to do with a different meaning of the original keyword (you know, like in a dictionary where a word or phrase has several meanings). This sudden change could have only come about because Google had enough data to decide that when people are searching for that particular phrase, most of them are in fact looking for something other than what the ecommerce site offered. It wasn't a case of the ecommerce site doing something wrong with SEO or the other sites doing something right with SEO. I guess the lesson there is make sure you really get your targeted keywords accurate because it's you against thousands or millions of searchers who know exactly what they want, and Google is watching them to tune the SERPS to that.
 
Back
Top