More Empirical Data From MY SEO Lab

I might release it as a service if it works well so I can't give too much detail out.
 
Last edited:
So what I can gather from the graphs, the most important data is where the blue average line differs at Rank #1 than from the rest of the line.
Would this be correct?
 
Can you tell us what is URL bytes ?
Is it how many chars is in domain or in URL ?
Does it mean that less chars in domain or in URL getting better position ?

It is how many chars are in the URL
 
Wow... ipopbb, this is some of the best stat porn I have ever seen, and it reinforces what I have suspected through casual observation.

Are you planning to do more of these monthly? I would expect that the results would be similar, barring any major algorithm changes from the Big G.

I don't know about monthly but planning to update in the near future and will release those to... for free of course... Everyone keeps expecting me to charge... no angle here... People in this forum can't possibly afford me. I took the sites I manage from $8M per year to $35M per year... I'll let you guess what they pay me to keep me from jumping ship to amazon, nordstrom, eddie bauer, REI, etc... I am giving back. That's my angle! ;)
 
Last edited:
Also what is with meta keyword ? Sites with less keywords in meta tag are getting better position ?

That's how I read it too... to me it looks like google gives you a hypothetical 10 points to divvy among your meta keywords... you can give 10 points to 1 or one point to 10... I noticed also a lot of top performers tune for 1 or 2 max search terms... I think this data supports that. Tuning a page for keyword density also supports that notion. You can only have 1 top performer ... tuning for too much is the same as not tuning.
 
Hm...I don't understand double charts.
For example one with number of words and images ?
Is it better to have more or less words and images ?

Google used to display page size in the SERPs... 63KB 130KB... I noticed way back then the there was a sweet spot for page 1 between 60KB and 220KB... to small and was authoritative enough and too big and its more than a human being can likely sift through... Google still has a sweet spot, but I don't know if they are looking at source size, or visible content size, or words, or sentences, or a balance between text and graphics... I just don't know... so I measure.

These are just the measurements I made. They have no baring from one graph to the next... the image measurement just measures images and the words per page just measures words per page... the fact they appear together is just how I took the screenshots.
 
So what I can gather from the graphs, the most important data is where the blue average line differs at Rank #1 than from the rest of the line.
Would this be correct?

Look at Max Title Matches... any more than 4 is a one way ticket out of the top 100 AND 1 match appears stronger than 4 if page 1 is your goal.

Even more important is that the rules appear to change from page 10-3 and page 2-1... TWO GOOGLE Algorithms. Which do you tune for? I tune long tail pages for page 3 algorithm for best revenue effects and splash pages for page one for traffic gen. I don't know why I'm the only one who sees it. Google's best kept secret is that there isn't 1 algorthm... there are 2!
 
Hey man, any news from your lab ?
 
Hey man, any news from your lab ?

Hoping to recompile these charts soon, but I have to finish adapting my money makers for the new Google First... I have a TON of awesome new things to share once I have data and revenue to back them.

I'm stealing some darkgrayhat techniques from Target! Who woulda thought! ;)

Can't wait to share, but it will still be some time before the results are available.

Be patient.
 
Wow, funny enough these statistics are PRECISELY consistent with my own rules of on-page optimization for the most part!

What a valuable share, thank you so much for contributing to this already awesome forum.

Two quick questions though:

1) Can you define HTML Bytes and Visible Bytes? Is that per page or per domain?

2) What's the formula for "title/url/etc matches"? I'm having a hard time understanding how you get 0.X values :)
 
1) Can you define HTML Bytes and Visible Bytes? Is that per page or per domain?

I don't analyze websites... just the result pages as they are linked from the SERPs. HTML Bytes is the document size in bytes of the result page. If you were to view source on a result page and count the characters you would measure what I labeled HTML bytes.

Visible Bytes = HTML bytes - HEAD block - SCRIPT blocks - COMMENT blocks - STYLE blocks - HTML tags

in concept... the maximum number of characters of text content you are capable of displaying from the result source. The number assumes pages don't cloak, but they all do at least a little.


2) What's the formula for "title/url/etc matches"? I'm having a hard time understanding how you get 0.X values :)

Keep in mind that this data represents hundreds of searches. If you calculate that average number of times the search terms appeared in the title you will most likely get a fractional amount. When you look at the charts understand that they are measuring occurrences as hundreds of different searches trend towards result 1 on page 1.

Once you wrap your mind around that nut then the data gets exciting. ;)
 
ipopbb, I think that google analyze only part of page that is unique compared with other pages on that domain.
For example it won't look in header, sidebar, footer and other elements that are the same or similar on all pages.
This is just guess.
Most of my sites today are on WP and maybe it's case with wordpress blogs where main content is usually in "content" div, so it's easy for them to find it.
Maybe G measure only content (kw density) inside that main part of page.


It could be good if you can somehow scan few pages inside domain and try to clean those repeated elements. I know it could take a lot of time to code that.
 
ipopbb, I think that google analyze only part of page that is unique compared with other pages on that domain.
For example it won't look in header, sidebar, footer and other elements that are the same or similar on all pages.
This is just guess.
Most of my sites today are on WP and maybe it's case with wordpress blogs where main content is usually in "content" div, so it's easy for them to find it.
Maybe G measure only content (kw density) inside that main part of page.


It could be good if you can somehow scan few pages inside domain and try to clean those repeated elements. I know it could take a lot of time to code that.

I used to think that too. Visible matches largely accounts for that, but truthfully... How many times have you or others been searching for something that appeared only in a single link on the bottom or on the side of the result page. Happens a ton to me still when I look for rapidshare links.

Google is indexing the whole page. It becomes more apparent when you analyze a much broader collection of pages. In fact... I'll start a whole new argument on BHW... I don't think a link at the top of the page carries anymore weight than a link in the footer. They both work equally as far as I can measure. Maybe my tools aren't precise enough to tell the difference, but for any practical purposes a link is a link to me anywhere on the page.
 
Keep in mind that this data represents hundreds of searches. If you calculate that average number of times the search terms appeared in the title you will most likely get a fractional amount. When you look at the charts understand that they are measuring occurrences as hundreds of different searches trend towards result 1 on page 1.

Once you wrap your mind around that nut then the data gets exciting. ;)

I understand that part, and that's why it's already exciting :p

I just wanted to know the general concept of how you calculate it...

For example:
Let's say the query is "white shoes for women"
And let's say the page's title that we want to analyze is "Great deals on white shoes for women"

Now, what's the keyword title match coefficient?
Is it 1 because the keyword occurs once throughout the title?
Is it ~0.6 because in the Title 4 words out of 7 match the search query?
What if shoes was mentioned twice in the Title, does that affect our number?
etc..

I hope I made my question a little bit more clear this time :)
 
I can't wrap my mind around the "nut" and I am excited! lol

Really an enlightening thread, even for being new here.
 
I'll start a whole new argument on BHW... I don't think a link at the top of the page carries anymore weight than a link in the footer. They both work equally as far as I can measure. Maybe my tools aren't precise enough to tell the difference, but for any practical purposes a link is a link to me anywhere on the page.

Hmm, I am not sure about this one, I have heard and read a lot about the difference in weightage of link juice coming from different sections of the page. Now I don't run any tests of my own, but I know the engines can differentiate http://www.seobythesea.com/?p=1093. I have even heard of sites getting rejected from Yahoo due to having too many footer links.

Also, as you know many people create themes for WordPress and insert their money site links in the footer for getting backlinks. So I decided to analyze one of them, it's called http://www.bromoney.com. Most of the themes they create have links to this particular site and it's inner pages, they target terms like "internet banking" etc. But I never see the site ranking on the first page. I have analyzed the top 10 sites for the keyword term and even though this site seems to have a better backlink portfolio with higher PR, mozRank, domain authority etc. than few of the to the top 10, it does not rank in top 10. And 95% of the backlinks to this site are footer links which were embedded in the themes that they distributed and this leads me to believe they are being treated differently. Now this is only one example, but I have seen this often. Footer links are surely not worthless, just that they may have differences in how they pass link juice (maybe relevancy plays a bigger role etc.), atleast that's what I think.

If I may ask, what setups do you have to measure such a thing and what sort of tests do you run?
 
I understand that part, and that's why it's already exciting :p

I just wanted to know the general concept of how you calculate it...

For example:
Let's say the query is "white shoes for women"
And let's say the page's title that we want to analyze is "Great deals on white shoes for women"

Now, what's the keyword title match coefficient?
Is it 1 because the keyword occurs once throughout the title?
Is it ~0.6 because in the Title 4 words out of 7 match the search query?
What if shoes was mentioned twice in the Title, does that affect our number?
etc..

I hope I made my question a little bit more clear this time :)

I understand much better now. Good question! I don't test cases like that... My phrases are one and two word phrases. things like cars, games, diamonds, lawsuit, tennis shoes, gold watch, sterling silver, surround sound,etc... for it to be a match in my data it has to be an exact match including the whitespace. (I do normalize and collapse whitespace for pattern matching). My matches are case insensitive.

PS... those aren't actually my words... they are like my words.
 
If I may ask, what setups do you have to measure such a thing and what sort of tests do you run?

I link to a website from the splash page of my PR8 at the top. I link to another one from the splash page footer. They both attain PR3 in the same time for the same keywords without any additional link building. I don't see a difference. Like I said... maybe my tools aren't sensitive enough. There might be value in being at the top of a buzzillion PR0 pages, I haven't tested that. I don't see it in my pages with weight. Footer is my playground, it has always delivered for me.
 
I understand much better now. Good question! I don't test cases like that... My phrases are one and two word phrases. things like cars, games, diamonds, lawsuit, tennis shoes, gold watch, sterling silver, surround sound,etc... for it to be a match in my data it has to be an exact match including the whitespace. (I do normalize and collapse whitespace for pattern matching). My matches are case insensitive.

PS... those aren't actually my words... they are like my words.

I am sure the same experiments carried out with longer tail keywords, maybe 3 to 4 words, would surely have much different graphs. Also for longer tail keywords the on-page optimization matters more and changes to meta tags, content etc. would lead to visible changes in rankings. Since not everyone is trying hard to rank for these longer keywords, I feel you would get more natural results. Whereas with 1 to 2 word keywords, well everyone is doing nothing but copious amounts of link building and on-page optimization would barely contribute to any change in rankings.

Over how long of a period did you carry out these experiments?

Also do you have probability distribution graphs for the data?
 
Back
Top