validseo
Senior Member
- Jul 17, 2013
- 928
- 551
SEO Content Tuning Experiment
Summary:
• Keyword stuffing works
• H1 not as important as people claim
• Keywords near top of page not so important, but not unimportant
• Nofollow links are much stronger than do follow links in on page content
• Keywords in alts tags do matter quite a bit
• Short domains do rank better
• Keywords in your forms... yup they matter
• Strong, bold, em BAD!!!! Underline Good???? Maybe.
• And a zillion more things in the data... I'm giving you ALL the data.
What I did:
I measured 200 different content factors for each of the top 10 Google search results for 25 VERY niche high competition keyword searches. This post contains some of the interesting highlights as well as ALL of the data.
About the sample:
I chose the niche of "DUI Attorney" because it is extremely high competition. I wanted to measure content tuning factors in a niche where the websites would be financially motivated to invest in getting FREE organic clicks. At $50-$200 per click in most major metros these websites would be very motivated to get every free click they can. I chose 25 metros so each search was similar to "dui attorney phoenix" or "dui attorney Chicago" etc. So the sample got lots of different websites with extremely similar search intent. The local block search results were ignored. For this study I am only interested in what is trending in the organic results.
Sample Size:
25 search terms in a very specific niche.
About the correlation:
Correlation % is calculated based on what percentage of the sample did the average of the first 5 results trend with the average of the second 5 results in favor of the factor. So for a factor like "Domain Name Length" where shorter is supposed to be better we want to see the average domain length decrease as you approach number one across all the keywords and visa versa for factors that should increase with rankings.
Because of the way we calculate and use correlation you should know that a correlation of 50% is random. It literally means that the factor trended properly for half of the sample. So the higher the correlation gets above 50% the more interesting it gets. When the correlation is smaller than 50% it most likely means that there was insufficient data on the factor in the sample.
If you wonder why the correlations fall into quantized amounts it is because of the sample size. Common fractions of the sample produce those like values.
Correlation != Causation:
Yes. Yes. I know. But unless you work at Google on their algorithms then correlation is the best empirical data you are going to get. Plus the people who gripe about correlation often forget that if it doesn't correlate then it probably isn't causal either SO correlation can vastly narrow the field of THINGS you need to guess about. Another saying anti-correlation people should remember: "If it doesn't correlate then you aren't using data." I'd also accept "If it doesn't correlate then you are just blind guessing."
About the Charts and Tables:
Most of the charts and tables I present in this post show the min, average, and max values for the specified factor across all 25 searches FOR the ranking the result had. The tables show the values and the charts show the lines, but the most important number is the correlation % which indicates how well or poorly the factor performed across the whole sample.
About Image Attachments:
BHW only allows 8 image attachments. I have a lot more than 8 things to show. There is also a KB size limit to the attachments. So I combined them into a single image. So I apologize but you'll have to open the image in a new window so you can see the charts and data as your read the post.
Disclaimer:
I did use software I created and some day hope to offer. I am not selling it currently and fully intend to have a paid membership when I do. I am just sharing the data and insights I am able to produce with the BHW community first.
Lets Begin
Myth: Keyword Stuffing Doesn't Work
Saying the keyword more appears to matter... It trends 22% better than random! It makes sense. At its core, Google Web Search is a typical document indexing engine. This kind of solution has existed since the 70's and they all suffer from the same kinds of limitations. For example, when all other things are equal whoever says it more wins. Look at those average and max values across the samples. That's a lot of stuffing, but at $170 CPC if it helps is it worth it? The myth that Google solved keyword stuffing and that it doesn't work is TOTALLY BUSTED. That doesn't mean it isn't risky and you wont get punished for it. You might.
Myth: H1 is the most important heading
In terms of Bing this myth is totally true. In terms of Google it is totally busted. H1 tags only correlated 84% of the time. H2 correlated 88%. H4 correlated 84%. When I treated H4-H6 as a group they correlated 96%... H1-H6 only correlated 88%. H1-H3 only correlated 80%. This makes sense to me. I believe Google uses Bayesian methods to identify web spam based on a training set. I bet that training set exploits the heck out of H1 tags thus diminishing its value as a signal. Just my theory... but H1 being most important... BUSTED
Myth: It Helps To Have Keywords Near Top
"Keywords near the top" correlates weakly... only 14% better than random. That is not a strong factor. I tested both full source and text with HTML removed. They produced the same result. As a factor it is plausible, but weak.
Links
Holy Guacamole! I guess you want a fair amount of no follow links on your pages. If we use the "do follow" correlation as the control then nofollow links are 20% more important! It is interesting that keyword matches in "on page links" don't correlate as strongly as either the number of do follow or nofollow links. Honestly, I'm not sure what this means.
Images
1. Keyword matches in image alt text trends 42% better than random! I guess a photo (with relevant alt text) is worth a thousand words (of content)!
2. Clearly the keyword matches in the alt text matters about 20% more than just having the images with alt text.
3. Since it is technically an image... the favicon... look how smoothly that average transitions... As factors go this is pretty good, trending 26% better than random. There are a lot of simple little things like this that you can and should add like privacy policy, terms of service, a copyright, apple touch icons. Everything a typical spammy web page might opt to not include.
Forms
There seems to be something about having a web form that uses your keywords. I guess that demonstrates some sort of intention to address a visitor need.
URLs & Domains
1. Keywords in the URL... only trends 10% better than random. Lower than I thought it would be.
2. Short domain name? Yup! There appears to be a pattern that shorter is better.
3. And keywords.html appears to have value too!
Page Size
Strangely the number of words and number of sentences in the page content correlated at 50% or completely random... not a factor, but the kilobyte size of page did trend 30% better than random. It appears you want to be right around 150Kb in page size. Weird.
Emphasis
Strong tags, bold tags, and em tags failed to correlate. They all 3 had values. It makes me wonder if over using these hurts your rankings. If so, it would be one of the few things I've ever measured that actively hurts rankings. The one outlier was italic tags which did trend as a factor 42% better than random. That's huge. I can't help but bring up my Bayesian theory again here. That could explain it.
Wrapping Up
As I promised, below is a link to a zip archive containing all the data. I hit the attachment limit for this post. In the archive there is a spreadsheet containing the 200 measurements for each result of the 25 search terms. There is an aggregate.xlsx file that contains the correlations across the whole sample set. There is A LOT more to discover in the data and I encourage you to take a closer look.
Let me know what you think?
Thanks!
All the data in a 6MB zip file: https://drive.google.com/file/d/0B9Xgfy4uuUsrdDN5b0xHV2F1WGs/view?usp=sharing
Summary:
• Keyword stuffing works
• H1 not as important as people claim
• Keywords near top of page not so important, but not unimportant
• Nofollow links are much stronger than do follow links in on page content
• Keywords in alts tags do matter quite a bit
• Short domains do rank better
• Keywords in your forms... yup they matter
• Strong, bold, em BAD!!!! Underline Good???? Maybe.
• And a zillion more things in the data... I'm giving you ALL the data.
What I did:
I measured 200 different content factors for each of the top 10 Google search results for 25 VERY niche high competition keyword searches. This post contains some of the interesting highlights as well as ALL of the data.
About the sample:
I chose the niche of "DUI Attorney" because it is extremely high competition. I wanted to measure content tuning factors in a niche where the websites would be financially motivated to invest in getting FREE organic clicks. At $50-$200 per click in most major metros these websites would be very motivated to get every free click they can. I chose 25 metros so each search was similar to "dui attorney phoenix" or "dui attorney Chicago" etc. So the sample got lots of different websites with extremely similar search intent. The local block search results were ignored. For this study I am only interested in what is trending in the organic results.
Sample Size:
25 search terms in a very specific niche.
About the correlation:
Correlation % is calculated based on what percentage of the sample did the average of the first 5 results trend with the average of the second 5 results in favor of the factor. So for a factor like "Domain Name Length" where shorter is supposed to be better we want to see the average domain length decrease as you approach number one across all the keywords and visa versa for factors that should increase with rankings.
Because of the way we calculate and use correlation you should know that a correlation of 50% is random. It literally means that the factor trended properly for half of the sample. So the higher the correlation gets above 50% the more interesting it gets. When the correlation is smaller than 50% it most likely means that there was insufficient data on the factor in the sample.
If you wonder why the correlations fall into quantized amounts it is because of the sample size. Common fractions of the sample produce those like values.
Correlation != Causation:
Yes. Yes. I know. But unless you work at Google on their algorithms then correlation is the best empirical data you are going to get. Plus the people who gripe about correlation often forget that if it doesn't correlate then it probably isn't causal either SO correlation can vastly narrow the field of THINGS you need to guess about. Another saying anti-correlation people should remember: "If it doesn't correlate then you aren't using data." I'd also accept "If it doesn't correlate then you are just blind guessing."
About the Charts and Tables:
Most of the charts and tables I present in this post show the min, average, and max values for the specified factor across all 25 searches FOR the ranking the result had. The tables show the values and the charts show the lines, but the most important number is the correlation % which indicates how well or poorly the factor performed across the whole sample.
About Image Attachments:
BHW only allows 8 image attachments. I have a lot more than 8 things to show. There is also a KB size limit to the attachments. So I combined them into a single image. So I apologize but you'll have to open the image in a new window so you can see the charts and data as your read the post.
Disclaimer:
I did use software I created and some day hope to offer. I am not selling it currently and fully intend to have a paid membership when I do. I am just sharing the data and insights I am able to produce with the BHW community first.
Lets Begin
Myth: Keyword Stuffing Doesn't Work
Saying the keyword more appears to matter... It trends 22% better than random! It makes sense. At its core, Google Web Search is a typical document indexing engine. This kind of solution has existed since the 70's and they all suffer from the same kinds of limitations. For example, when all other things are equal whoever says it more wins. Look at those average and max values across the samples. That's a lot of stuffing, but at $170 CPC if it helps is it worth it? The myth that Google solved keyword stuffing and that it doesn't work is TOTALLY BUSTED. That doesn't mean it isn't risky and you wont get punished for it. You might.
Myth: H1 is the most important heading
In terms of Bing this myth is totally true. In terms of Google it is totally busted. H1 tags only correlated 84% of the time. H2 correlated 88%. H4 correlated 84%. When I treated H4-H6 as a group they correlated 96%... H1-H6 only correlated 88%. H1-H3 only correlated 80%. This makes sense to me. I believe Google uses Bayesian methods to identify web spam based on a training set. I bet that training set exploits the heck out of H1 tags thus diminishing its value as a signal. Just my theory... but H1 being most important... BUSTED
Myth: It Helps To Have Keywords Near Top
"Keywords near the top" correlates weakly... only 14% better than random. That is not a strong factor. I tested both full source and text with HTML removed. They produced the same result. As a factor it is plausible, but weak.
Links
Holy Guacamole! I guess you want a fair amount of no follow links on your pages. If we use the "do follow" correlation as the control then nofollow links are 20% more important! It is interesting that keyword matches in "on page links" don't correlate as strongly as either the number of do follow or nofollow links. Honestly, I'm not sure what this means.
Images
1. Keyword matches in image alt text trends 42% better than random! I guess a photo (with relevant alt text) is worth a thousand words (of content)!
2. Clearly the keyword matches in the alt text matters about 20% more than just having the images with alt text.
3. Since it is technically an image... the favicon... look how smoothly that average transitions... As factors go this is pretty good, trending 26% better than random. There are a lot of simple little things like this that you can and should add like privacy policy, terms of service, a copyright, apple touch icons. Everything a typical spammy web page might opt to not include.
Forms
There seems to be something about having a web form that uses your keywords. I guess that demonstrates some sort of intention to address a visitor need.
URLs & Domains
1. Keywords in the URL... only trends 10% better than random. Lower than I thought it would be.
2. Short domain name? Yup! There appears to be a pattern that shorter is better.
3. And keywords.html appears to have value too!
Page Size
Strangely the number of words and number of sentences in the page content correlated at 50% or completely random... not a factor, but the kilobyte size of page did trend 30% better than random. It appears you want to be right around 150Kb in page size. Weird.
Emphasis
Strong tags, bold tags, and em tags failed to correlate. They all 3 had values. It makes me wonder if over using these hurts your rankings. If so, it would be one of the few things I've ever measured that actively hurts rankings. The one outlier was italic tags which did trend as a factor 42% better than random. That's huge. I can't help but bring up my Bayesian theory again here. That could explain it.
Wrapping Up
As I promised, below is a link to a zip archive containing all the data. I hit the attachment limit for this post. In the archive there is a spreadsheet containing the 200 measurements for each result of the 25 search terms. There is an aggregate.xlsx file that contains the correlations across the whole sample set. There is A LOT more to discover in the data and I encourage you to take a closer look.
Let me know what you think?
Thanks!
All the data in a 6MB zip file: https://drive.google.com/file/d/0B9Xgfy4uuUsrdDN5b0xHV2F1WGs/view?usp=sharing
Attachments
Last edited: