Authenticity (% unique 3-5 word phrases >60%)
I wish I understood this better. Maybe somebody who understands language and the way that Google works might be able to tell me exactly what this means. It is something to do with the uniqueness of the work and the way that Google parses text in 3 to 5 word blocks or phrases. Once again there are caveats regarding the inclusion or exclusion of stop words and the way in which these blocks are put together.
My shorthand version would be; "Unique content is better".
But I know there are some real experts out there who may well be able to give me and my readers are much better overview of what this actually means. Full disclosure here - I'm not really sure myself.
I'll add that this does not preclude spinning and re-purposing of content. Spinning well and keeping your work readable and contextual seems fine. However it must appear unique when published and meet the language and context rules here as well to add real value to your SEO campaign
It refers to word n-grams. n-grams are the basis of detecting similar or duplicate text. Copyscape uses an n-gram comparator algorithm but I would not bet that Google uses the same algorithm, though is probably still based on n-grams.
Given two sentences for example:
1. I would like to
buy myself some nice shoes for the
Halloween.
2. Since
Halloween is coming, I am going to
buy myself some nice shoes.
That is a 5-gram (n-gram, n=5) that is common. It represents 41% of the sentence. In practice, the algorithms for similarity can be smarter or dumber (copyscape one is pretty dumb) and required computational power increases greatly if you want to use smart ones. For example it could also check for unigrams (1-grams/words) that are common in the sentence, increasing the similarity metric. Plus other stuff. Also number of false positives can be pretty high. Beyond a certain threshold however you can assume it is spun/non-unique content. All this works if you actually do a one-to-one comparison but Google would have to compare every single new web page with all the trillions of indexed pages in its index, or at least with the shitload that match the same corpus (words/n-grams/topic). That would be impossible. Hence, Google either does not detect duplicates like this or it uses a different approach (i suspect the latter) which is possible because of the nature of their reverse index storage system.
From my experience the problem of detecting computer generated content is unfeasible though. Yes, you can find shitty auto-generated content but you can't find properly generated content. The computational resources required to generate content that is grammatically correct and is made of real-world n-grams (n-gram based Markov chains) are not very high. Not to mention that depending on the requirements you have, even simple sentence mixing/randomization gets indexed and ranks even though each sentence in the article is copied from some other article.
Badly/insufficiently spoon content actually can perform far worse than sentence mixing. Won't go into the technicalities of why that is so. Properly spun content however is and will be for the foreseeable future, completely fine. The point is not to spin blindly or superficially but to do it properly so that what results is an article that is unique enough and also is readable, makes sense and is useful just like a hand written article would be. I have wrote some threads and posts on spintax uniqueness, including a case study on a 500 words spintax at different spinning complexity. You may want to look it up if you're interested in the numbers (uniqueness % and number of articles that can be generated safely).
I hope that answers the questions you had. Also, great post BTW. Rep+ because was a good read! (edit: I didn't, seems I can't and have to spread some rep around first) I disagree with some things (obviously) and I will mention them here briefly without much explanation, so when you read them, remember SEO is very relative
1. Benefit of EMD - still strong, people just didn't adapted their SEO for EMDs - I have a slightly modified strategy for EMDs just to be on the safe site but never ever had EMDs penalized. Then again, many people do all sorts of strange and stupid things with their sites so it would be impossible to know for sure why others got hit so widely.
2. N0F0llow - I prefer <30%. I have sites with <10% doing great too. never did tests so I have no idea if 30% or 50% works better, I do have a hunch however that a difference here would be mostly irrelevant. what I can tell you for sure is that a high % of n0f0llow is bad, especially combined with an increased link velocity (=footprint of link spam).
Regarding the other points, I either agree, have no idea or think there is a potential for error because of the impossibility to properly isolate the factors/variables being analyzed and their relation to other factors. In other words, a result you see might seem affected by X when in fact it is affected by Y or a construct along the lines of (X+C-D)*Y where another factor (Y in this case) is a multiplying agent (has most influence) and might only apply if another factor or group of factors (e.g. W, T, U) produce a certain result. Or whatever other different reason that would make one believe something that is not accurate or correct.
As a final conclusion, what I can say is pretty much everything works. The key is to finding the balance and implementing things the way it works not the way it kills your sites. That being said, Google is a biased asshole - they make rules and then after a couple of years they penalize you for following them.
Roses are red
Violets are blue
I don't give a fu*k what Google wants me to do