How Google Calculates Duplicate Content

Preon

Senior Member
Joined
Apr 6, 2016
Messages
992
Reaction score
1,724
There's an article over on seroundtable.com (https://www.seroundtable.com/google-dupe-detection-canonicalization-30376.html) which talks about how Google calculates duplicate content.

Essentially they create a hash of the content of the page and compare it to existing hashes. Any matches are seen as duplicate content.

This also means it's probably quite easy to fool. A few minor changes should result in a completely different hash value.

However, it's likely we don't have the full picture and there's probably more at work here. I wonder if it utilizes indexed document matching in order to detect partial matches. So, even if an article has its paragraphs rearranged, it would still come up as a 100% match as the content is the same.....

I'm not sure how best to test?!
 
Especially if it's done sentence by sentence :smirk:
 
Well, that's a massively bad thing to put in the public space for them, a good thing for us.

It could be a situation where they make a checksum for every xxx words so they can get some partial matching, or it could be a custom checksum where they reduce every word to 1 letter and smush it together. This would help them to establish partial matches really well.

However, why it's good for us is that we now understand spun content is truly superior and doesn't need that extremely high level of spinning. This makes the spun articles reasonably readable and may pass off as a bit of bad grammar but as a unique article without sweating over synonyms and such.

Slowly but surely google is a falling company. Launch your search engine boiz, the market is about to be litty.
 
They are talking about duplicate content on the same website - canonicalization.

They are not talking about content that has been copied and published on a different website.
 
They are talking about duplicate content on the same website - canonicalization.

They are not talking about content that has been copied and published on a different website.

Could also be used across the wider web. Or at the very least highly suggestive of the reason, they don't penalize duplicate content across domains.
 
They are talking about duplicate content on the same website - canonicalization.

They are not talking about content that has been copied and published on a different website.
I would imagine they would use the same methodology.
 
I'm also having this problem, I really want to filter out the same information as quickly as possible, hope someone can help me.
 
What we need to think and rethink is how good is the spinning technology right now ? They are good, right ? Almost as good as unbelievable spun content which can match original at any stage provided we review it manually else there are still chinks with them and can be spotted. So do we really think the AI of google cannot spot it and will believe in some hash#. Those are just theories.
 
And you
Well, that's a massively bad thing to put in the public space for them, a good thing for us.

It could be a situation where they make a checksum for every xxx words so they can get some partial matching, or it could be a custom checksum where they reduce every word to 1 letter and smush it together. This would help them to establish partial matches really well.

However, why it's good for us is that we now understand spun content is truly superior and doesn't need that extremely high level of spinning. This makes the spun articles reasonably readable and may pass off as a bit of bad grammar but as a unique article without sweating over synonyms and such.

Slowly but surely google is a falling company. Launch your search engine boiz, the market is about to be litty.
You think readers will engage with spun content?
 
What we need to think and rethink is how good is the spinning technology right now ? They are good, right ? Almost as good as unbelievable spun content which can match original at any stage provided we review it manually else there are still chinks with them and can be spotted. So do we really think the AI of google cannot spot it and will believe in some hash#. Those are just theories.


Somewhere along the lines google have said, sites with entirely copied content are treated as spam sites as they add no value. Then we started spinning articles. We started getting really complex with it as we needed large numbers of articles to keep up with the demand of GSA.

Google understood the task at hand, GSA spams and it spams crappy articles. I believe they took some form of action against this. I haven't been keeping up. But the action they took If i remember correctly employs the use of N-Grams to understand the context of the article and a few other things to try and understand the grammar. If the grammar is too trash then your article can be deemed as trash.

With this publication, he makes it known that they use hashes to compare articles(which makes sense as this is the least resource demanding way to do it) and that suggests, we no longer need deep nested spins. Especially since we now aim to build low volume strong links. Better quality spun content might have higher index rates.
 
And we all missed that part where he said to rank well in search engines a minimum word count of 3000 expected :cool: we are still stuck on spinning:D
 
Google creates a Checksum for each page, meaning a unique fingerprint based on the words of a particular page. By using checksums of multiple pages, Google can identify the pages that have similar content. To do so, Google collects small-sized data derived from a set of digital data with a purpose to identify flaws that may have occurred during the time of transmission or storage. Additionally, checksums verify the integrity of data available, but it may sometimes fail to examine its authenticity.
 
There's an article over on seroundtable.com (https://www.seroundtable.com/google-dupe-detection-canonicalization-30376.html) which talks about how Google calculates duplicate content.

Essentially they create a hash of the content of the page and compare it to existing hashes. Any matches are seen as duplicate content.

This also means it's probably quite easy to fool. A few minor changes should result in a completely different hash value.

However, it's likely we don't have the full picture and there's probably more at work here. I wonder if it utilizes indexed document matching in order to detect partial matches. So, even if an article has its paragraphs rearranged, it would still come up as a 100% match as the content is the same.....

I'm not sure how best to test?!
You can fool Google if that is duplicate content or not, but if there is no new information added, and everything you said adds no value for a particular query, chances are that you will have difficulties to get ranked high.
 
You can fool Google if that is duplicate content or not, but if there is no new information added, and everything you said adds no value for a particular query, chances are that you will have difficulties to get ranked high.
It's probably not as simple as you're making out.

Take Joe Blogger, an expert in Accordions, who writes perfectly written guides on choosing the finest accordions, his written word is like honey for your eyes. He posts his articles on his brand new blog, accordionbuyer.com.

Along comes Max Spinner, a writer for a media conglomerate who owns one of the largest music websites in the world. Max takes Joe's articles, spins them very lightly, and posts them on 'numberonemusicblog.com'.

Who will outrank who?
 
It's probably not as simple as you're making out.

Take Joe Blogger, an expert in Accordions, who writes perfectly written guides on choosing the finest accordions, his written word is like honey for your eyes. He posts his articles on his brand new blog, accordionbuyer.com.

Along comes Max Spinner, a writer for a media conglomerate who owns one of the largest music websites in the world. Max takes Joe's articles, spins them very lightly, and posts them on 'numberonemusicblog.com'.

Who will outrank who?
I'm not making it simple, just tried to simply answer the matter, which is definitely not simple at all, especially having in mind that Google isn't that good at recognizing duplicate content, at least at the moment.

When it comes to your question, of course, the answer should be Joe Blogger. On the other hand, it depends on other factors as well. For example, if Max Spinner's site is better than Joe's: well-optimized, good looking, fast-loading, and with a great UX, Google will probably find and rank it more easily.
 
I'm not making it simple, just tried to simply answer the matter, which is definitely not simple at all, especially having in mind that Google isn't that good at recognizing duplicate content, at least at the moment.

When it comes to your question, of course, the answer should be Joe Blogger. On the other hand, it depends on other factors as well. For example, if Max Spinner's site is better than Joe's: well-optimized, good looking, fast-loading, and with a great UX, Google will probably find and rank it more easily.
Absolutely should be Joe Blogger.

But my experience and I'm assuming yours, would suggest the other outcome is more likely. Given the other site has far more authority. Google has said themselves that duplicate content can outrank the original if the original website has trust issues.

I'm not trying to be an arse about it. But, Google is not especially smart nor fair, just look at how well Pinterest dominates the SERPs using what is essentially duplicate content.
 
Absolutely should be Joe Blogger.

But my experience and I'm assuming yours, would suggest the other outcome is more likely. Given the other site has far more authority. Google has said themselves that duplicate content can outrank the original if the original website has trust issues.

I'm not trying to be an arse about it. But, Google is not especially smart nor fair, just look at how well Pinterest dominates the SERPs using what is essentially duplicate content.
Definitely, that's what I said, Google is everything but fair here.

On the other hand, if you are the one to create great original content, and have your website well-optimized, you will be the winner, just like Joe Blogger, if he did the same. Besides this, if his content is so great and shareable, and he has a great site, chances are that nobody could out-compete him for the keywords he focuses on. Even with the imperfect Google
 
Definitely, that's what I said, Google is everything but fair here.

On the other hand, if you are the one to create great original content, and have your website well-optimized, you will be the winner, just like Joe Blogger, if he did the same. Besides this, if his content is so great and shareable, and he has a great site, chances are that nobody could out-compete him for the keywords he focuses on. Even with the imperfect Google
If he's new though, he has no authority.

I have firsthand experience of my content being copied by a larger site months after the original article was posted. The copied content outranked my original. This is without spinning.

FYI, it was streetinsider.com that copied my content, but they kindly took it down when I emailed them.
 
Back
Top