Lot of dis-information here, and muddled shit from that fat prick Matt Cutts. Google can't fookin read ffs, they have scripts written in whatever language, python or something else. If you were using some document level analyzer that could match blocks of similar text, that might be possible, but theres loads of shit like this. For example in general human language we use the same words and phrases over and over again. Other times you get people copy pasting articles, news sites, press release sites etc etc.
On the other hand let's say you have 10 sites, numbered 1-10. Site number 3 was the original content the others are duped, but the first crawl was on site 4 and site 5 was the first link, how the fuck is that palace of bullshit that is google gonna figure that out? As in which is original? People tend to think Google is borg. I say BS, corporate centric businesses are often like this.