google attracts some of the brights minds in coding, as a matter of principle & professional vanity, no engineer is going to do shit manually, when there are solutions already available to scale with code.
2 years you could use google doc's for readability test algo's like ARI & flesch?kincaid to cleanup your auto-gen, but they removed it as it was giving to much intel away.
when you run crappy spin thru readability tests it stands out like dogs balls, so if you where to do a statistical analysis of total posts on site for readability, and it pulls a site-wide general low score....thats a likely autogen footprint.
n-gram databases google released in 2005 to outside researcher where huge, now they could be 1000's of Gb, natural language processing is one of the top 3 areas of research at google.
as a noob layman, with no training in linguistic, when you spend 5 sec looking at most ALN posts you can't help but trip over examples of phrases/sentences that drunk homeless bums on acid don't even use...thats a likely autogen footprint.
its pretty much statistically fool proof when a site has 1000 posts, that xx% has shit readability & phrases that epileptic monkeys in a washing machine spin-cycle could'nt even type, autogen is at play.
-datacenter scale statistical linguistic machine learning- VS -TBS autospin-
whos going to win?