- Oct 5, 2009
- 2,363
- 1,207
Is it normal that on average 99% of my harvested URLs are dupes? I'm going for a spam list so they're all highly unrelated (as in totally different niches) keywords, and yet I'm having an enormous about of duplicates. It's depressing to watch my 1m scrape lists go down to a few thousand.
Yeah it can be depressing, but still, at least it's all automated and you just have to click buttons.
Couple have given you some tips. Here's another thought - if you are scraping all four engines, you're gonna get a lot of overlap from them. Nothing wrong with that, and sometimes that diamond in the rough is only indexed by Y! or bada bing or teh googz, but sometimes for a quick and dirty scrape I just harvest from one of them (guess which one? that would be teh guh-oogles). Your dupe rate will be lower that way if you want to try it.
Also I've noticed that BE you end up with a lot of dupes just because BE sucks so badly.