Scrapebox 90% Duplicate Urls / Domains

baedorf

Registered Member
Joined
Mar 6, 2011
Messages
78
Reaction score
2
Hi,
today i had harvesting for 6 hours.
Result 2 Million Urls.
But After removing duplicate Urls and Domains i got only 30K Results.
How can i keep if i harvest 2 Million Urls 50%?
everyone is telling keywords are Key.
but i dont know how i should do it otherwise.
I have used a dictionaryunique wordlist without unique Results.
thanks for all help! keep scrapebox Junkie!
 
How deep into Google are you scraping? If your going max e.g. 9999 it will show lots of dupes so try less if that's the case.
 
Hi,
where can i change this setting?
Results i have set to 1000.
I thought google and Yahoo could only scrape First 100 Results?
 
Also, you need to use more specific keywords. General keywords that are similar are going to turn up the same results.
 
I'd say thats normal. You can only scrape first 1000 results. Out of those many expect that many duplicates.
 
You need to use bigger word and phrase dictionaries in your scraping. It can often take days or even weeks depending on your server speed to build massive lists. 100K - 1MM unique urls in size.
 
Personally I'd change it to top 100.

Hi,
where can i change this setting?
Results i have set to 1000.
I thought google and Yahoo could only scrape First 100 Results?
 
You need to use bigger word and phrase dictionaries in your scraping. It can often take days or even weeks depending on your server speed to build massive lists. 100K - 1MM unique urls in size.

Exactly.... welcome to the wonderful world of Scraping. Use many many specific keywords, mash them up with other keywords..... wait for days, then watch 75% go up in smoke as dupes.
 
wide range of keywords both long and short with a big footprint.
Your always gona have alot of dupes, part of the game dude.
 
This is even more true now that google shows pages of the same sites over and over again if you manually search a term and click on pages 2, 3, 4, etc.

I'm no sb expert, but I'm good, and yet I've had a great comment on a high pr edu page removed because sb found that same site for a completely unrelated search on a different project. sb seems to be almost useless for this now, at least for scraping google, but the other engines don't allow as many specific parameters that you used to be able to use in google.
 
Back
Top