What Should be best scrapebox settings for scraping web 2.0

alishakapoor

Junior Member
Joined
Sep 23, 2015
Messages
181
Reaction score
58
Hi Everyone,

I have 10 private dedicated proxies for my scrapebox. After scraping for 2 days I hardly get any credible results. Only 10k urls scraped. Can anyone guide me where I'm going wrong or what setting should I change so that I can get some credible results out of it.

Thanks
 
Well most proxies nowaways would hardly work on Google scraping as the big G is quick to block the IP's access to their site, once they sense automated queries being sent to their search engine. You can try expanding the number of proxies on your list and go for Yahoo and Bing, and they still do work accordingly. Otherwise, if you have proxies that work on G, you can in fact search for 'past' search results, for specific date/year, and you'll find alot more results from that.
 
If you want to scrape google directly, you will need a bunch of private proxies to get a few millions of blogs. Also your list of keywords is important. Forget about niche related keywords because most niches are not to be found on expired (deleted by their owners) blogs.
The other way is try the other search engines that are not as strict with connection number. Some even get their results from google, so you can benefit from that massive index. Some will need a few adjustment but there are plenty of tutorials on how to scrape with scrapebox.
 
Hi Everyone,

I have 10 private dedicated proxies for my scrapebox. After scraping for 2 days I hardly get any credible results. Only 10k urls scraped. Can anyone guide me where I'm going wrong or what setting should I change so that I can get some credible results out of it.

Thanks

Well its hard to say as you didn't really share much info about your setup. Like if you have an insane delay setup, or if your proxies are blocked, or how many you have or what engines your scraping or what kind of footprints you are using etc...

So without lots of details I can't really give accurate help, I can only guess.

But my biggest recomendation would be to try bing. I assume blocked proxies is your biggest issue and bing has VERY lenient bans. So start there. Else check out your help >> show error log (assuming your using scrapebox for windows and not scrapebox for mac) and go to the harvester error log. What are the errors?

503 and 302 for google are ip bans
 
Well its hard to say as you didn't really share much info about your setup. Like if you have an insane delay setup, or if your proxies are blocked, or how many you have or what engines your scraping or what kind of footprints you are using etc...

So without lots of details I can't really give accurate help, I can only guess.

But my biggest recomendation would be to try bing. I assume blocked proxies is your biggest issue and bing has VERY lenient bans. So start there. Else check out your help >> show error log (assuming your using scrapebox for windows and not scrapebox for mac) and go to the harvester error log. What are the errors?

503 and 302 for google are ip bans


Thanks loopline for the revert,
Well, my setting are pretty much default only exception is delay time set to 20 secs. I can say my ips are not blocked because even if I complete and check results, google does show no. of crawled results (its not zero). I’m scraping Google, Bing, Yahoo and Ask, footprints used are site:blogspot.com (keyword), site:tumblr.com (keyword) and so on.
 
That doesn't make a lot of sense, bing is super lenient. Its possible there are not that many results of course. Are these public or private proxies?

I would uncheck the use proxies box and try a sample run with just bing, is it in line with the results your getting or much better or worse?

Don't stress about an ip block, I know of people scraping over 1 million results from bing with no proxies.
 
That doesn't make a lot of sense, bing is super lenient. Its possible there are not that many results of course. Are these public or private proxies?

I would uncheck the use proxies box and try a sample run with just bing, is it in line with the results your getting or much better or worse?

Don't stress about an ip block, I know of people scraping over 1 million results from bing with no proxies.

Thanks alot Loopline, scraping only bing worked for me. Though I would still like to figure out a way where I can scrape google without getting blocked.
 
Back
Top