xrummer maybe king BUT ..... AND ITS A BIG BUT
for what you pay for scrapebox 57$ for everything it does
+ free updates free addons + premium addons its not far
behind xrummer ...
The next version of ScrapeBox has an additional custom harvester:
Settings:
Harvester:
So you can train it to harvest from virtually any website with a search engine. Also it can harvest and test proxies, then refresh them every X minutes while harvesting. The original harvester hasn't changed if you still want to use it, you can switch between them in the settings.
Guys
Are there any instructions on how to use this part of scrapebox - I think it is what I am after.
I have 15000 sites that I am trying to scrape inner pages from with G and it gets only 10% through and stops...
Any help would be much appreciated
Thanks in advance guys
Ill do up a vid on the new harvester, but its pretty basic. You just go to settings and tick off to "Use Custom Harvester" and then it will use it when you click harvest.
However its not going to solve your google problem. Google is quitting because its blocking the IP(s). You need to set your harvester connections to 10% of your proxies. If you have 20 proxies, set your connections to 2. Etc....
You can also do a big scrape of internal pages, and then feed them all into the link extractor and extra internal links, and then feed the results back in and extract internal links again. Do that a few times and you will have a nice internal index of pages.
I don't know what are you talking about saying Scrapebox > Hrefer.
If you are going to do small, specific scrape - Scrapebox will help you, but you need private proxies for Google - all kind of shared and public will only make you nervous...
If you are going to collect list's for GSA/ any other link bulding tool, go for Hrefer... really, the cheaper version of Xrumer - $290 - still comes with hrefer. If you have to setup a huge scrape - scrapebox will fail with huge keyword list, scrapebox will fail with refreshing your proxies all the time. But hrefer will:
automatically refresh your proxies each X minutes
wont crash at huge keyword lists
wont fail with inurl: scrapes - even with public proxies and google!
will be 10 or more times faster than scrapebox
you can use sieve filter, example:
you want to scrape xpressengine's only, you setup a sieve filter to deny all urls not containing "document_srl" - clear results and month of scraping totally afk...
As for failing in URL - that's a function of Google banning the proxies, not to do with the tool.