Scrapebox Harvesting: Always low results. Suggestions?

With regards to harvesting URLs, I've now figured it out.

Only use Yahoo and do not use proxies.

The Yahoo API allows each IP to do 5000 queries per day; however; Scrapebox uses about a dozen APIs which rotate every request.

Basically - you should be able to harvest around 5 million URLs a day using this method, which for the majority I think will be more than enough.

I did a scrape earlier today from about 2000 keywords and got about 500,000 URLs... unfortunately, after duplicates were removed, I was left with about 140,000 but that just means I need to use more keywords or use a more diverse range of keywords. Either way, it's just a case of scaling it up to get more results - whereas before it was a proxy issue.

You could use private proxies if you wanted more than 5 million. This means not using Google, but even Scrapebox themselves have said just using Yahoo is the best approach.

Dude, thanks for the info. That did the trick :)
 
It's always a rule to have faster proxies to get good harvesting results. Even though I got private proxies, I'm just using them for commenting and not scraping so it would not be Google banned. So faster public proxies is a way to go. But they become dead fast too.
 
With my foot print manually typed in Google I get 2500 results but in Scrapebox with private proxies I get 400 results? Is this a proxy issue?
 
With my foot print manually typed in Google I get 2500 results but in Scrapebox with private proxies I get 400 results? Is this a proxy issue?

Did you try click through to the end of Gs serps manually to doublecheck? Sometimes after 200-300 results you get the odd "repeat search with omitted results included" message showing up in G.
 
My advice is sod Google off and just use Yahoo without any proxies. It's worked wonders for me.

I harvested over 900,000 unique URLs last night, currently posting to them 150,000 at a time - via public proxies (I don't use Scrapebox source).
 
Did you try click through to the end of Gs serps manually to doublecheck? Sometimes after 200-300 results you get the odd "repeat search with omitted results included" message showing up in G.

Yes, that was it! Thanks.
 
I tried disabling multi harvester and now it is fetching some results from Google. Unfortunately I cannot use Yahoo as I need to scrape some number data using Google's ".." operator (example age from 25..30) . Any idea if there is an equivalent operator for Yahoo?
 
Do you guys use AOL too? Was wondering whether to include it or not
 
Scrapebox + Yahoo harvesting doesn't appear to work very well at all now.

Even with 150+ proxies and 10 connections it only gets a few thousand at a time.

Yahoo must have severely limited things on their API.
 
Scrapebox + Yahoo harvesting doesn't appear to work very well at all now.

Even with 150+ proxies and 10 connections it only gets a few thousand at a time.

Yahoo must have severely limited things on their API.

Same here.
 
When doing the no proxy Yahoo method, should I ramp down the max connections? I had them at 14 and seems I got banned within about 30 seconds, but I harvested a good bit too.
 
With regards to harvesting URLs, I've now figured it out.

Only use Yahoo and do not use proxies.

The Yahoo API allows each IP to do 5000 queries per day; however; Scrapebox uses about a dozen APIs which rotate every request.

Basically - you should be able to harvest around 5 million URLs a day using this method, which for the majority I think will be more than enough.

I did a scrape earlier today from about 2000 keywords and got about 500,000 URLs... unfortunately, after duplicates were removed, I was left with about 140,000 but that just means I need to use more keywords or use a more diverse range of keywords. Either way, it's just a case of scaling it up to get more results - whereas before it was a proxy issue.

You could use private proxies if you wanted more than 5 million. This means not using Google, but even Scrapebox themselves have said just using Yahoo is the best approach.

Tried this and was banned in seconds on yahoo. Alternatively I used 5 decent public proxies picked up through sb, checked them, made sure they were under 3000ms and ran them through google and im at 20,000 urls with just 93 keywords right now. Looks like yahoo is banning all my proxies quick too....

Perhaps im doing something wrong? Using all standard settings.
 
Last edited:
Scrapebox + Yahoo harvesting doesn't appear to work very well at all now.

Even with 150+ proxies and 10 connections it only gets a few thousand at a time.

Yahoo must have severely limited things on their API.

Ah....figures im late to the party. Shoot....
 
footprints
proxies
timeouts
connections

thoae are the first things you should be looking at and adjusting
 
No it still works. Just disable multi-threading or turn connections down to about 2 while increasing the time between searches to about 5 seconds. I'm still working on improving this method too.
 
You're going to need a lot more than 40 proxies if you're wanting more than a few hundred results.

Look for proxy sources here on the forums. I found a BIG list (needed a little cleanup) but it's done me well and I can get 2k-3k working public proxies on every scrape. Of course then checking all of them to see if they work is kind of a bitch... But they're free.
 
Tried to harvest today a very simple pattern:

Code:
intitle:keyword inurl:path

The proxies are public, however they are unique and regular SB user doesn't have access to them. So they don't saturate that fast.

First impression: Yahoo sucks. try the pattern above in Yahoo search and it will bring 0 results, try the same in Google and you will get 1000. I am speaking about the normal search - without the Scrapebox.

Second impression: Google harvester doesn't work as it should. Even after disabling proxies and setting threads to 1 the harverster got stuck after about 1200 results. My IP wasnt banned, I was searching freely from the same PC using browser. With proxies it didn't work at all, i checked some of the proxies and they weren't banned.

FYI
 
Last edited:
I am experiencing similar problems. I have owned Scrapebox for 6 months, but havent used it a lot (big mistake).

Now that I want to actively use it to scrape blogs where to comment to, I can't scrape basically anything.

Also, I've noticed that if I put in the keywords for example intitle:Barack Obama, most of the results will not have that in the title. Which is very odd to me.

1/ I have noticed the same . I am getting a lot of URLs , which does not have the KW in the URL, even though I used INURL operator.
I am guessing that as soon as the quality of the proxies start to deteriorate, you are starting to get not that "adequate" results.

2/ Lately I have been getting really Bad harvesting results. A week or two ago it was all good.
2/1 With Using Public Proxies ( not form internal SB list) The Google Pass rate is about 10-20 %.
IP test goes even lower.

I have used 2 additional sources of proxies:

1/ 125 filtered proxies daily-the result is the same. Google pass test ~10-20%

2/ Another plugin that scrapes proxies and filtered them -even with these proxies My PASS rate is 10 % or so.

I do not want to mention the sources of the Proxies, cause I am not sure they are to blame in the first place.

I use the Local Library, so I guess most of the time my external IP appears as the Public IP. Unless, somehow The Library router itself shuts me down or My Internal IP leaks out to Google. :).., all of these are just my theories though...:):eek:


When I Harvest URLs in SB without proxies, the results are awesome.., of course not until I get shut down.., which is normal.:p


After all of these issues I am thinking of getting Private Proxies and give it a try.

I use deep searches - long Footprints with G Operators , harvest targeted URls for only manual commenting to my main sites.

Any ides or suggestions on some of the above mentioned?

Appreciate,
 
Back
Top