With regards to harvesting URLs, I've now figured it out.
Only use Yahoo and do not use proxies.
The Yahoo API allows each IP to do 5000 queries per day; however; Scrapebox uses about a dozen APIs which rotate every request.
Basically - you should be able to harvest around 5 million URLs a day using this method, which for the majority I think will be more than enough.
I did a scrape earlier today from about 2000 keywords and got about 500,000 URLs... unfortunately, after duplicates were removed, I was left with about 140,000 but that just means I need to use more keywords or use a more diverse range of keywords. Either way, it's just a case of scaling it up to get more results - whereas before it was a proxy issue.
You could use private proxies if you wanted more than 5 million. This means not using Google, but even Scrapebox themselves have said just using Yahoo is the best approach.
Dude, thanks for the info. That did the trick