Just emailed support regarding a bug crash report for the broken links checker, is there anyway I can recover the progress scrapebox was making before the crash? I was doing +1k URLs.
No, unfortunately not, there is no live logging in the broken links checker.
Thanks man. Where is the retry error settings? I honestly can't find them anywhere.
Well that would be because I fouled up, there is no retry in the link extractor, sorry. I was thinking of the link checker settings. However you can export all as rows, and then just excel to sort the status column. Then you can grab all the errors and then use them again.
How many links ~ per second/minute it can scrape with 10 proxies ? I'm thinking about buying vps and getting 30 proxies, but if 10 would to the job then it would be fine.
Not really sure, I don't measure the speed, but 10 proxies would go pretty fast, it should be pretty decent though. The VPS probably won't help much, but more proxies would. I assume you running it on a home machine now? Google is light weight on pages, so adding more proxies to your home machine would allow you to use more connections and go faster. So unless your home connection is just Very slow, like less then 2 or 3Mbit then buying a vps won't likely help yo u scrape faster, but it would let you post faster and link check and other things that are more bandwidth and resource intensive.
Do you know estimated time for v2.0 to be released?
I have no idea, and I don't think Sweetfunny is going to try and post that. I can tell you from having hired out my own tools, nothing ever goes as planned. I can also tell you that in times past when Scrapebox was planning to release things they might tell me date X and then they were either early or late. Because little things sometimes that you can't forsee will set you back, while sometimes you can get thru other things quicker then what you think. Something this large could have any number of things go wrong.
I would "suspect" that the first time they post saying "its going to be released at X" it will be withing 1-4 days prior to the actual release, that or it will be "Its available now" But I know they are working on it non stop, so its always forward progress. They have 5 years of work to redo, so I can't even imagine the work involved in starting from the ground up.
Can someone help me?
I've got 10 dedicated private proxies and 4 connections for google harvester. But it seems google bans my proxies very very fast. I can't even scrape beyond 1000 urls for a lot of different keywords. ( no not 1000 urls from one keyword, but 1000 from all the keywords). Is it because I scrape with "inurl:hidden" footprint? Also I use "inurl:%KW% -2014 -2013 -2012" and I get no results but that could be of the banned proxies.
Are you putting your keywords in quotes like that? Like are you actually using
"inurl:%KW% -2014 -2013 -2012"
or are you using
inurl:%KW% -2014 -2013 -2012
?
With 10 dedicated private proxies and 4 connections you will get all your proxies banned. I wouldn't think it would be within the first 1000 results, but its entirely possible, especially if you have been banned multiple times before. 10 proxies and 1 connection is too many with those kinds of footprints.
When using advanced operators like that you would want to either use 20-30 proxies with 1 connection
or go to settings and uncheck "use custom harvester" and "use multi threaded harvester"
and then set the delay drop down to RND (in the lower right hand corner of scrapebox)
and then go to settings >> adjust rnd delay range and set the min to probably 5s and the max to30 and start there and go up if need be.
Also if you are using
"inurl:%KW% -2014 -2013 -2012"
with the quotes, then your getting no results because non exist. Check your resulting footprints in browser to see if they return results or not.
What is your goal there, to minus out the years of 2012 - 2014? Because its also going to be ineffective at best. Its not going to minus them out based on googles timeline for those results, its just not going to give you pages that actually have those dates on them, which will help, but if a pages was found in 2012-2014 but doesn't list the year it was published on the actual page then its still going to be returned.
If you want true date timelines based on googles time frame search then you would need to build a custom string in the custom harvester for your time frames.
Yes your going to ask, "how do I do that". Im going to make a video and post it, but Ill not post a hundred screenshots in this thread in case thats not what your talking about. Also there is some reference here:
http://scrapeboxfaq.com/how-can-i-c...-the-automator-or-custom-harvester-for-google
But bear in mind that requires you use the custom harvester, you can NOT use the custom harvester AND use a Delay. So you would need to add 10-20 more proxies and set the custom harvester at 1 connection for that to work. Not completely sure how this will work in Scrapebox 2.0, and I will probably wait till it comes out to make a video on the time span, but if you want more info I can post screens.