i am getting a problem.
i loaded 15 private proxies , connection limit = 10
then loaded 12k keywords , i was able to get 600k results in 12 hours. it said harvest complete , but it left 9k keywords as "not completed"
i not need to scrape fast , i need to scrape slow (i won't mind scraping for 10 days constantly )
other settings (i was only scraping google)
search time = 1year old
multi threaded harvester proxy retries = 20 (max)
time out for harvester = 90 seconds (max)
can anyone tell how to let it to scrape all keywords ?
(i know there are many other scrapers to scrape fast but i am in slow scraping and i can't afford to lose even one result )
(it happens to every time sb leaves 50% of keywords) (i check proxies evey time after scrape and they are google passed )
-=-
Your connections are 1000 percent too high. Put connections to 1, not 10. You want a ratio of connections vs proxies at 1:10 or 1:15+ depending on your query. So you have 15 proxies so connections should be no more then 1. 20 proxy retries will be fine, since you only have 15 proxies anyway.
Hey any idea why the scraping stops after just couple of keywords?
I load like 20 keywords and it scrapes 2 and says " Harvesting Completed"
Go to settings >> use multi threaded harvester - and uncheck this. Also under that same settings menu uncheck the use custom harvester. Then try and harvest, there will be a status column.
What errors are you getting in the status column?
Here is a video that goes over how to do this and other troubleshooting for the harvester that might be helpful.
http://youtu.be/2QaLWgTXsRo
Thanks for your response. This is a lengthy process. I like the way scrapebox makes it easy to check pr. Just load urls, click on check domain or url pr, which allows you to quickly remove undesired domains from harvester. I am actually looking for something similar for index checking, trimming to root makes the process longer. I am not looking to indexed de-indexed domains, but to get rid of them and submit those urls whose main domains are indexed to the indexer service. Not sure if I have made myself clear enough.
Yes you made your self clear, and if you go reread my process and try it you will see that it will only be submitting the urls that have your links on them for indexing, not the root domain or any other domain on there.
Also I understand that this will take you an extra 2 mins a day to do, but if scrapebox goes in and adds shortcut features on everything that everyone has requested, menu items would be so long they would fall clear off the screen, and thats even if you had a 60 inch monitor. The point is, should scrapebox add such features, that are already available, to save you 2 mins, but that would bloat the software, cause possible bugs/crashes/conflicts for other users, and then cost them hours and hours and hours of support time, which in turn means hours and hours and hours less that they can develop the software, so you can save 2 mins? No.
Nothing personal, but keeping the software trim and running smoothly is more important to me then saving 2 mins. I have multiple processes that I would love shortcuts for like your saying, but thats just me. Ive seen people ask for these little short cuts a thousand times (no joke), literally the menu systems would fall clean off the screen if they were all added.
Go reread my process, it will do exactly what you want, providing you an end list that has urls that have your links on them, but that reside on domains that are not currently indexed. And it will take you less or the same time to do then to respond to this post.
Plus I don't work for scrapebox, so perhaps they will deem your addition important and add, thats up to them to decide, Im just saying that you can already do it, and it won't take long at all.