spiritfly
Regular Member
- Apr 30, 2011
- 302
- 132
Im using back connect proxies, but I pulled all 2200 ish footprints from GSA, added them on 24 hour google (which netted 200K results) and scraped google. 100 connections and also 25 connections. When it gets near the end the threads do decrease but nothing like your saying, it reaches the point where all keywords are active in a thread and decreases from there till its done. Its quite fast for me.
Ill keep testing, but I can't reproduce what your talking about, grant it Im working on lower connections, but as you say it should be worse. Although you are running 5000 connections right? I mean at a point you are going to have 5000 active queries and its going to reach the end of the list of keywords and the last 5000 are going to have to finish, which may take a bit, but I would assume not hours like you are experiencing.
That's right. I'm using 5000 threads (now experimenting with 7k and 8k). I guess it's a combination between using public proxies and large amounts of threads. My guess for this is: when the keywords get less than the max threads set some of those public proxies go bad and scrapebox is waiting for the timeout until it declares it bad and switches to another proxy. So from 5000 public proxies probably a lot go bad and scrapebox has to wait for each one some time until it can switch to the next one. And if it switches to another bad proxy it will still have to wait until it gets to the next one. And the more bad proxies + the more threads = the more the waiting. This has to be the cause for this.
I'm getting like 2k urls/s average until the point where the threads start to decrease. From there on until the end the average speed goes to around 50 urls/s. I really have to do something about this, but unfortunately this process is completely in scrapebox control. I cannot add any external script or anything really to stop the process before it starts decreasing the threads. I'm still thinking how can I get around this.
I really can't see how to solve this. It represents a real issue for public proxy users. Maybe there should be different modes of scraping in scrapebox when using public vs private proxies. Private proxies are a bit expensive for what I'm trying to accomplish and many users are still using public proxies for scraping. Can anything be done to improve public proxies scraping with massive thread number?
Last edited: