working like a turd - rule 1 public proxies , if u ever scrape proxies there are never
more than 5k or so live any given time always been this way so scraping 92k as
test base well was a bit of a waste.
judging from your results if i have this correct - very few proxies worked
urls where shitty to rubbish which ive been saying all along. right
my service has averaged this for 2years.
you get the idea ..... now last 2 updates i get this sht and a support that never replys
now as of last update i get this - i stand by they ballsed something up
and if u define working better than it was while your own results show less
than 1k urls a minute thats shocking that tool should be and has done
around 100k urls min since launch ......1k isnt working better 1k is down
right trash , with respect
Because we share a proxy tester in common I will start there. When you run a search tree you basically have two types of search that you can do. These are a breadth first and a Depth first. For my own scraper I use a depth first search. I want as many IP's used by proxies as I can latch on to. Frequently, prior to deduping with my scraper I will pull 290k IP's; after deduping the number ranges between 40k and 110k. I am aware that the vast number of the IP's will be dead. That is the least of my concerns at this point. Getting usable proxies for scraping is my concern. Using the depth first search and the URL that I scrape from I may very well pull a number of the proxies that was in your morning list. Again usable proxies and not number of IP's is the goal. To illustrate the principle, I fired my scraper up and puled about 200k IP's tp test. After deduping, I had about 82k proxy IP's.
Here is the Bleach Proxy Checker result of testing those IP's:
As you can see, I have 1489 live proxies. Most of those will either be banned by some search engine, or be in some spam/bot list. But because I am only using them for scraping, I do not care what spam or bot list the may be in. When I dup a list with 82k proxies into Gscraper I frequently have 700 - 800 proxies that will work for my purpose. Most will be on port 80.
I agree that at any given time there will only be about 5000 proxies that the public can access and about three times that if you count the government, military, and closed proxies. The government, military and closed proxies I want to stay well away from.
Testing that large number of proxies, and using 5000 as a baseline, I have just scraped up 30 percent of the proxies. To me, it is worth it; to another, who knows.
(edit add: the number of proxies that would have passed Bleach would have been higher if I did not have Malwarebytes and PeerBlock blocking undesirable proxies)
------
Getting out of the proxy issue and into GScraper.
My comment that it works better than it was is based on a scraper that scrapes slowly is better than a scraper that does not scrape at all. The highest number of URL's that I have gotten for a while is somewhere in the neighborhood of 2000 URL's/min. So do I think GScraper is in top form? Very far from it. But these are not the only updates that have been bad; there have been several. Ever since somewhere around version 1.2.3.8 Hawke got married. Since that time Hawke has had less time to spend on Gscraper, and as a consequence GScraper has become progressively worse. Hawke has tried to buffalo people by adding new features to cover up the product deterioration.
It was the deteriorating product that caused me to purchase ScrapeBox along with some of the issues I discussed in PM.
When Gscraper was first released, support for Gscraper was good. Since that time, product support has become non-existent, and sometimes downright rude when he does answer. GScraper has enough telemetry in it to shame Microsoft and Windows 10. There is very little that is done on GScraper that is not sent to China.
I do not know what is going on in Hawkes personal life, and do not care. I suspect that he has gotten a job, or that he is over his head in coding. I do not remember the guys handle, but Hawke had a programmer working for him when the product was released. about the version given above, the programmer quit and Hawke was pretty upset about it. At one point, though weak in the wallet, I had offered to buy GScraper from Hawke. He never responded. As long as he has suckers paying him $66/mo. for his poor proxy service, he has some money coming in. However, because of the way that GScraper works using a proxy for one query and picking up the next proxy round robin, your service is better. I am not a customer of yours either. I know your reputation, and I know your proxies from the ones you publicly post.
As I told you in PM, I am reaching a point where I will have to make a decision that may contradict my ethics regarding GScraper. If I must make that decision, then it will be in favor of the community, though the code will be to an older version. Sometime between version 1.2.3.8 and 1.3.0.6 Hawke changed the Obfiscater from the DotNet obfuscater to dotNet Reactor. It is not hard to unpack, but it is time consuming. Presently, I am giving Hawke the opportunity to repair his product.
I really do not have a lot nice to say about Hawke or the current operation of GScraper.