- Jul 13, 2008
- 2,219
- 5,431
With the actual stats I have no insight what is happening. There isnt even a log after stopping. And I dont know if google was requested for each keyword and keyword result page...
Maybe im a perfectionist and want to have all results that are available... but at the moment the multiharvester gives no insight what its doing. Thats not good for me...
Ok here's the thing right, the multi-threaded harvester was written totally from scratch it wasn't simply adding a bunch of thread to the old one it's a total ground up rebuild.
Here is Sciborg's screenshot scraping at 2,133 URL's per second when it was set to 100 connections, and since it's been upped to 500 connections and i've seen just over 4,000 URL's scraped per second (1/4 Million URL's in 60 seconds).
Granted i've only ever used a few scrapers, but i've never seen numbers like that before so ScrapeBox is probably one of the fastest scrapers that exist. Doing that in about 3 days, plus piling on granular stats when the kinks are still being ironed out wouldn't be wise there would of been complaints everywhere and problem tracing would be a nightmare.
As you seen there was a socket timeout issue, IPC getting swamped etc and this is why the old harvester is still available in tandem.
There's numerous problems to solve when reporting what 2k threads are doing, in different states across 4 engines.
<insert Chinese proverb about patience>