- Jul 13, 2008
- 2,219
- 5,431
I am killing search engines
Q miss: What is the best value to put to harvest an amount for per keyword?
When I put 1000 times google let me harvest for about 10 keywords then it stops and block my ip for a while.
so what is the best number i should put?
check this out
![]()
Holy shit lol, you harvested almost half of Bings entire index.
So running it on the server is fast? Good stuff.
With Google not much can be done when using a single IP except change the delay so you don't scrape it as fast. The only good way to scrape large amounts of data from Google is proxies.
With Bing and Yahoo, ScrapeBox is using their API's which allow far more queries.
Sweet, Is there a way to "suppress" previous blogs that I've commented on that I don't want to duplicate the comment?
Maybe a way to load a previous blog file in a "suppression" window so when Harvesting new blogs it don't load any that are in the suppression window.
Thanks.
Sure you can do that in 2 ways, you can suppress either the URL's only.. To do this you need to have URL lists you have commented on saved on your PC. When you have scraped fresh URL's to comment on, "Select the URL lists to Compare"
When you click that, if there is any URL's sitting in the harvester that are also found on your old comment files they will be removed from the harvester leaving you with just unique URL's.
The other way, you can block whole domains by putting the domain name in your Blacklist.txt file. When you do this, you won't be able to harvest or comment on any URL belonging to that domain.
For instance i have Matt Cutts URL in the default blacklist so nobody spams him.
Proxies are problematic for me. I'm getting Error 404 messages when I check proxies,
usually a "connection aborted on request" message. Any ideas on how to fix, or should I obtain my own proxies?
Have you updated to v1.9.21?
This was released about ~10 hours ago, and contains a fix for the external judge which was causing the 404's.