ScrapeBox Issue with harvesting

MichaelAdv

Newbie
Joined
Mar 6, 2024
Messages
21
Reaction score
3
I just got scrapebox, I set it up with about 14k Keywords to scrape for websites, Using 10 proxies by MPP (MyPrivateProxy)
when scraping google wont scrape above 200-300.

when I went into the error log this error kept repeating itself - -1 Error receiving data: (12002) The operation timed out.

Tested proxies they passed tests.
They are new bought today.

Scrapebox wouldn't scrape at all from bing aswell.


Anyone has a clue?
 
It was scraping perfectly for me. Contact the Scrapebox support, they will assist you within 24 hours. or else, here loopline is there, he will definitely give a solution for this.
 
So I've received this answer from scrapebox support -
Hello

Thanks for contacting us.

The error "The operation timed out" is a network error, it's likely due to the proxies timing out or you may have far too many connections running than the proxies can handle. It's going to be difficult scraping 13,800 keywords from Google or any engine with just 10 proxies, so you will need to use a delay. Also Google will typicially only returns 200-300 results per keyword. Please see

Please go to Settings >> Harvester Engines Configuration >> Select the engine that's not working, then click "Test Engine" Do you get search results returned here? Then click the Next button do you also get results for the next page.

If you do then the reason is you may be being blocked, Google is quite restrictive on advanced operator searches like site: and other engines don't support such operators. Or your proxy/VPN may be the source of the errors because the "Test Engine" feature doesn't use proxies.

If not then please go to Settings >> Harvester Engines Configuration >> Import >> Download Default Engines From Server.

This will update the engines file to ensure you have the latest version and all are restored to default. You can also modify the engines to work however you like such as adding completely new engines and modifying existing engines.

Regards,
ScrapeBox Support.

Tried everything they've said I limited the harvester connections to "1" and updated the entire file.

still not working anyone has a clue? I'll attach a photo below of the issue 1710850251042.png
 
I just got scrapebox, I set it up with about 14k Keywords to scrape for websites, Using 10 proxies by MPP (MyPrivateProxy)
when scraping google wont scrape above 200-300.
This is usual behaviour. Even though Google may say they have thousands of results available, they'll only generally show a maximum of 300 for any query. Just try the search yourself and you'll see. To eke out some more results you can try add some keywords to your search queries, for example stop words (like a, and, the etc) or negative keywords (-keyword, -another keyword). Scraping with these methods will give you some different result sets.
 
So I've received this answer from scrapebox support -
Hello

Thanks for contacting us.

The error "The operation timed out" is a network error, it's likely due to the proxies timing out or you may have far too many connections running than the proxies can handle. It's going to be difficult scraping 13,800 keywords from Google or any engine with just 10 proxies, so you will need to use a delay. Also Google will typicially only returns 200-300 results per keyword. Please see

Please go to Settings >> Harvester Engines Configuration >> Select the engine that's not working, then click "Test Engine" Do you get search results returned here? Then click the Next button do you also get results for the next page.

If you do then the reason is you may be being blocked, Google is quite restrictive on advanced operator searches like site: and other engines don't support such operators. Or your proxy/VPN may be the source of the errors because the "Test Engine" feature doesn't use proxies.

If not then please go to Settings >> Harvester Engines Configuration >> Import >> Download Default Engines From Server.

This will update the engines file to ensure you have the latest version and all are restored to default. You can also modify the engines to work however you like such as adding completely new engines and modifying existing engines.

Regards,
ScrapeBox Support.

Tried everything they've said I limited the harvester connections to "1" and updated the entire file.

still not working anyone has a clue? I'll attach a photo below of the issue View attachment 330628

Were you ever able to resolve this issue? I'm running into a similar problem where it's basically stopping after the first search and only harvesting a few URLs.

I've been using the legacy harvester until recently, but I've had to switch from US backconnects to worldwide, and it doesn't seem like adding new custom engines works since I'd need to edit the query for US search results. I tried duplicating the default Google engine, and even the copy won't return any results when the default does, so it seems like there may be a bug with creating new engines in general with the legacy harvester.

I tried using the beta harvester, but I'm running into a similar issue to you; it just stops.
 
Sometimes, search engines update their algorithms or interfaces, which can disrupt ScrapeBox's ability to scrape data.
 
Sometimes, search engines update their algorithms or interfaces, which can disrupt ScrapeBox's ability to scrape data.
That's understandable, but that's why scrapebox moved towards continued development of the beta harvester. I've been corresponding with their support a bit and they believe the current version should work. I'm using residential backconnects and the proxy manager tests show responses within 2-3 seconds and passing google tests, so maybe it's an issue of how the harvester handles these types of connections since the IPs rotate on request?

I ended up recording a short video last night for their support explaining and showing the issue in detail with the engine and harvester configuration. I hope to hear back from them soon, and so far, their support has been very responsive.
 
The issue might be related to your proxies or the connection timeout settings in Scrapebox. Even if your proxies pass the tests, Google can block them quickly if they detect suspicious activity. Try reducing the number of simultaneous connections in the harvester settings to ease the load on your proxies.
Also, increase the timeout value under Settings > Connections and Timeouts. For Bing, ensure the custom harvester engine is up to date, as outdated engines can cause scraping errors. If the issue persists, try testing with a smaller batch of proxies or switching to different ones to rule out proxy quality issues.
 
The issue might be related to your proxies or the connection timeout settings in Scrapebox. Even if your proxies pass the tests, Google can block them quickly if they detect suspicious activity. Try reducing the number of simultaneous connections in the harvester settings to ease the load on your proxies.
Also, increase the timeout value under Settings > Connections and Timeouts. For Bing, ensure the custom harvester engine is up to date, as outdated engines can cause scraping errors. If the issue persists, try testing with a smaller batch of proxies or switching to different ones to rule out proxy quality issues.
Thanks, but those settings are not applicable to the Beta harvester, which their support informed me of. At the moment the beta harvester has no such controls and threads appear to be limited by the number of proxies. Not sure what happened to their support, but I have not heard a word since providing them the video showing and explaining the issues in detail three days ago.
 
shame the scrapebox team abandoned there bhw sales
thread after 14 years, just upped and vanished
That's odd, especially since they have a large customer base here and still seem to be supporting their product to some extent.
 
I thought I'd post an update. I did receive an email from Scrapebox support this evening, and I guess after reviewing my video, they were able to replicate the issue I was having with the current beta harvester. I was told the information was passed along to their dev team, so I'm hopeful we'll see an update pushed within the next week to resolve this issue.
 
Back
Top