Im using the vanity name checker, and after checking for some domains, im getting a "socket error" message.What seems to be wrong here?
A socket error occurs generally when something forcibly closes the connection, it could be a firewall, bad proxies etc.
Firstly try it with and without proxies. Addons only pull proxy data from scrapebox on startup, so if the use proxies box is checked or unchecked on startup and any proxies loaded, will only be refreshed in the addon if the addon is restarted. So try unchecking the use proxies box and restart the addon, does it still happen?
If so then it's probably something firewall or antivirus related.
Its basically process of elimination. If it happens without proxies its probably local, if it happens only with proxies, then its probably a proxy issue.
src:
http://scrapeboxfaq.com/what-does-socket-error-mean-in-an-addon
whats the Problem with google Harvesting? after harvest 3k-4k it will not harvest any url.How can i solve this issue,I have 20 us Private proxy and all are fast.Is there any wrong with my Scrapebox setting?
I have set 3 connection for google harvest and i have 30 sec timeout,I was change this setting also,But google don't harvest no more than 3000-4000 url,Why?? whats the correct setting for 20 private proxy?
Please let me know As Soon As possible.
View attachment 49578
Try 1 connection, 3 is too many and your ips are being banned.
one more thing, how can i scrape for all the pages a website have?Internal link extractor?
Yes, you could do that, or ideally
do a site:domain.com
to get started, that will pull a bunch of the internal pages, then you can do the link extractor on internal, and them keep loading the results back in and pulling internal links. Also you can use the sitemap scraper addon if there is a sitemap.
I m stuck harvesting from Bing. Both Google and Yahoo completed but Bing found around 20k plus results and its still showing that its harvesting but nothing moves. Bandwidth, URLs are all stopped. If i press stop harvesting it still doesnt stop and i have to end the process in order to stop it.
This has happened all the times in the recent weeks i ignored it as the harvested URLs are saved. But its getting annoying.
Any help?
Also some of the harvested URLs have this added at the end /RK=0 which gives a page not found error.
Its a known issue that bing locks up sometimes on some setups in the multi threaded harvester. If you want to use bing then use the custom harvester, there is no issue there. Also the custom harvester has multiple advantages, such as 20+ engines it can scrape from and in your other case the ability to download a new engines file. So scrapebox can update the file and you can download it when an engine changes something and it causes something like your /RK=0 to append to urls. Else with the multi harvester its hard coded and you have to wait for a update of scrapebox. Video I did on that: