davidlagunas
Registered Member
- Aug 7, 2013
- 81
- 15
Well it sounds like your proxies are blocked.
Are you using the custom or detailed harvester?
Custom
Well it sounds like your proxies are blocked.
Are you using the custom or detailed harvester?
I have doubt hope you can let me know whether it's possible do it through scrapebox
I have a list of domains with extensions and is it possible to check whether it's possible to register them ?
Custom
do you have a discount Op?
Hello,
Yesterday I've purchased ScrapeBox because it has a feature to scrape from the "Dogpile" search engine.
Now that I got Scrapebox I went to the "costume harvester configuration" window to check the Dogpile engine out, but when I click "Test engine" it says "No links could be retrieved"....
I've tried making it work by changing the engine's settings but with no success...
Please help me out with this, I need the dogpile engine...
Cheers
I've just updated the Dogpile engine, so if you restart ScrapeBox it will automatically download the latest engine definitions and it should work.
The proxy retries was on maxSo you can try going to settings >> connections timeouts and other settings >> more harvester opttions - and turn the proxy retries all the way up. Because whats happening is many of your rotating proxies are no doubt blocked and when the retries are reached the keyword is skipped.
You could alternatilvey try the detailed harvester, which is built for accuracy. It will take longer but likely yield more results. But a lot of your proxies are blocked which is why its happening.
The proxy retries was on max
thanks for the solution much appreciated I got one more doubt
is there a way to compare two domain list and save the only difference ?
Hey, I've got a question about proxy usage...
is there a feature to wait between search engine queries for each proxy?
for example proxy1 has to wait 60 seconds between each search it does.
This is really impotent to me so my proxies wont get banned.
Cheers
Yes, you can use the detailed harvester and use the delay option. That is a delay per keyword not per proxy. You can go into the settings >> harvester engine configuration - and then set a delay for each engine. This delay is per query.
Both those work in the detailed harvester and they do combine, so if you set a delay per query and a delay per keyword you will wind up with both stacking.
You can also go to settings >> connections timeouts and other settings >> more harvester options >> proxy change interval - and set this. This is how often you force a proxy change. So you could force a proxy change after every request if you want. Mind you if you use the custom harvester the delay option isn't there, but forcing when to chnage a proxy may achieve what you want in the end as well.
I read several times still a little confused.
This will lower the speed of scraping significantly unnecessary.
The proxy change interval is well... I can't call it reliable because I don't really know how many queries the proxy does in a set amount of time.
so really there's isn't such feature for fast scraping because the details harvester only uses 1 thread
Making a feature that creates search engine queries timers for each proxy could really save them from getting banned, no matter how many threads you use.
For example:
you have 10 fast private proxies and you really don't want them to get banned, you could make a safe bet and set that each proxy would search every 60 seconds.
That means 10 search engine queries every 60 seconds.
and if you're willing to risk it to have faster scraping you could lower the query delay for each proxy.
and of course you can get more proxies to make it scrape faster with the proxy delay.
This is just a safe bet feature that the proxies won't get banned, resulting in faster scraping in the long term.
could you guys impalement this feature?
I think this feature would greatly benefit scrapebox and its users.