So say I harvest 1000 results. I click on "remove duplicate domains" and sometimes as many as 400 out of the 1000 results are removed. So I never could actually achieve 1000 unique domains in a single search - if I am using it right?
So my suggestion is that as soon as the harvester finds one URL from a domain, it ignores subesquent ones, if selected.
Correct but the 1,000 result limitation per query is a limitation with Google, Yahoo & Bing not with ScrapeBox.
So it makes no difference if ScrapeBox removes duplicate domains on the fly, or you do it by pushing the "Remove Duplicate Domains" button you will always get 1,000 minus however many duplicate domains there are if all you want is unique domains. Example..
Search: test
You can harvest 1,000 URL's right, and test.com is first so in order to exclude that domain from being harvested in subsequent pages and get a full 1,000 unique domains we would have to again query
Search: test -test.com
This now performs the query again eliminating that domain so we have a full 1,000 unique domain sample to harvest from for the keyword. If there is 400 duplicate domains, this means another query with 400 -domain.com tacked on to the keyword to filter all your found ones.
A lot of work and calculations and it's not real practical. You know a certain portion will be removed if you elect to filter the list and only have unique domains.. So harvest more than you need to compensate for this, that's the best way.
Make sense now, Google allows you to harvest 1,000 URL's total so if you want to filter this down you have to expect the number to reduce.
Thanks - all working now!!
Thanks mate, and no problems.
Edit:
v1.1.10 Uploaded
- Pause Button added for all posting/harvesting operations
- Export URL's & Pagerank to .xls Excel Format
- Some rewriting of Proxy Checker
- Ping Mode increased to average of 20 pings per second
Please be careful with Ping Mode, you could possibly bring down shared hosting with it (DOS) so use a delay.