rupptscheck
Newbie
- Jan 15, 2009
- 5
- 0
Hello there
I have a question about harvesting URL's whenever i do harvest url's when it finished with the task let's say i have in total 30k harvested url's then what i do is to check for duplicates url's and duplicates domains....
The problem is that after doing this i have almost everything duplicated and i end it up with only 500 url's or less from 30k ? is that normal ?
Why so many dupes, this has something to do with footprints or something like that ? anyone has the same problem ?
Thanks...![]()
very normal! cause you are scraping from 4 different search engines. With the domain one only use it if you are promoting less then 2 or 3 sites. I say that cause if you have 3 domain urals and your promoting 3 sites, it will post each on that website but not the same page since you took out dup pages.
i stick with taking out all dups pesonally..
Hi there was no return button or i just quit the page. What can i do now?
Hey man Thanks for your answer !
The thing is that i'm only harvesting url's from Google not using the others because i didn't wanted any dupes, but after harvesting google i got this mass duplicates problem over and over again .....
Din't understood what you're saying about the domains ? i'm just trying to promote one page ... i though it was because the keywords and url's i'm attemting to get are not in English language but i did the same in English and i get the same results after 50k url's harvested i remove dupes and i get only a few hundred's why ? :bawling:
Hello there
I have a question about harvesting URL's whenever i do harvest url's when it finished with the task let's say i have in total 30k harvested url's then what i do is to check for duplicates url's and duplicates domains....
The problem is that after doing this i have almost everything duplicated and i end it up with only 500 url's or less from 30k ? is that normal ?
Why so many dupes, this has something to do with footprints or something like that ? anyone has the same problem ?
Thanks...![]()
think of it this way...... if you dont take out duplicate domains and the scraper picked up 30 links off it from the same site, you will bomb the site. A webmaster might not even notice a comment 1 time, but on 10 pages it might be enough to piss them off. So its best to get rid of all dups in my eyes unless you have more then 1 site in your list to promote.
I have also experienced this, one time, I scraped around 50k blog list then after filtering the duplicate domains, only about 2k blogs are left.
I suggested a feature before that the filter should be done while scraping is in action but I think it would be too difficult for sweetfunny to do this.
What I'm doing right now is I list as many keywords as possible and I select all keyword sources. I limit results to 100 then choose google, yahoo, bing and aol as search engines. With this settings, I can scrape around 100k URLs then after the filter, about 20k - 30k will be left. There will be a lot of duplicates but hey, at least the results will still be a lot.
So my suggestion is to go by the numbers![]()
Just want to say your article is striking. The clarity in your post is simply striking and i can take for granted you are an expert on this subject. Well with your permission allow me to grab your rss feed to keep up to date with forthcoming post. Thanks a million and please keep up the ac complished work. Excuse my poor English. English is not my mother tongue.
Just have to say scrapebox is awesome! Support is really awesome too and there's updates almost everyday. There's a million uses for it if you're creative and it has literally cut down days of work into just minutes!
Although those of you who are just copying the comment examples in this thread really suck and are going to ruin it for the rest of us if this gets out of hand. Like the moron who keeps spamming me with comments like this:
I am really struggling to find proxies to use (paid). Can anybody point me in the right direction of some decent paid proxies that allow users to use Scrapebox?