9to5destroyer
Regular Member
- Nov 14, 2011
- 360
- 208
i'm in the process of a massive scraping mission i'm scraping all urls of certain sites then extracting data using custom tools.
i'm having problems with certain sites mainly sites with subdomains eg health.site.com
this is only on certain sites which i dont really understand.
1.if i scrape site:site.com keyword it returns next to no results but if i google site:site.com it returns over million indexed subdomains.
2. on other sites with subdomains following the same process using around 100k keywords(very varried) even if i scrape 3-4 million urls when i remove duplicates i only have about 30000 urls.
like i said the process is working perfectly on most sites but i'm having trouble on only a few any ideas?
i'm having problems with certain sites mainly sites with subdomains eg health.site.com
this is only on certain sites which i dont really understand.
1.if i scrape site:site.com keyword it returns next to no results but if i google site:site.com it returns over million indexed subdomains.
2. on other sites with subdomains following the same process using around 100k keywords(very varried) even if i scrape 3-4 million urls when i remove duplicates i only have about 30000 urls.
like i said the process is working perfectly on most sites but i'm having trouble on only a few any ideas?