Scraping Websites

agag2

Supreme Member
Joined
Feb 17, 2009
Messages
1,303
Reaction score
260
Hi

If i wanted to scrape news websites and extract ALL content from the site - how would I go about retrieving the urls for all content on site in absence of an xml file or sitemap?

Thanks
 
Like all scrapers do: by following links...
You can use: httrack. 1 minute install, 1 minute to understand how it works (at least for the basics), and then you can download the website :)

Have Fun.
 
The one and only beast SB(scrapebox) will help you to find the desired sites.
 
GSA SER has an option "crawl online" which crawls as many levels as you wish.

Regards
 
Back
Top