If i wanted to scrape news websites and extract ALL content from the site - how would I go about retrieving the urls for all content on site in absence of an xml file or sitemap?
Like all scrapers do: by following links...
You can use: httrack. 1 minute install, 1 minute to understand how it works (at least for the basics), and then you can download the website
This site uses cookies to help personalise content, tailor your experience and to keep you logged in if you register.
By continuing to use this site, you are consenting to our use of cookies.