I designed my own scraper with a PHP recursive http_get to access 60,000 records from a non-profit search engine database in 2004 (the same year, I built a merger just like in Scrapebox for Google AdWords, Google got rid of spreadsheets the next day). Google claims to index 500 pages every 3 days for new websites, each new domain, with 0 Pagerank, then 500/day with Pagerank>=1. This is powerful as the 99.99% of named individuals, their professions and/or their businesses, who are not well known, look up or others look up that information, corresponding to a seed keyword or "near me" connection probably several times a month. With a new domain/website a day, that's at least 3 million impressions a month, and probably hundreds of thousands of visits that accumulate over time.
There is a need for recaptchas and captcha libraries to index Web pages with OCR robots that will post text to visible text fields, and hit "Post" or synonyms for similar words, to create a MLM sale's pitch, to Facebook, PBN's, Symposium, etc.