SerpAPI launches it's own scrapable non-Google index

Steptoe

Elite Member
Jr. Executive VIP
Jr. VIP
Joined
Aug 9, 2017
Messages
5,667
Reaction score
6,970
Looks like SerpAPI have been putting all the data it's been scraping to use as a secondary purpose - they've built their own web index of apparently billions of pages. Preview is available here: https://serpapi.com/search-index-api

I couldn't find any info on what crawler they're using to source their own data (if indeed they are - they could just be relying on scraped data), and this is the only official post I could see from SerpAPI mentioning it: https://serpapi.com/blog/reddit-promised-its-users-an-open-internet-then-it-closed-the-door/

Still, always good to have another data source!
 
An additional independent index would also be helpful particularly for research and checking pages outside of Google i am interested to see how big and current the index will become relative to the existing search APIs.
 
Looks like SerpAPI have been putting all the data it's been scraping to use as a secondary purpose - they've built their own web index of apparently billions of pages. Preview is available here: https://serpapi.com/search-index-api

I couldn't find any info on what crawler they're using to source their own data (if indeed they are - they could just be relying on scraped data), and this is the only official post I could see from SerpAPI mentioning it: https://serpapi.com/blog/reddit-promised-its-users-an-open-internet-then-it-closed-the-door/

Still, always good to have another data source!
These kinds of data are as useful as they are fresh. Any idea how fresh or recent they are?
 
These kinds of data are as useful as they are fresh. Any idea how fresh or recent they are?
Looks like they haven’t documented that part yet. The Search Index API is still marked as preview, and I can’t see any crawl timestamp or recrawl schedule for the underlying index in their docs

They do have a no_cache=true option, but that only bypasses SerpAPI’s search-result cache, which expires after an hour. It doesn’t mean the pages in the index were freshly crawled

So for now I wouldn’t assume it’s real-time. A crawl date per result or a published recrawl policy would make the freshness much easier to judge
 
Back
Top