Wayback machine pros: how to crawl at large scale?

cooooookies

Senior Member
Joined
Oct 6, 2008
Messages
1,133
Reaction score
283
Hi,

What I would like to do: get the last active root page of about 10M expired domains. Is this possible? I am not necessarily asking for software, I would do the crawling and processing myself. I am asking for the correct way?

1) Are there any rate limits imposed on wayback machine?

2) Is there any programmatic way - via API? I tried the wayback machine availability API, but it is always returning empty for the snapshots archived.

3) Often, the domains were parked in the past and are expired now. In that case the last snapshot is bull. How would you get the last "real" snapshot?

Thanks and cheers
cooooookies
 
Back
Top