Site not scrapable, is it possible?

gm90

Junior Member
Joined
Feb 15, 2018
Messages
160
Reaction score
45
Hi. I want to scrape a website but sadly, they added a cloudflare captcha worldwide, even for legit users.
To bypass the cloudflare captcha I found the origin IP of the server, changed my file host so I can bypass the Cloudflare. But here, after 12 scraped page, I get the error "too many requests" from another plguin. Of course a solution for this would be to use proxy, but proxy ignore my file host so they try to open the website through cloudflare and I don't have privilege to change their file hosts to point the request to the origin IP. So Is this site not scrapable for real?
 
Adding a wait time wouldn't work?
At the moment my delay was 15 seconds and was blocked. I have not tried to increase this value yet but the site is big with thousands of page... increase this further would require months to scrape, looking for better alternatives
 
Back
Top