ANybody do web scraping at large scale?

Joined
Mar 23, 2022
Messages
15
Reaction score
6
I'm just wondering how you people do it since the same IP address tends to get blocked quite a lot.
Do you do IP rotation, slow down the process, or use residential proxies? What has worked for you guyss?
 
Most of our users rely on rotating residential proxies to avoid blocks on stricter sites. For platforms with lower security, they usually opt for datacenter proxies to maximize speed and cost efficiency. Slowing down request rates helps, but using rotating IPs makes the real difference.
 
I know a few people who order around 50–100 mobile proxies for these purposes. They’ve told me that it has become noticeably harder recently, though, so even with mobile proxies it’s not as straightforward as it used to be.
 
what website you want to do scraping? Maybe i can help and join you use as in experiment if it's appropriate for me.

I had done a lot of it in the past time. Millions of pages monthly at scale.
 
Yeah rotating IPs has worked better for me especially instead of pushing everything through one IP.
 
usually the more unique IPs per request you can do, the better, you spread server-side rate-limits that way. You can also start thinking of TLS fingerprinting (e.g. JA3) to avoid getting flagged by some more enhanced WAFs/CDNs, there are plenty of open source http clients for that nowdays.
 
Basically mobile proxies + constant rotation. Scraping at large can't happen without clean IPs.
 
I found out that slower requests + decent IP rotation is better as opposed to just rotating as fast as possible
 
LUXPROXY been pretty solid for some scraping work I’ve done, especially the residential sidee
 
If you’re testing different providers, luxhost lets you test the proxies first, which is handy before spending money.
 
Once scraping gets bigger, keeping the same IP for too long usually becomes a headache.
 
When going mass scale scraping I tend to use cheap ISP proxies, and rotate them while flipping multiple threads. All of this does somewhat depend on what kind of limitation the framework has set, some platforms are easier than others.
 
For data scraping I tend to go with the cheapest option available, usually you can get pretty far using cheap-o isp ips.
 
We have many users that use our scraping pools for exactly this. You don't need to slow your requests down if you have enough IP's in the pool
 
Back
Top