google meta data scraper

*Heracles*

Registered Member
Joined
Feb 3, 2020
Messages
64
Reaction score
12
Hi Folks, I am using the Scrapebox google metadata scraper to source emails from LinkedIn. I am having difficulty with finding the right proxy set up. Can anyone share the proxy set up they use to hide from the google overlords? Cheers
 
Sure, a couple of videos


and

Thanks Loopline, your a legend! Is the advice in these vids still accurate? Wasnt sure if google changes had changed your approach over the last few years.

Thanks for all your time advice over the past few months!
 
Thanks Loopline, your a legend! Is the advice in these vids still accurate? Wasnt sure if google changes had changed your approach over the last few years.

Thanks for all your time advice over the past few months!
Advice is 100% accurate from a concept standpoint.

That means that for private proxies its still about the ratio, which essentially equates to speed.

For rotating proxies the setup in scrapebox is still the same, although providers come and go, so I keep all the links in the video and video description updated as providers change.


What that means is for example in Jan of 2020 google made some updates. When they made those updates they significantly tightened down on blocking ips, especially for more advanced operators. Google makes changes from time to time but this one was big specifically on how fast they ban ips.

As a result if your working with a pool of ips then if everyone gets them all banned faster then no one gets any scraping done, and with private proxies it means the ratio goes up and/or delay. So you might need 100+ private/shared private proxies now with 1 connection to scrape google, and if you don't have enough you just need to use the detailed harvester with a delay.

At any rate, the concepts are the same as the videos, just the ratios and delays have changed since I recorded the videos. So basically if in doubt do 1 of 2 things

1 - private proxies - start at crazy high delays. Think of what you think is crazy high and then double it and start there. Then if it works, go down and keep going down till it stops working and you get blocked. Then you know your sweet spot. They un ban ips in 48 hours or less so no stress on getting the ips blocked.

2 - use rotating proxies for scraping. The downside here is that typically rotating proxies are less suited for posting and pretty much anything short of scraping.


so if your on a tight budget go for shared private proxies and tell the provider you want to scrape google so they can match you with people that do not want to scrape google. Then go uber slow and then you can use the same proxies to post.

If you have a small budget or a decent sized one then grab a small or medium rotating proxies package for scraping and then some shared private proxies/private proxies for posting. (assuming your doing posting or you need proxies outside of scraping).

Glad the videos and information is a help! :)
 
Advice is 100% accurate from a concept standpoint.

That means that for private proxies its still about the ratio, which essentially equates to speed.

For rotating proxies the setup in scrapebox is still the same, although providers come and go, so I keep all the links in the video and video description updated as providers change.


What that means is for example in Jan of 2020 google made some updates. When they made those updates they significantly tightened down on blocking ips, especially for more advanced operators. Google makes changes from time to time but this one was big specifically on how fast they ban ips.

As a result if your working with a pool of ips then if everyone gets them all banned faster then no one gets any scraping done, and with private proxies it means the ratio goes up and/or delay. So you might need 100+ private/shared private proxies now with 1 connection to scrape google, and if you don't have enough you just need to use the detailed harvester with a delay.

At any rate, the concepts are the same as the videos, just the ratios and delays have changed since I recorded the videos. So basically if in doubt do 1 of 2 things

1 - private proxies - start at crazy high delays. Think of what you think is crazy high and then double it and start there. Then if it works, go down and keep going down till it stops working and you get blocked. Then you know your sweet spot. They un ban ips in 48 hours or less so no stress on getting the ips blocked.

2 - use rotating proxies for scraping. The downside here is that typically rotating proxies are less suited for posting and pretty much anything short of scraping.


so if your on a tight budget go for shared private proxies and tell the provider you want to scrape google so they can match you with people that do not want to scrape google. Then go uber slow and then you can use the same proxies to post.

If you have a small budget or a decent sized one then grab a small or medium rotating proxies package for scraping and then some shared private proxies/private proxies for posting. (assuming your doing posting or you need proxies outside of scraping).

Glad the videos and information is a help! :)
Thanks again Loopline! Your advice is always appreciated.

Randomly over the weekend, I managed to get my settings to work with no proxy. It's slow but better than the proxies I was using so all good : )

Other than linkedin are there other sites you google scrape to find emails for people who do a specific job or business i.e Health & Safety folks / businesses.
 
Thanks again Loopline! Your advice is always appreciated.

Randomly over the weekend, I managed to get my settings to work with no proxy. It's slow but better than the proxies I was using so all good : )

Other than linkedin are there other sites you google scrape to find emails for people who do a specific job or business i.e Health & Safety folks / businesses.
I don't know of specific sites, but you could just spend a little time on google looking for directories of people for the niches you want, they probably exist.

glad you got scrapebox to work!
 
Hi Folks, I am using the Scrapebox google metadata scraper to source emails from LinkedIn. I am having difficulty with finding the right proxy set up. Can anyone share the proxy set up they use to hide from the google overlords? Cheers
what do you use this content for?
 
Back
Top