matessim
Junior Member
- Nov 22, 2008
- 164
- 74
Hi Guys,
I've written a scraper that can scrape Google fairly well and fairly fast and without getting blocked before reaching thousands/tens of thousands of requests in a rather short amount of time.
I'll be glad to release this to you guys and build it to do what anyone here is interested in it doing (it can scrape the URL's, the description, and the title at the moment, with possibly more features such as cached text of the page and preview icons for videos if anyone is interested).
I have reached the problem that i simply can't test it with multiple threads with multiple proxies since i can't seem to find any that work/haven't been blocked by google since they're hammering scrapebox/xrumer while i'm hoping to use them. If anyone can help with highly anonymous socks5 (https should be fine too) proxies, i'll be glad to continue developing this.
Here's a small taster:
Note: This is not via Google API or anything, It's actually a headless browser (with JS Evaluation, Google won't load well otherwise) that's visiting and parsing the page.
I've written a scraper that can scrape Google fairly well and fairly fast and without getting blocked before reaching thousands/tens of thousands of requests in a rather short amount of time.
I'll be glad to release this to you guys and build it to do what anyone here is interested in it doing (it can scrape the URL's, the description, and the title at the moment, with possibly more features such as cached text of the page and preview icons for videos if anyone is interested).
I have reached the problem that i simply can't test it with multiple threads with multiple proxies since i can't seem to find any that work/haven't been blocked by google since they're hammering scrapebox/xrumer while i'm hoping to use them. If anyone can help with highly anonymous socks5 (https should be fine too) proxies, i'll be glad to continue developing this.
Here's a small taster:
Note: This is not via Google API or anything, It's actually a headless browser (with JS Evaluation, Google won't load well otherwise) that's visiting and parsing the page.
Attachments
Last edited: