Ultimate scraper be like ?

GrandScraperG

BANNED
Joined
Jan 12, 2019
Messages
52
Reaction score
5
Hi Guys!
So, hope it's the right forum.
I would like to hear the crazy ideas for data scraper service, that is not exist these days.
I know services like scrape box and scrapinghub and i can assure my quality is much better.

Everything will be handled on my side including captchas and proxies.


With description on how you think it should work.
The one that have the most "likes" i will go with it.

I am an elite security researcher, specialist on reverse eng and bots.
Think about tools that you know are "impossible/difficult" to make.

That's going to be my contribute to this forum.
 
You sir, are a very humble person.
I stand behind my words.

I am going to make the best scraper in the world.
Let me know guys!

Scalable system with 10m+ requests daily for any site/mobile app
 
lol you think it's impressive ? it's not rocket sience to make a decent bot with a stealth system for web site with auth & captcha.
There are MANY protected apps and sites, that don't allow you to scrape anything.
That's my focus, breaking security
 
Scraping phone numbers & emails from instagram accs (backend one - NOT public)
 
Scraping phone numbers & emails from instagram accs (backend one - NOT public)
It can only be done, if while making a requests to their service backend via pure requests the response contains this data.
Otherwise, you need to hack their DB which is not the case we dealing with.
 
How exactly would your quality be better than existing solutions?

10M is really not impressive... do you think that this is a lot for existing services, e.g. for ScrapingHub?

Are you aware that you need a fully distributed system with linear scalability (e.g. the scrape job has 1B URLs, which of these are already crawled, like / of a site) (or how to store the results - you can't keep them obviously in memory)

Which technologies do you want to use?

Good luck!
 
How exactly would your quality be better than existing solutions?

10M is really not impressive... do you think that this is a lot for existing services, e.g. for ScrapingHub?

Are you aware that you need a fully distributed system with linear scalability (e.g. the scrape job has 1B URLs, which of these are already crawled, like / of a site) (or how to store the results - you can't keep them obviously in memory)

Which technologies do you want to use?

Good luck!
10M it's just a start, i can scale it as much as i want. Current projects not require more than that.
I have some tech in python and some in golang. it depends on the level of research of protection.
 
Can you reverse engineer recaptcha to get the g-recaptcha-key from javascript?
 
Can you reverse engineer recaptcha to get the g-recaptcha-key from javascript?
It's not profitable project, google change very often and solving captcha via 3th party solution is soooooo cheap. :)
 
Create Google places scraper that will use just Http requests, I'm wondering how will you handle javascript, will you post your code too?
 
Create Google places scraper that will use just Http requests, I'm wondering how will you handle javascript, will you post your code too?
definitely not sharing the code.
Can you explain more, which places should scrape and who will be interested in the solution ?
 
Back
Top