jimBlackHat
Newbie
- Jan 14, 2015
- 3
- 1
Hi,
right now I am developing a web scraping service where people can scrape pages like yellow pages, ebay, amazon and all kind of shops for data they need. Everything runs on my server.
A bit more detailed: People register on my site an pay for the service. They now can scrape data from pages and download this data as excel, csv, xml, etc. I simply turn unscructured data into lists.
I have talked to a lawyer and he said that it would be a potential risk to start such a service in germany (this is where I live) because of "competition regulations" and "copy right".
Now I am looking for possibilities to avoid getting sued by shop owners that my potential customers will scrape.
Some ideas:
- What about setting up a limited (LTD) company?
or
- Use a web hoster thats not located in germany for registerting the domain name and running the server?
There are already several websites that are providing such kind of scraping service. I wonder how they manage this problem.
Any ideas are welcome!
EDIT: Another solution. Always use open proxy server and fake the user agent. Good or bad idea?
Thanks,
jim
right now I am developing a web scraping service where people can scrape pages like yellow pages, ebay, amazon and all kind of shops for data they need. Everything runs on my server.
A bit more detailed: People register on my site an pay for the service. They now can scrape data from pages and download this data as excel, csv, xml, etc. I simply turn unscructured data into lists.
I have talked to a lawyer and he said that it would be a potential risk to start such a service in germany (this is where I live) because of "competition regulations" and "copy right".
Now I am looking for possibilities to avoid getting sued by shop owners that my potential customers will scrape.
Some ideas:
- What about setting up a limited (LTD) company?
or
- Use a web hoster thats not located in germany for registerting the domain name and running the server?
There are already several websites that are providing such kind of scraping service. I wonder how they manage this problem.
Any ideas are welcome!
EDIT: Another solution. Always use open proxy server and fake the user agent. Good or bad idea?
Thanks,
jim
Last edited: