Hi !
I reup this thread because I have stop my project from years, but I have restart it !
I have implemented all this one:
- Extract a very large kind of things in any webpage: links, images, sounds, emails, proxies, PDF, Words, Paragraph, Discogs links.
- Extract tweets.
- Use custom Regexes...