[GET] Get Article - The Article Scraper

Looks nice, What format should the proxies be in. I tried :
IP:PORT:USERNAME:PASSWORD

But that seem not to have worked

No private proxies for now, just IP:PORT. I will try to add them in the next release. So for the next release I will add try to include these features:

- private proxies(no socks5 though)
- more sites to scrape from(will use the ones suggested from noermanto)
- limit how many articles to scrape from each site ?
- for now it scrapes only first 100 (ezine,goarticles) and first 15(articlebase) articles. Should I increase this or that number is more than enough ?
 
I could definitely use something like this and I'm sure a lot of other people can too. Very nice share.
 
Thanks for sharing this tool for bhw community.
Might come in handy for plenty of us.

//Jim
 
Thanks ,

I use yourprivateproxy, so it would be great if i could use them with this tool.
Would it be possible to have the titile at the top of each article and not only the file name. REason is i like to spin these before sending them out, having the titile at the top of each article will expedite my spinning as i dont need to copy paste it into the article.

With regards to the number of articles; instead of just scrapping the first 100 or whayever it is, maybe consider a input field where one can specify the number of articles required, and then instead of it going out and pulling the 1st 100 or whatever is specified why not let it randomly fetch the specified number of articles

Great tool

No private proxies for now, just IP:PORT. I will try to add them in the next release. So for the next release I will add try to include these features:

- private proxies(no socks5 though)
- more sites to scrape from(will use the ones suggested from noermanto)
- limit how many articles to scrape from each site ?
- for now it scrapes only first 100 (ezine,goarticles) and first 15(articlebase) articles. Should I increase this or that number is more than enough ?
 
Last edited:
Alright, time for the next build. Features added:

- Private proxies IP:PORT:USERNAME:PASSWORD (not tested though,should work, no socks4/5)
- Added 8 more directories which makes total of 11

Plans for next week:
- Set number of articles to scrape from each site
- Include title on top of text file
- Random article scrape

In case of any bugs let me know :)

Virustotal 0/43
Code:
http://www.virustotal.com/file-scan/report.html?id=9e470a4141873ce0b25f9dbb3782572148836d6a60ac491e4e4d92cc3c835ca3-1294345321

DL
Code:
http://www.mediafire.com/?dnt32zuj91ajbhk
 
Hi,

This is cool, works perfectly on XP 64 and fast scraping from 512Kbit internet connection, thank you.
 
Works brilliantly.

One small request - any way it could retain the paragraph formatting rather than outputting single slabs of text?
 
Works brilliantly.

One small request - any way it could retain the paragraph formatting rather than outputting single slabs of text?

That might work if the article contents are saved in html rather than txt. Gonna test it now

edit: indeed saving in html, keeps all the formatting including aligning, quotes, images, etc. exactly like seen from the browser, but just the article frame
 
Last edited:
Would proxies scraped with scrapebox work well with this or should I stick to using my private proxies?
 
Would proxies scraped with scrapebox work well with this or should I stick to using my private proxies?

Scrapebox proxies are mainly for scraping SE and blocked by SE, so they should work fine here. I personally use about 2 public proxies per keyword, so if I have 10 keywords that's 20 proxies and according to the http sniffer, almost all articles are saved successfully without the article site complaining about too traffic coming from me
 
Hi

I am interested in working with you to further promote this, have send you a PM
 
Whoooa. Just what I was looking for, even though I don't know which side I am on. Quality over quantity or vice versa.

Thanks for this!
 
you are awesome, this is likely the best program of it's kind. thanks!
 
Nice software. You could really turn this into something special by building more features into it.
 
Gave +1 REP to you. Awesome tool. Please add more directories or put option to load a directory list please....:D
 
That might work if the article contents are saved in html rather than txt. Gonna test it now

edit: indeed saving in html, keeps all the formatting including aligning, quotes, images, etc. exactly like seen from the browser, but just the article frame

So ihow do you save in html?

Hope it works on ezine. :)
 
This is really VERY useful already...thanks for sharing it with us.

In addition to the planned changes (title on first line of output, ability to specify # of articles to scrape per keyword), is it possible you can also add an option to output all articles to a single directory instead of the default of creating subdirectories? That would be really sweet.... :)
 
Back
Top