[GET] Free content for any niche. (get it before it's gone)

Does anyone know the name of the Firefox addon to scrape and download the content? Also, is there a way to do this using Scrapebox?
 
@Brix
fatboy made a bot for free and if it now worked then we can't say any thing against him.

move on.
 
Does anyone know the name of the Firefox addon to scrape and download the content? Also, is there a way to do this using Scrapebox?

Asked Loopline regarding this. He mentioned that it is not capable of doing so. You're better off with what other members suggested above (proxyswitcher (not sure if this is the one) + download the mall addon on firefox.). Or download free trial of Seo Content Machine. DL the lists given above and run them on SCM's URL content scraper. A proxy would last around 500 articles so you'll need a lot if you plan on scraping a lot.
 
And here I was thinking I heard the last of this shitty thread.

LOL. Before yesterday, I had NO IDEA who "fatboy" was. If he has a good reputation, awesome for him. If he's contributed to this forum for "years", even better. I've done the same shit on WF. A LOT of people contribute to forums dude. And personally, I can't stand when members are singled out for all the good they do.

No one is singling me out, kvmcable mentioned me as they thought you were throwing your toys out the pram about me. If you were trying to slate someone else then I have no doubt whatsoever kvm would of done the same for them.

I want it to be clear, I do not respect someone just because they've been here for years or have a high post count. Because there are people on here with lots of posts and thanks who use their reputations to troll and abuse newbies. So I respect people based on my own perceived value of what they share. And I might have high standards but that's fine with me.

If you think about it everyone is owed respect UNTIL they make it obvious they do not deserve it, in my opinion living by that rule will be a lot better, however you appear to go from the other direction that everyone means nothing to you until they live up to your standards.


Finally, if fatboy really had the "skills" you speak of, shouldn't he have known there were addons that already do what his bot does?

Do you know every single addon in all the browsers? If you read back I pointed out that all you really need is a bit of bash script to do it all. Not everyone likes running scripts, not everyone likes browser plugins, not everyone likes running exe files - horses for courses.

Isn't that the type of shit that "skilled" programmers are suppose to know? Don't you find it ironic that the person who ultimately found the best method wasn't even a "programmer" at all? Cause as highly as you seem to think of him, even AFTER I apologized, you don't seem to think highly enough that he can handle his own criticism. To the point where you have to have his arguements for him! LOL.

Can you point out anywhere where I say that I am a skilled programmer? I use Ubot FFS, thats drag and drop. Do I have a programmers mindset, yes, do I have the skills to use multiple languages - depends how you look at it. I put stuff together using perl, php, ubot and little bits of .net as I try to move away from Ubot as the company behind it is dragging it to the bottom if a murky sea. If you want a skilled coder, you wouldn't want to pay pennies for a bit I write, you want skilled look at the likes of TomPots or DarkPixel, they are in my eyes the skilled programmers.......I am a hobby hacker, I have a 9-5 (no matter how shit that is) to pay the bills, coding, scripting, hosting and sysadmin is my hobby for beer money. Never made a secret about that.

I'm trying to keep fatboy out of this now but if you want more than an apology you're out of your mind.

Then lets do it, the thread has served its purpose, bitching all day isn't going to pay your bills.......
 
@Brix
fatboy made a bot for free and if it now worked then we can't say any thing against him.

move on.

Speak for yourself.

If the bot really "worked" like a real scraper you wouldn't see members still making posts asking how to mass download the articles. And I wouldn't have gotten 6 pm's in the last 12 hours from people saying "I used fatboys bot but it won't do what I need, can you please share your method?" So it works for small, limited scrapes, but not for mass downloading.

Some of you guys are really close though. Just use download them all by firefox, a proxy switcher, proxies, and if you need anything more than 25k articles you will need some macros to automate the process. Just be sure to write "cache" in the filter part of the DTA so it only downloads cached versions. Right click, download, wait till the addon stops working (I was averaging 300 articles / proxy), switch proxies, repeat over and over till you got what you need. It's very simple and works better than that bot (imo). And again, I appreciate him taking a whole 20 minutes to build the bot but some people here wasted DAYS on end just figuring out this stupid simple method.

Then lets do it, the thread has served its purpose, bitching all day isn't going to pay your bills.......

I agree. I'm done.


-BB
 
Last edited:
Speak for yourself.

If the bot really "worked" like a real scraper you wouldn't see members still making posts asking how to mass download the articles. And I wouldn't have gotten 6 pm's in the last 12 hours from people saying "I used fatboys bot but it won't do what I need, can you please share your method?" So it works for small, limited scrapes, but not for mass downloading.

Okay - last reply before I unsubscribe from this thread.......

The bot worked fine when Yahoo Voices was still up and going, I scraped a load and others did to. Yes, now you have to scrape cached links its as much use as tits to a bullfrog. I can put proxy switching in it quite easily, but to be honest your way of doing it will probably be just fine.

At a guess you could probably use a hundred or so proxies, select a random one, download 50 articles (Google seems to captcha at 70 - 75), switch to another proxy, repeat.....by the time you get to the end of your proxy list Google would of probably chilled its beans enough to use the first one again.

Loop that bitch up and you will be happy.

Like I said, thats my last reply on the thread / bitchfest.

Enjoy the weekend!
 
Okay - last reply before I unsubscribe from this thread.......

The bot worked fine when Yahoo Voices was still up and going, I scraped a load and others did to. Yes, now you have to scrape cached links its as much use as tits to a bullfrog. I can put proxy switching in it quite easily, but to be honest your way of doing it will probably be just fine.

At a guess you could probably use a hundred or so proxies, select a random one, download 50 articles (Google seems to captcha at 70 - 75), switch to another proxy, repeat.....by the time you get to the end of your proxy list Google would of probably chilled its beans enough to use the first one again.

Loop that bitch up and you will be happy.

Like I said, thats my last reply on the thread / bitchfest.

Enjoy the weekend!

Ah, I did not realize you built the bot when voices was still up. For some reason I thought the database went down July 1st but I had to check and realized it was Aug 1st. Makes perfect sense now.

Anyway, enjoy your weekend too. And congrats on winning Jr VIP. :)
 
I have posted a article on my blog but it didnt get index, i tried everything. That article is already deindex from Google 3 days ago. Any solution ?
 
I have posted a article on my blog but it didnt get index, i tried everything. That article is already deindex from Google 3 days ago. Any solution ?

post it on a blogger/wp blog.
-=-
 
I've been trying to do this, It only redirects me to yahoo's home page?

Anybody else finding this?

Don't click on title link ,you will see drop menu in right corner,click on it than chose cahed.If it redirect you than it is good because it mean it is down,point is to get downed article which you can find only cahed.
 
any idea when yahoo voices articles will deindex? I ran some through smallseotools plagiarism checker and articles still exist in google

Some of them are already deindexed. On my niche, 2 of the 100's articles that I (manually :( ) got are posted on my site!

Check regularly!
 
Well, I got to this thread a couple of days late, but it is a interesting read.

The basic technique is a great find. I am not planning on using the raw articles because who knows who else will be using the raw articles and therefore you would have duplicate content.

This technique reminds me of scraping Amazon's Cloud:

site:.s3.amazonaws.com "keyword"

site:.cloudfront.net "keyword"

What you get here are mostly PDF files, but you might find an unlocked bucket.
 
Glad the bitching fest is over

Fatboys bot works flawlessy and is essentially idiot proof. Its designed to scrape target urls text body. It doesnt apply the google.webcache to links automatically, you need to manually do it.

In regards to the content

I understand the concept here that once content clears the cache its essentially "unique". Though you would think (you would think) google would have some kind of safeguards or backup-backup cache here for being aware of duplicate content. If not for everyday sites then simply due to the size of the de-indexing, and the fact its yahoo. Google must have known that scraping would be a concern of a site of this magnitude.

Would be interested to see if people have these articles not just ranking but ranking well.

my 2 cents
 
PM me if you can scrape niche related articles. Will pay for your time obviously.
 
Okay, I'm tired of PMs. I'm not going to release my private method, but this also works.

Scrape the URLs. There are a couple of good posts that give you a tutorial. I think someone dropped a 200k+ yahoo voices URL txt file here. Use that if you're too lazy.

Download Notepad++ and open the txt file.

Press ctrl+f and press the "Replace" tab

Find what: http://
Replace with: http://webcache.googleusercontent.com/search?q=cache:

There you go. You got 200k+ links of google cached yahoo voice articles.

http://www.metaproducts.com/mp/mass_downloader.htm

Download that. Import your URL list and press download. Don't make thread count too high or its going to skip over a lot of articles.

Make sure you have lots of proxies or a couple of good VPNs

Don't PM again about this. FFS people, I have never coded in my life and I found two methods.
 
I've got the 12k niche related articles I wanted so I'll throw out some tips. For those who are pulling from Google cache if you add a &strip=1 operator to your URL, it removes pretty much all rich formatting and leaves you with just the article (minus a few menu's on the bottom)
 
Let me clarify a point, I believe that the OP is suggesting you use these for a base of SPUN content. So you splice, spin, and reuse the articles. Not for direct use. Still not THAT useful because spammy garbage content is barely useful today. Two years ago this would have been GREAT, but not anymore IMHO.
 
Yes this content is useless, stop using it immediately, if you use it Google will give you the death penalty, so step away from this method and move on... leave this content alone
 
Back
Top