Peisithanatos
Registered Member
- Mar 9, 2013
- 59
- 23
It is working fine, thank you fatboy. The only question is to change IP to avoid Google blocks.
So I was busy on here talking to 3 random "programmers"...
Really?! Check out my itrader and see for yourself how badass I am dude![]()
I can do it for you, though I dont I will be able to build a standalone program that scrapes and parses html, I can scrape and get the articles for ya.
It seems that Google blocked my IP for too many requests lol. Guess I was too fast for Google
I will continue to work with this to see if I can manage to automate it as much as I can.
It's quite simple look up gopher protocol http://en.wikipedia.org/wiki/Gopher_(protocol)Thanks for the share man. Content has ALWAYS been my major limiting factor and I wish we could get a more comprehensive collaboration going on this. Although it might be smarter to do in private like the VIP section.
AFAIK, there are thousands, if not hundreds of thousands of databases with free books, articles, and all sorts of shit that is not published on Google.
Like here's 1 example - http://www.monmouthcountylib.org/index.php/research/alpha-content-example
It's a local online library. But I had to register in person and show my ID to get a passcode to access the database. These are suppose to be "public" databases but they're really not because they force you to show ID before you can access the databases. Same thing with many college databases. You can't access them unless you're a student.
1 thing I've always wanted to do is start something called "The Content Project". And start it over in the VIP section.
My plan is pretty simple.
1) Get 100 volunteers to start.
2) At least 25 of those volunteers I want to be college students.
3) Anyone who's graduated college (or not gone at all) would register at their local libraries online database and get passcodes.
4) Anyone in school would share their username and passcode to access college databases.
5) Then basically, anyone who volunteers 1 database would get access to all 100 databases (or more).
To avoid too many people from using the content, if you do not submit at least 1 database and passcode then you can't access any.
I've even thought about paying a developer, then using proxies and a massive harddrive to download all the content from all these databases then pool it into 1 massive private database.
And like I said, nobody could access the database unless they help build it. If their passcode expires it would be associated to their username and they'd lose access to ALL 100+ databases.
Now obviously, this wouldn't benefit everyone on the forum. And it shouldn't. It would only benefit those who actually get off their ass, go to their library and get a passcode. And imagine what benefit this would provide to BH marketers. Imagine having passcodes for 100 databases FULL of content that isn't published on google. IMAGINE what type of benefit that could provide for the few who participate. And IMAGINE if we could pool all that data into 1 database?
Shit would be crazy. You wouldn't have to write an article again for the rest of your life.
-BB
dude, no offense, but your bot is not working. it scrapes only the article names, the txt files have 0 bytes. I believe you when you say it worked but for now it isn't.
https://chrome.google.com/webstore/detail/scraper/mbigbapnjcgaffohmbkdlecaccepngjd?hl=en
http://webcache.googleusercontent.com/search?q=cache:
VT:
Code:https://www.virustotal.com/en/file/3a1798b8989bd12042b0fd05fa4d095ea655ee6de42a104d933c01a987ca2a86/analysis/1404587091/
Download:
Code:https://www.mediafire.com/?mq97920dl46a77w
The Bot is Working absolutely fine With the Cached link of the articles.
Here is what i did to get 100 articles in my Niche.
First Make google show 100 results on a single page... Go to Search settings to do so.
Install this Chrome Extention to Scrape Similar Links
HTML:https://chrome.google.com/webstore/detail/scraper/mbigbapnjcgaffohmbkdlecaccepngjd?hl=en
Now Search> site:voices.yahoo.com "Your keyword here"
This will Give you 100 Results on the search page. Just Right Click on the the first result, you should see "Scrape Similar" Click on it. Thats it You have got all 100 links Which will be saved on Google Docs. you can delete the sheet later on.
Now These are Original article link so just add ThisBefore Each Url Using MS Excel or anything... Thats it now we have Cached Url of the articles, Cross check once in your browser... Load them in Fatboys Bot and it will do its job.HTML:http://webcache.googleusercontent.com/search?q=cache:
Bot Link
when you have all cached urls , you don't need the bot . Any http downloader will do the job . Check my earlier post .
when you have all cached urls , you don't need the bot . Any http downloader will do the job . Check my earlier post .
Yep. Furthermore, you don't even need the cached urls in the first place.
There is an addon by firefox that works much better. You set results to 100, set the addon to filter "cache", click 1 button and it downloads them all super fast. You do need a proxy switcher. But I was averaging 250-300 articles per proxy and was able to scrape 17,500 articles using less than 70 g proxies. But if you need to scrape like 100,000+ articles you need a better system.
when you have all cached urls , you don't need the bot . Any http downloader will do the job . Check my earlier post .
This is hilarious. There are so many wanna be "programmers" on this forum it's unbelievable.
Fatboy uploaded a list that by now, probably more than half of those links are gone. Then he builds a "bot" that doesn't even work to download his own list. And instead of people calling him out on it, they thank him for sharing a broken bot!? If I had seen that thread when it was created, I would have called him out on it, and everyone could have had a working bot weeks ago. But I'm kinda glad he failed, because now people won't abuse the articles.
This is exactly why I'm learning visual basic. I have a regular developer I work with who could have coded this bot in less than an hour. But you gotta request a project a couple weeks ahead of time. So I was busy on here talking to 3 random "programmers" who didn't know what they were doing, most my articles were vanishing, and last minute a guy who doesn't know shit about coding... shows me a method to mass download the articles.
I'll give you guys a hint. You DON'T NEED A BOT. A certain billion dollar company (not Google) already has an addon that can do this. You just need a proxy switcher and you can scrape and download them all.
Problem is I can't reveal the addon cause the guy who shared it with me requested I don't. But if you look hard enough, and think very clearly with common sense chances are you'll find it.
-BB
I@fatboy - Nevermind. I just figured it out and deleted my response. I don't know why the bot works this way but it takes about 10 minutes just to start doing anything. I tried it again but this time I accidentally forgot to close the program (got a phone call), came back 10 minutes later and noticed it had started scraping articles. So the mystery answer to my question seems to be "press start then ignore the program for 10-15 minutes until it starts running". Sry for complaining it's just seems like a bug or something.
The Bot is Working absolutely fine With the Cached link of the articles.
when you have all cached urls , you don't need the bot . Any http downloader will do the job . Check my earlier post .
I just read through this entire thread and when I hit this post I was like "what an arrogant twat". Sorry but you got a shitload to learn mate. You flamed a respected and generous BHW member while leeching a method like a know-it-all from someone much smarter than yourself. Then you have to come on here and tell everyone the leeched method works great but you can't share it with your, "I got some candy and screw you guys", look at me post.
I voted against your promotion to JrVIP because you were such a new member. Others argued you were mature and a nice guy. Well this one post confirmed just what I suspected. How dare you flame Fatboy while having no skills of your own.
Before anyone says this post is sour grapes because BnB won't share his candy; imagine my scraper with over 5 thousand google passed proxies ability to harvest this crap if I wanted it. I certainly don't need any methods from BnB but he sincerely chapped my ass putting down a well respected BHW member that has contributed so much to BHW over the years. BnB learn some respect mate. BTW, you can edit and even remove your posts as a JrVIP.
Wget in a bash script is all you really need, bit a lot of people won't go down to that for whatever reasons - the plugins and bots are just for convenience.
I just read through this entire thread and when I hit this post I was like "what an arrogant twat". Sorry but you got a shitload to learn mate. You flamed a respected and generous BHW member while leeching a method like a know-it-all from someone much smarter than yourself. Then you have to come on here and tell everyone the leeched method works great but you can't share it with your, "I got some candy and screw you guys", look at me post.
I voted against your promotion to JrVIP because you were such a new member. Others argued you were mature and a nice guy. Well this one post confirmed just what I suspected. How dare you flame Fatboy while having no skills of your own.
Before anyone says this post is sour grapes because BnB won't share his candy; imagine my scraper with over 5 thousand google passed proxies ability to harvest this crap if I wanted it. I certainly don't need any methods from BnB but he sincerely chapped my ass putting down a well respected BHW member that has contributed so much to BHW over the years. BnB learn some respect mate. BTW, you can edit and even remove your posts as a JrVIP.
any idea when yahoo voices articles will deindex? I ran some through smallseotools plagiarism checker and articles still exist in google