Hello, i scraped my own list to GSA (articles) using scrapebox, but i have a problem to filter it. I was trying to find some tutorials and tried to figure it out by myself, but still got a big problem. I was trying to filter it by outgoing links,external links, how many pages are indexed, TF, CF etc. etc. i was using a lot of tools but i couldn't find any metrics how to recognize not spammed domains in bulk. Because it's easy to check anchor text of this domain, but you have to do it manually. Can you recommend me anything? I was thinking also about bad words filter
Slightly confused about what your asking.
You scraped a list of urls to post to and your asking how to filter it based on certain metrics?
I mean Scrapebox is awesome, don't get me wrong, you don't have to look far to see I have not a word bad to say and I think my over 100 videos prove I have plenty of good to say. But you can do all this in GSA. You just have to set it up in your project and GSA will do all this on the fly.
I mean you can use Scrapebox, but its a lot more steps to get to the same end, personally I try to work smarter not harder. I would just set it all up in GSA and then let it do its thing. So what if it takes a few mins longer, its still going to be quicker then you trying to do all this in Scrapebox and compile it all together, not because scrapebox is slow, but we as humans aren't 100% efficient.
That said, if your scraping article targets for GSA in general you will notice its going to create a new page to post your content. So with your article and your link on its own page, you probably don't need most of those filters.
Also note if you start using a load of bad word filters and all these metrics and such you liable to eliminate your entire list or most of it. I sell lists and Ive had more then 1 user getting exactly 0 links because they set things so aggressive they eliminate 99.999999999% of the entire internet. (ok maybe I exaggerated a tiny bit, but they eliminate 100% of potential targets GSA SER can post to anyway).
The last note is that this is the scrapebox sales thread. your question was about Scrapebox and thats fine, but now that I have directed you towards GSA this isn't the thread for 300 GSA followup questions. However its a big forum and we have room for those questions in the right thread.
Okay, thank you for your response. My proxy issue is indeed sorted out now and that was indeed as I suspected me doing something wrong.
For point 3, I still have the same issue. Downloading the default engines hasn't changed anything. I didn't think it would have seeing as I had already tried to install a fresh copy.
I have a similar result scraping yahoo instead of bing:
http://i.imgur.com/S3Hf9qo.png
I typed up a nice long response and my browser decided to foul up and delete the whole thing (Im sure it was my fault but who wants to take the blame for that? Far better to blame the pc, LOL)
Anyway, I can't reproduce this and Ive spent a good long while trying. Try it with and without proxies, custom harvester and detailed harvester as well, does it do this in every case?
Do you have the engine so.com and sogou.com at the bottom of your engine list? You should or your not using the latest engine file.
Have you tried kicking your machine? just kidding, lol
Maybe you can give screens of the settings >> connection timeouts and other settings - all 4 tabs and the settings >> harvester engine configuration - and click on 1 of the engines like bing or yahoo that are giving you trouble.
Do other engines like deeperweb or google work?