New ScrapeBox Plugin [Automator]

If I get it right, you mean some sort of "If only a domain is returned, remove the url" and "If the url has a path (domain/path), but the path is less than x character long, remove the url"?

exactly, first part is correct "If only a domain is returned, remove the url" and "If the url has a path (domain/path)" keep the url.
but second part is just another solution/idea for it, but its fine as long it return url (domain/path) not only domain so you got the first part correct.

you can apply both options if you like but what i meant and my aim is the first part between quotes "to remove domains returned from harvester and only keep domain/path/".

i was thinking of this few months ago before the realase of automator, to be applied to scrapebox past harvester options itself.

Thanks
 
If it was free, maybe more people will love and appreciate it:D .. this is my opinion about that
 
If it was free, maybe more people will love and appreciate it:D .. this is my opinion about that

You don't think SB provides enough value for $57? This is literally the only program that I've EVER purchased that I'm still using years later. I think SB would still be a great value at $200 and I would keep it if it were $50/month.
 
If it was free, maybe more people will love and appreciate it:D .. this is my opinion about that

Id say you are either ignorant of the value that scrapebox provides, your a cheapskate, or your just don't value others hard work. "Thats is just my opinion about that."


+Rep to you for being a sensible person BHopkins
 
exactly, first part is correct "If only a domain is returned, remove the url" and "If the url has a path (domain/path)" keep the url.
but second part is just another solution/idea for it, but its fine as long it return url (domain/path) not only domain so you got the first part correct.

you can apply both options if you like but what i meant and my aim is the first part between quotes "to remove domains returned from harvester and only keep domain/path/".

i was thinking of this few months ago before the realase of automator, to be applied to scrapebox past harvester options itself.

Thanks

The feature to remove the url when it is just a domain will be in the next update of Automator.
 
I want to use this tool to harvest continually. How do i go about this?

What i would like to do is as follows:
load proxy list from file
harvest from google and yahoo with my own custom footprints
once it's harvested 1,000,000 urls, i'd then like it to remove duplicates, and save this list to a file
I'd then like it to remove from the footprint list, the footprints it had success at harvesting with, and save the failed footprints yet to be harvested to a file

I'd then like the process repeated

With it loading the footprint list file, and harvesting again

This seems pretty straight forward, but I can't see a way in which I can save a footprint list within the program!!??
 
I want to use this tool to harvest continually. How do i go about this?

What i would like to do is as follows:
load proxy list from file
harvest from google and yahoo with my own custom footprints
once it's harvested 1,000,000 urls, i'd then like it to remove duplicates, and save this list to a file
I'd then like it to remove from the footprint list, the footprints it had success at harvesting with, and save the failed footprints yet to be harvested to a file

I'd then like the process repeated

With it loading the footprint list file, and harvesting again

This seems pretty straight forward, but I can't see a way in which I can save a footprint list within the program!!??

Well there is no "built in cycle feature" so you would have to load the job and then just save it and merge off tons of copies of it. So you would have to guess how many cycles it would take to make it thru the keyword list, or else it wouldn't' finish.

There is no "export failed or uncompleted footprints/keywords" at this time.

Why not just let it scrape up all the footprints, get X millions, and then go into the harvester sessions folder when its done and just use the dupe remove addon to manually remove duplicates? Would only take a couple of mins.
 
I made a new video for the Automator, might help people understand better.

 
Last edited by a moderator:
Cool. I am thinking of buying this now.
Thanks for the video loopline.
 
Cool. I am thinking of buying this now.
Thanks for the video loopline.

Your welcome mate, this has saved me so much time its not even funny. And I find new uses for it as time goes on.
 
If you want to get a keyword list done as good as possible, you could try unticking the use multi threaded harvester, in the settings menu. This would be slower, but the single threaded harvester is built for accuracy, while the multi threaded harvester is built for mass speed. , at the cost of occasionally skipping some keywords.

While the multi harvester skips keywords that fail too many times, the single threaded harvester loads all footprints/keywords into an array and works thru them one by one. If proxies fail it keeps tyring. So it will literally just sit there endlessly trying new proxies, until it gets results for your keyword. So you get 100% completion every time, but it is a "single threaded" harvester, so its slower.

Option B for that would be that if you find you typically run a list 3 times to get what you want out of it, just open up 3 instances of scrapebox at the same time, load the same automator job file into each one and let it scrape the same list in each of the 3 instances at the same time.

Then when you are done, use the dupe remove addon to merge all results and remove duplicates. Its not especially resource efficient, but its time efficient. I do this when trying to make sure I get everything out of list when posting for AA lists etc...

thanks loopline, top tips here

was wondering how people get a complete keyword list scraped without the export uncompleted keyword function in the automator

just takes some creative thinking i see

However, if I were you, and using it for GSA, I would just let sbox load in the full keyword list and then just spend a couple mins manually using the dupe remove addon at the end to take all the results from the harvester sessions folder and merge them, remove duplicates and then split them whatever size chunks you want.

But I feel like that would be easier to work with then splitting the keyword list, since your taking it outside of sbox.

curious... if url list is taken outside of sb, why not split the keyword list?
 
thanks loopline, top tips here

was wondering how people get a complete keyword list scraped without the export uncompleted keyword function in the automator

just takes some creative thinking i see



curious... if url list is taken outside of sb, why not split the keyword list?

You could, you just have to touch it more times. Like if you split a large keyword list into 4 pieces then you have to run it 4 times and then do import and compares on each list to remove urls that you had in previous lists. Or just merge all the results together and then remove dupes.

So you have to "Touch" and take time to interact with scrapebox all that many more times, vs just loading in one big keywords list, manipulating the urls after harvest and then spiting the end list. I like to set and forget it.

Like I had a non urgent project so I loaded in a massive keyword list and some private proxies and set connecitons to 1, just because I wasn't in a hurry. That was a few days ago, its got 19 million urls harvested and counting. I haven't touched it since. Its like all the gadgets you see sold on infomercials. Add water and dirt and set and forget it and come back to a gourmet meal in 6 hours it slices, it dices it chops, it spins, it even does your laundry on the side! HAHA I dunno. Thats just how I think, less touch from me the better.


~~~~~~~~~~~

As for your first Q, I just don't care about completed keyword lists. Meaning If I load in 500K keywords and use private proxies, my speed and success are consistent. Private proxies are key to this, set connections at 10%-20% of your proxies, so you have 100 proxies set connections to 10-20. That adds enough of a delay that they never get banned.


So if I load in 500K and 1000 keywords total don't complete, I don't care, cause I got 499K keywords and that more then enough for me to deal with. High efficiency comes from the private proxies though and whats left I just don't care about cause the whole concept is about mass and when I scrape 120 million urls, I don't care about the 300K that I missed. ect...

Thats the way I see it. If I only had 100 kewyords well that would be different, but Id anticipate that with private proxies that I would get a 99-100% success rate on scraping those 100 keywords anyway.

So for me its about changing the picture. I try to look at what some people say , such as "how do you get a keyword list scraped efficiently without exporting uncompleted keywords" and say 1.) Do I need to be effecient with it?
2.) Why are so many failing in the first place, is it my proxies? Is it my settings?

That way I arrive at the point where it is efficient and I might only be missing a few keywords and then the need for the keyword export is mute. But thats not the driving factor for me. For me the driving factor is, that by arriving at the point where I don't need to export keywords and then reimport and scrape again, I have increased my effeciency and I don't need to touch scrapebox as often and it doesn't take as long to run.


At the end of the day many have said it but most recently I read it in Shoemoneys new book.

"Time is all each and every one of us has. How we use our time is what differentiates us all in the end." ~Jeremy Schoemaker

Great read by the way, highly crass, but Id recommend "The Shoemoney Story" by Jeremy Schoemaker to anyone in IM "if they are mature enough to read past some of the crassness and ignore the "ideals" that you may not agree with. The life skills and business skills are the golden nuggets and its a fun read, as he says so often in the book, he blurrs the lines, his book blurrs the lines between "business/self help" and "a story book".


Anyway.. Dunno how I got off on the book but the point was the quote. I don't seek to make the tool work for how I want to use it, I seek to use my time better and to do that I try and change the picture of what I think the problem is that is keeping me from using the tool more effeciently. - Another Great book for that is "Thinker Toys"

I forget who said it but they said something to the effect of

"The only two things that influence how you shape your life and who you become are the books you read and the people you meet". Hence all the book mentions. You sound like a person who is the type of person who helps themselves. Some people don't care to be helped they want you to do it for them or hold their hand, you seem like your going somehwere so I dropped the book mentions.

Ok, enough about books in the software thread, lol Cheers!
 
So for me its about changing the picture. I try to look at what some people say , such as "how do you get a keyword list scraped efficiently without exporting uncompleted keywords" and say 1.) Do I need to be effecient with it?
2.) Why are so many failing in the first place, is it my proxies? Is it my settings?

That way I arrive at the point where it is efficient and I might only be missing a few keywords and then the need for the keyword export is mute. But thats not the driving factor for me. For me the driving factor is, that by arriving at the point where I don't need to export keywords and then reimport and scrape again, I have increased my effeciency and I don't need to touch scrapebox as often and it doesn't take as long to run.

brilliant, never thought of it that way!

guess my mind was stuck on how to automate and adjust the next run for uncompleted keywords... never did i think to reverse it and maximize success rate first!

At the end of the day many have said it but most recently I read it in Shoemoneys new book.

"Time is all each and every one of us has. How we use our time is what differentiates us all in the end." ~Jeremy Schoemaker

good reminder. here is another good one -> "imperfect action is always better than a perfect plan (which really doesnt exist)"

trying to take imperfect action daily!

"The only two things that influence how you shape your life and who you become are the books you read and the people you meet". Hence all the book mentions. You sound like a person who is the type of person who helps themselves. Some people don't care to be helped they want you to do it for them or hold their hand, you seem like your going somehwere so I dropped the book mentions.

thank you for the encouragement!!! will take it to heart :)



---------
question
---------

just one more ;)

Private proxies are key to this, set connections at 10%-20% of your proxies, so you have 100 proxies set connections to 10-20. That adds enough of a delay that they never get banned.

been playing with sb settings and trying to maximize sb so i don't have to touch it throughout the day too

adjusting the number of connections from my tests definitely affects temp bans

i did some tests and found not much difference between setting the "Adjust RND Delay Range" (only tested 1-2 mins v.s. 1-3 mins though). you found the same?
 
brilliant, never thought of it that way!

guess my mind was stuck on how to automate and adjust the next run for uncompleted keywords... never did i think to reverse it and maximize success rate first!



good reminder. here is another good one -> "imperfect action is always better than a perfect plan (which really doesnt exist)"

trying to take imperfect action daily!



thank you for the encouragement!!! will take it to heart :)



---------
question
---------

just one more ;)



been playing with sb settings and trying to maximize sb so i don't have to touch it throughout the day too

adjusting the number of connections from my tests definitely affects temp bans

i did some tests and found not much difference between setting the "Adjust RND Delay Range" (only tested 1-2 mins v.s. 1-3 mins though). you found the same?

Thats a good quote as well. All too often I get tied up in making "the worlds greatest plan of attack" for any given element and don't actually ever get anything done. :)


I don't actually use the RND delay, in fact I dont even use the delay of 1-10 seconds very often. (on a grand occasion.) The only time I use any delay is if I am not using proxies and Im not in a hurry and I just want to harvest slowly with my IP.

But the Delay only applies to this:

Single Threaded Harvester
Ping Mode
Page Rank Checker
Email Grabber

There might have been one more thing added, can't remember, its been a while since I double checked. But generally you wouldn't need a delay otherwise anyway.

So if you using it with the multi threaded harvester or another function you can set delays all you want and they won't fire.

But regardless by setting connections to no more then 20% of your proxies, err... like if you had 10 private proxies set connections to 2, then it creates essentially, a artificial delay as scrapebox rotates thru the ips. It works out perfectly just long enough that IP bans are almost non existant.

Plus you can do big harvests, like I plan out which instances are using which proxies on which servers and they all get dedciated private/shared proxies for their specific tasks that are repetitive. So I set the connections appropriately for each, that way at no time is 2 instance accidently using the same proxies for harvesting from google which would in turn double the connections and push me past 20% or 1/5th and then cause an IP ban.

But I did take 30 proxies or something that were already being used and shove it into a new instance as a test and set 1 connection. So not really pushing me over by much out of 30 proxies and let it set. I dunno how long it ran, but it stopped at 1 connection on private proxies at 45 million results harvested. And that was only because the server got disconnected from the internet completely due to a hosting issue. No doubt it would still be going.

The point simply being, too many people get hung up on "I need to use 50 connections to harvest or 200 so it doesn't take forever" well by the time they constantly refresh their public proxies, recycle their uncompleted keywords, and touch it all the time (less now with the customer harvesters auto refresh on the fly, but still have to mess with uncompleted keywords) that 2 days later they have their list and mine finished in a day and a half and I only had to touch it once. If that makes sense.

Anyway, as irony would have it, think out side the box, with scrapebox. If you read someone saying "do this method" reverse engineer that method and do the opposite or change/vary it. No one writes their best stuff and shares it, if its really good, they keep it close to the vest.

Thats how I approach anyway, figure out how it works, so that I can apply it in non standard ways. I know a lot of other people do good like that too. Keeping getting imperfect work done!

I also like

Ready Fire!, Aim.

and

Do something, Do Anything. I think that is from "eat that frog" also a good book. It has a swiss cheese approach I take a lot of times when Im unmotivated as well and by the time you get in the middle of it, you all of a sudden have the motivation. Anyway, the quotes and books go on forever, but imperfect action is key. I need to go do some of that right now, lol.

Cheers!
 
Back
Top