New ScrapeBox Plugin [Automator]

not happy i baught scrapebox and scrapejet because it was not automated. well pizzled. Luckily $20 ain't a lot of money lol.

You and BHopkins mentioned that it's possible to limit the results - this was also the only reason for me to buy it (its atm a little bit limited - I would prefer a simple scripting possibility). But perhaps I'm blind, but I can't find the possibility to limit the harvested results. There is a possibility to limit the results per keyword (1-1000) (as normal SB would offer it too), but I don't see any possibility to limit the harvested url's in total.

I work with huge keyword lists (for example 170k) per footprint. This needs multiple runs because not all keywords could be searched/harvested completely. When I let SB hit the 1 mio border then I can get the results from the related folder but I can't see which keywords are already done and which not. This forces me to manually stop the harvester at about 950k urls, save found results, save open keywords, clear everything and continue till I've worked trough the most part of the keywords. Is there a workaround for this? I'm sure I'm not the only one with this problem...
 
You and BHopkins mentioned that it's possible to limit the results - this was also the only reason for me to buy it (its atm a little bit limited - I would prefer a simple scripting possibility). But perhaps I'm blind, but I can't find the possibility to limit the harvested results. There is a possibility to limit the results per keyword (1-1000) (as normal SB would offer it too), but I don't see any possibility to limit the harvested url's in total.

I work with huge keyword lists (for example 170k) per footprint. This needs multiple runs because not all keywords could be searched/harvested completely. When I let SB hit the 1 mio border then I can get the results from the related folder but I can't see which keywords are already done and which not. This forces me to manually stop the harvester at about 950k urls, save found results, save open keywords, clear everything and continue till I've worked trough the most part of the keywords. Is there a workaround for this? I'm sure I'm not the only one with this problem...

I can't find this either. This is the one thing I want to use it for. SB won't allow more than 1 million harvested results. I want to stop at 1 million, remove dupe domains then start again using my custom footprints.
 
I can't find this either. This is the one thing I want to use it for. SB won't allow more than 1 million harvested results. I want to stop at 1 million, remove dupe domains then start again using my custom footprints.

Exactly - thats the only thing I'm wanting too and thought a lot time about writing some automation software for this scenario myself (but don't have the time to experiment). I don't get why is this so hard to implement, especially with the possibility of an automator plugin. If someone is able to get this done via automator plugin please let us know.

The currently offered functions are quiet limited compared to the functions of SB itself - for example loops aren't possible and a lot other functions aren't choosable.

What I also would like is a possibility to have a growing list of proxies checked every day and exported as "anon proxies" and "google proxies". This is possible to do but to have these lists seperate I have to play around with the textfiles (wrote a simple batch which merges the old anon list with the new proxylist) themself AND do the whole process two times with SB (because at the moment there is only the possibilit for "google proxies" or (if unchecked) "anon proxies".
 
I will say that SweetFunny is excellent about updates. I expect the automator plugin to be excellent soon enough.
 
I will say that SweetFunny is excellent about updates. I expect the automator plugin to be excellent soon enough.

I hope so :) SB is of course just great for this price and had really regular updates. But such a function (related to harvesting) would be a killer feature.
 
I think I used the SB contact form on their website to post a bunch of Automator concerns, centered on harvesting. I got a reply and a promise that some of them would be implemented. I suggest y'all do the same.
 
I think I used the SB contact form on their website to post a bunch of Automator concerns, centered on harvesting. I got a reply and a promise that some of them would be implemented. I suggest y'all do the same.

You're completely right - I've did it too and explained everything as detailed as possible.
 
You and BHopkins mentioned that it's possible to limit the results - this was also the only reason for me to buy it (its atm a little bit limited - I would prefer a simple scripting possibility). But perhaps I'm blind, but I can't find the possibility to limit the harvested results. There is a possibility to limit the results per keyword (1-1000) (as normal SB would offer it too), but I don't see any possibility to limit the harvested url's in total.

I work with huge keyword lists (for example 170k) per footprint. This needs multiple runs because not all keywords could be searched/harvested completely. When I let SB hit the 1 mio border then I can get the results from the related folder but I can't see which keywords are already done and which not. This forces me to manually stop the harvester at about 950k urls, save found results, save open keywords, clear everything and continue till I've worked trough the most part of the keywords. Is there a workaround for this? I'm sure I'm not the only one with this problem...

There is a few ways you can go about this.

Firstly you can guesstimate how many keywords it takes to get about the 950K mark. Then just split your keyword list into that many lines. So lets say, for simplicity sake, that you use 170K keywords/lines per a footprint like you said. And lets say thats 35K keywords/lines gets you about 950K results.

Just open up the dupe remove addon and in the file splitter section load it in. Then set it to 35000 and it will split your keyword file into 5 parts, with 35K in the first 4 and the rest in the last one.

Then just create your automator job to do what you want and select the 35K file. Then after its done, save it. Then merge it back in 4 more times, and on each of these for, just change the input file to one of the other 4 files.

Load in scrapebox and walk away. Follow?


The quick way to determine your keywords (sort of quick) is to just do a scrape in scrapebox manually for your 170K keywords. Take the results and divide them by your 950K mark. Then divide your 170K into many files.




Another way to do it, is to just scrape your keywords in scrpebox manually, and then use the dupe remove addon to mass dupe remove them and break into however big of a file you want, and then use the automator to do the rest.


The reason that it can't do more then 1mill lines is that scrapebox harvested urls function uses a microsoft designed grid. The grid doesn't allow for more then 1 million lines. Since the automator is not its own "harvester/poster etc.." it just tells scrapebox to "go" or "do", so it still uses scrapebox to do everything. Addons are external programs and for a script to tell an exe (scrapebox) to launch another exe (the dupe remove addon) and then feed it data and then get data back from it, has virus written all over it. Aside from all the extra coding, getting anti-virus programs to allow it could be a nightmare, so Id guess thats a ways off.

At first I was like, hmmm thats annoying to have to deal with splitting my keyword files into so many smaller files. I mean I came across this in the beta, while I was still wrapping my head around the addon. At that point I had to split keyword files into 5K lines or smaller to get it to work. - Needless to say the automator is much more robust then it was when I first messed with it :) So you dont' have to go near that small now, because you can load the file into scrapebox.

What I took away from that was that I was getting ahead of myself. I dont' know what your doing with 170K keywords, but I assume you want to post to the content (otherwise you wouldn't care if you had to manually remove urls, if you were using the data outside of sbox), and you scrape up 5 million results and remove dupes and get down to 2.6 million results and you load that into the poster, it might work, or it may well crash scrapebox with an out of memory error.

Ive found that by being forced to limit my output on scrapes to 1 million, that it works out better, I typically get about 300K-400K to post to, after my dupes are removed and thats a nice size list to ensure that it always posts and completes (because crashing in the middle of an automator job, and you come back 3 days later when you expect it to be done and find it crashed 2 days ago, and you thought all would be good, because you qued it up for several days worth of work - thats annoying, Ive done that).

So the whole 1 million issue actually saves me a lot of time and hassle on crashes from me not thinking thru how large of an end result list might get loaded into sbox in some other function that I tell the automator to do.


Now, it would "seem" ideal to have scrapebox stop after X results harvested and remove dupes and then keep going. That might be ideal, so we could tell it the ultimate list size we want to post to, but the current issue is thats not how scrapebox was built. Automator just tells scrapebox what to do next. A lot of scrapebox core code had to be rewritten just to make the automator work in the first place, because sbox wasn't made for to be automated to this level. It just wasn't forseen.

So to get scrapebox to stop in the middle, based on a number of urls, and then do other functions, and then come back to havesting etc... It would require a lot more overhaul of the core scrapebox. I helped sweetfunny beta test the automator and new sbox updates for quite some time before this came out. Scrapebox is built on several years work, and lots of things depend on each other. So as they would send me an update and do X, then Y and Z would break. Id send that info back and then they would fix Y and Z and C and R would break.

The point is, we did a Lot of back and forth in order so they could program scrapebox and the automator to work to where it is today. So to start rewriting scrapebox to the level it would take to make what you want work, could take massive amounts of time and testing etc... At that level, they "Could" possibly start rewriting the whole thing or a new program all together in Delphi 64bit code (can't convert whats there, have to start over due to Delphi limitations) and then 1 million file limits and out of memory errors would not even be an issue, and we wouldn't be having this discussion.

The reason of course that they haven't done that is, because it took years to get sbox to where it is and to start over in 64bit, while maintaining all the free support and license transfers and keeping up with market changes and still trying to improve the current program, is not even viable. Its a time war, same thing we all face, only 24 hours in a day.

Which brings up the next point, as The Editor kindly asked everyone to email support and fill their inbox with the same requests, and thereby cause them to have to reply to the same concerns over and over again and actually use up there time so they can't actually do any development work towards adressing any of the concerns... well you can see the issue. Its asking people to create another time issue, which prevents Sweetfunny from even trying to address any of your concerns in the first place.

Im sure that wasn't the intent, and you are thinking of it more of a "well if 10 people send in the request, they are 10 times more likely to do it" mentality, but thats not needed. They read the forums and lots of features are in sbox just because 1 person asked for them. Sweetfunny is a great judge of looking at an idea and seeing the potential. They love the scrapebox software and working on it, no one else is providing the support and value that they do for $57. Thats because they actually care, this was after all, their own personal software that they used, and decided to sell. Its their baby, if you will. Thats why they work so hard on it, so they want it to improve. I know they have Lots of things they want to do for it, but they havent' yet been able to, because of time restraints.

Anyway, point being, "some" things that I have thought I wanted solved regarding the automator, actually would have worked against me down the road. I know, I tried them manually, lol. Not to say Sweetfunny won't do something at some point, but everyone on BHW need not send them the same email with the same concerns. I want to actually see the program developed, which takes time, because I make money off it. But wasting time on emails costs us all, they read the forums every day anyway.
 
Last edited:
Thanks a lot for your detailed instructions and multiple ways to get it done Loopline (thx and rep given). I tought about the splitting solution before too, this would be a way, but this way lacks in one single point (but is usable): I have to go multiple times over the same keyword list to get it done as good as possible. With the manual process it's easy to reuse the not processed but with this strategy I don't see a possibilty to handle this.

Thanks also for explaining the 1 mio urls problem - I tought it's a 32 bit problem (memory adressing..) but this makes sense of course. It would be really great if the output isn't forced to the grid component and could written directly into a file (or multiple files) and dedublicate it in the same process (there are always a lot duplicate urls).

EDIT: Saw your edit now. I don't use the url's for SB itself, I'm using them for GSA SER. So only huge list of possible "targets" dedublicated in one file are my main target.
 
Last edited:
Thanks LoopLine for the explanation. Thanks and rep from me too.

I don't use SB for posting, but for large harvesting. Do you think there will be a time where I can load a huge keyword list in, have Automator harvest 950k, stop, save, then start with the next line in the list?

That's all I want it to do, run automated till 950k, stop, save then start again.
 
Exactly - thats the only thing I'm wanting too and thought a lot time about writing some automation software for this scenario myself (but don't have the time to experiment). I don't get why is this so hard to implement, especially with the possibility of an automator plugin. If someone is able to get this done via automator plugin please let us know.

The currently offered functions are quiet limited compared to the functions of SB itself - for example loops aren't possible and a lot other functions aren't choosable.

What I also would like is a possibility to have a growing list of proxies checked every day and exported as "anon proxies" and "google proxies". This is possible to do but to have these lists seperate I have to play around with the textfiles (wrote a simple batch which merges the old anon list with the new proxylist) themself AND do the whole process two times with SB (because at the moment there is only the possibilit for "google proxies" or (if unchecked) "anon proxies".

You can do "loops" just add the same command squence more then once. OR just save off the job and then merge it back in, Merge it back in 10 times and its going to run in a loop 10 times. Essentially. Plus this way you get granular control of all the elements of each loop, if you want it.

Can't help you out on the proxies. I just use private proxies to scrape and keep my connections to 20% of my proxies and I can scrape all day long, no need to waste time harvesting and testing etc... I occasionally use proxygos service too, and let him do the work, if I jsut need to scrape with tons of instances at once, but private proxies are so fast, I dont' often need to do this. And Im pushing millions and millions of urls a week thru sbox.

The dupe remove addons "merge" function would do it though, the whole dupe remove addon is line by line argument based, it doesn't care if the file has urls or not.


Thanks a lot for your detailed instructions and multiple ways to get it done Loopline (thx and rep given). I tought about the splitting solution before too, this would be a way, but this way lacks in one single point (but is usable): I have to go multiple times over the same keyword list to get it done as good as possible. With the manual process it's easy to reuse the not processed but with this strategy I don't see a possibilty to handle this.

Thanks also for explaining the 1 mio urls problem - I tought it's a 32 bit problem (memory adressing..) but this makes sense of course. It would be really great if the output isn't forced to the grid component and could written directly into a file (or multiple files) and dedublicate it in the same process (there are always a lot duplicate urls).

EDIT: Saw your edit now. I don't use the url's for SB itself, I'm using them for GSA SER. So only huge list of possible "targets" dedublicated in one file are my main target.


Glad thats of a help. :) Yeah, Ive had issues in the past with my browser "losing" crashing, BHW not responding etc.. and loooong posts are lost. I don't mind typing, but typing the same thing 2 or 3 times is annoying, lol. So I often stop half way thru, copy to clipboard, post, and then edit. Most of the time people aren't reading it as fast as you jumped on it. :)

If you want to get a keyword list done as good as possible, you could try unticking the use multi threaded harvester, in the settings menu. This would be slower, but the single threaded harvester is built for accuracy, while the multi threaded harvester is built for mass speed. , at the cost of occasionally skipping some keywords.

While the multi harvester skips keywords that fail too many times, the single threaded harvester loads all footprints/keywords into an array and works thru them one by one. If proxies fail it keeps tyring. So it will literally just sit there endlessly trying new proxies, until it gets results for your keyword. So you get 100% completion every time, but it is a "single threaded" harvester, so its slower.

Option B for that would be that if you find you typically run a list 3 times to get what you want out of it, just open up 3 instances of scrapebox at the same time, load the same automator job file into each one and let it scrape the same list in each of the 3 instances at the same time.

Then when you are done, use the dupe remove addon to merge all results and remove duplicates. Its not especially resource efficient, but its time efficient. I do this when trying to make sure I get everything out of list when posting for AA lists etc...

However, if I were you, and using it for GSA, I would just let sbox load in the full keyword list and then just spend a couple mins manually using the dupe remove addon at the end to take all the results from the harvester sessions folder and merge them, remove duplicates and then split them whatever size chunks you want.

But I feel like that would be easier to work with then splitting the keyword list, since your taking it outside of sbox.




Thanks LoopLine for the explanation. Thanks and rep from me too.

I don't use SB for posting, but for large harvesting. Do you think there will be a time where I can load a huge keyword list in, have Automator harvest 950k, stop, save, then start with the next line in the list?

That's all I want it to do, run automated till 950k, stop, save then start again.

Your welcome, and thanks for the rep and thanks guys!

I don't know if you will be able to do it in automator at some point, but you could do this now without much work. Just do your harvest of X millions of urls, and then just use the dupe remove addon to split them into 950K chunks. (and optionally remove duplicates as well).

Why 950K chunks?
 
Last edited:
Man, I never even knew there was such a plugin. This will make one of my jobs so much easier.

And a rank tracker too!!! I really need to start taking better notice of things.
 
Automator Plugin 1.0.0.5 new features:
- Limit harvester results
- Remove urls containing/not containing.

Not sure if anyone realizes but having the automator being able to "- Remove urls containing/not containing." is HUGE.
 
You can remove URLs that contain a certain set of keywords.
 
we also looking for option for automator to remove urls addresses itself less than x number of charachters or something similare to remove out the domain homepages it harvest sometimes. because many blogs put a brief summary of blog posts on home page and its not the actual blog post page.
 
we also looking for option for automator to remove urls addresses itself less than x number of charachters or something similare to remove out the domain homepages it harvest sometimes. because many blogs put a brief summary of blog posts on home page and its not the actual blog post page.

If I get it right, you mean some sort of "If only a domain is returned, remove the url" and "If the url has a path (domain/path), but the path is less than x character long, remove the url"?
 
Back
Top