Permanently Closed Marketplace Sales Thread

Status
Not open for further replies.
Thanks for the quick reply man,
Yeah, the single harvester is for the delay and for checking why my harvester stops
It just seems that it uses all the proxies until all of them get 404..
Version 1.16.4
Will the multi harvester not remove proxies that receive 404 ?

Your welcome. No the multi harvester will just keep cycling thru the proxies endlessly, even if they return a 404 a thousand times. However the multi harvester does not work with the delay, so you can set it to 1 connection but there will be no delay.
 
I have changed to my new vps 10 days ago and submitted activation from that vps. it worked for few days and now it's asking me for activation details on same vps. i purchased using a stealth paypal and no longer have access to paypal email. i have sent a email to support with all details and that bug file.

Is this a normal error or is this only happening with me ?
-=-
 
Sometime between now and December 31st 2015. Hehe. Honestly though I don't know. I even asked them, because Im going to do some videos and I was going to take on some other projects and stuff and didn't want to load up too many things at once

I asked for an ETA in another thread - so I'll edit it. :) Looking forward to the update. Adding some thanks too for your vids. GREAT JOB!
 
Great, I will try that, thanks!

your welcome!

I have changed to my new vps 10 days ago and submitted activation from that vps. it worked for few days and now it's asking me for activation details on same vps. i purchased using a stealth paypal and no longer have access to paypal email. i have sent a email to support with all details and that bug file.

Is this a normal error or is this only happening with me ?
-=-

Its only happening to you, but its something support will have to sort for you, because its licensing related.

I asked for an ETA in another thread - so I'll edit it. :) Looking forward to the update. Adding some thanks too for your vids. GREAT JOB!

Thanks for the thanks!
 
I can't wait for 2.0, dreaming about It!

Maybe I can cut down on my now $800+/m for proxies (24/7 scraping) :O

I would pay to get In on any sort of beta. 64 bit FTW, goodbye 1.8GB memory limit.

LoopLines 2.0 tease video Is fap worthy.
 
I can't wait for 2.0, dreaming about It!

Maybe I can cut down on my now $800+/m for proxies (24/7 scraping) :O

I would pay to get In on any sort of beta. 64 bit FTW, goodbye 1.8GB memory limit.

LoopLines 2.0 tease video Is fap worthy.

LOL. Do you really spend $800 on proxies a month? Thats hefty.
 
LOL. Do you really spend $800 on proxies a month? Thats hefty.

Yep! Not only for scraping though, a few other tools I use. Can't afford to have any downtime, so I fork out the cash. Proxy budget has been Increasing every couple months due to workload.

Being able to use auto updated cloud proxies could cut that down a bit. I guess we will have to see between now, and the end of 2015 by your oh so tight Scrapebox 2.0 ETA :)
 
Last edited:
Yep! Not only for scraping though, a few other tools I use. Can't afford to have any downtime, so I fork out the cash. Proxy budget has been Increasing every couple months due to workload.

Being able to use auto updated cloud proxies could cut that down a bit. I guess we will have to see between now, and the end of 2015 by your oh so tight Scrapebox 2.0 ETA :)

Yeah thats a super tight ETA, I like to be accurate. :D

Have you tried things like reverse proxies back connect proxies (change the ip on the back end every 10 mins, I think they are private proxies) or there is also proxy rack, which just changes the ip on every request. I use proxy rack and its pretty impressive. Its slow, and public proxy based I believe, but it works well for the things I use it for and I have the 200 connection pack and its still faster then using filtered public proxies because they do all the work on the back end.

Also there is solid proxies or something like that, they are newer, and they utilize an api where you can request a new proxy. So they give you a fixed IP just like reverse and proxy rack, but then when the ip becomes blocked by google/search engine/site your using it with, you can programmatically request that the ip be changed on the back end. Its a nice mix, because proxy rack changes it on every call and reverse changes it every 10 mins but with the solid solution (or whatever its called, can't remember) its changed on demand based on when you want it. So you could specify based on X error request a new proxy.

Thats not integrated with Scrapebox, but you could build your own api tool that just requests a new ip every X often based on your experience of when you need it. Anyway, just tossing it out there, might save you some cash with some of those.

When the automator comes out in 2.0 I plan to build my own automator file that runs that grabs proxies from my own custom sources and filters and kicks them out to a file every X mins that I can then share out to my servers via dropbox. Maybe anyway, I just got done implementing a completely automated scrape and filter setup between multiple servers and its running great with paid stuff so not sure I will need it. Its an option anyway.
 
Yeah thats a super tight ETA, I like to be accurate. :D

Have you tried things like reverse proxies back connect proxies (change the ip on the back end every 10 mins, I think they are private proxies) or there is also proxy rack, which just changes the ip on every request. I use proxy rack and its pretty impressive. Its slow, and public proxy based I believe, but it works well for the things I use it for and I have the 200 connection pack and its still faster then using filtered public proxies because they do all the work on the back end.

Also there is solid proxies or something like that, they are newer, and they utilize an api where you can request a new proxy. So they give you a fixed IP just like reverse and proxy rack, but then when the ip becomes blocked by google/search engine/site your using it with, you can programmatically request that the ip be changed on the back end. Its a nice mix, because proxy rack changes it on every call and reverse changes it every 10 mins but with the solid solution (or whatever its called, can't remember) its changed on demand based on when you want it. So you could specify based on X error request a new proxy.

Thats not integrated with Scrapebox, but you could build your own api tool that just requests a new ip every X often based on your experience of when you need it. Anyway, just tossing it out there, might save you some cash with some of those.

When the automator comes out in 2.0 I plan to build my own automator file that runs that grabs proxies from my own custom sources and filters and kicks them out to a file every X mins that I can then share out to my servers via dropbox. Maybe anyway, I just got done implementing a completely automated scrape and filter setup between multiple servers and its running great with paid stuff so not sure I will need it. Its an option anyway.

Thanks for the suggestions. I tried Proxy Rack more than a year ago (I think), I will have to give them another try for a month or two. Are you talking about solidproxies.com? Just a blank white page for me. Unless they have a payment/member page elsewhere I'm not aware of.

A big part of my proxy budget does come from social media though (managing and building pretty huge accounts). If you don't mind me asking, who do you use for regular private proxies (If any)? I have contacted a few providers and have gotten 50-100 private proxy samples for a week or two at a time, but they just have not been up to the task (and even gotten a few accounts banned, which thankfully I got back).
 
Thanks for the suggestions. I tried Proxy Rack more than a year ago (I think), I will have to give them another try for a month or two. Are you talking about solidproxies.com? Just a blank white page for me. Unless they have a payment/member page elsewhere I'm not aware of.

A big part of my proxy budget does come from social media though (managing and building pretty huge accounts). If you don't mind me asking, who do you use for regular private proxies (If any)? I have contacted a few providers and have gotten 50-100 private proxy samples for a week or two at a time, but they just have not been up to the task (and even gotten a few accounts banned, which thankfully I got back).

I use buyproxies.org for my regular private proxies, but I don't use those with social accounts. But yes its solid, and as noted below they work for me too.


Yes working fine for me too, its heavily scripted though, so could be an issue with the browser there maybe. Haven't actually tried them.
 
Hi
I have a static ip and so, if I don't use the proxy and will be banned, then can I scrape again without proxy or I will be banned to always ?


Thanks.
 
I probably did something wrong but I am using V2 .0.0.11 and scraped 14million tumblr blogs and I had auto remove dups enabled so I thought that would do that on the fly but it didn't. The text file is 2GBs and while it has imported into SB it is taking forever to dedupe it.

Is there a better setting that removes dups on the fly as it is harvesting to keep the size of the file down?

Isn't there a way to split up files once they hit a certain size?

Thanks
 
I probably did something wrong but I am using V2 .0.0.11 and scraped 14million tumblr blogs and I had auto remove dups enabled so I thought that would do that on the fly but it didn't. The text file is 2GBs and while it has imported into SB it is taking forever to dedupe it.

Is there a better setting that removes dups on the fly as it is harvesting to keep the size of the file down?

Isn't there a way to split up files once they hit a certain size?

Thanks


You did not do anything wrong. The auto remove function is only related to importing/loading urls, not for harvesting urls.
 
Hi
I have a static ip and so, if I don't use the proxy and will be banned, then can I scrape again without proxy or I will be banned to always ?


Thanks.

Probably not. I mean google will unban you in 12-48 hours on average. It will be short at first but the more times you get it banned the longer the ban will last.

But you can set a delay and then set it to RND and choose a random (RND delay range) delay under settings >> adjust RND delay range and put it to 55 min and 60 max seconds and then go to settings and uncheck the use multi threaded and use custom harvester. Then harvest and you "probably" won't get it banned unless your using some solid advanced operators.

But thats all with Google, other engines are up for grabs as I haven't tracked them all. Yahoo seems ban happy, Bing seems more tolerant, and loads of others in the custom harvester.
 
Hi
bought scrapebox today
I don't have access to my main paypal email now (where scrapebox bot sent me the email0
i donwloaded scrapebox and entered the details needed to activate it, and i insert my second email (not the one I purchased from)

where will the license be sent? to the email i specified when i activated the program or the one I bought from? if second - how can i get my license if i don't have access to my main paypal email now?

thx
 
Hi
bought scrapebox today
I don't have access to my main paypal email now (where scrapebox bot sent me the email0
i donwloaded scrapebox and entered the details needed to activate it, and i insert my second email (not the one I purchased from)

where will the license be sent? to the email i specified when i activated the program or the one I bought from? if second - how can i get my license if i don't have access to my main paypal email now?

thx

There is no "license" that will be "sent". The license details are your name, email and TID.

However you must use the email that the paypal payment was sent from, so right now probably nothing will happen as it will see you as having sent in the wrong details.

Click activate and enter your name and the paypal email and your paypal transaction/receipt ID.

Then you can mail

support (at] scrapebox [dot) com

and ask that they change your license mail to your 2nd mail or whatever you want, unless you normally have access to your paypal mail in which case it doesn't matter. In any case make sure you save the name/paypal license email/TID for future reference.
 
I have a problem

I try to use my own Ip to scrape google

But every time I try to start scraping there is 320 error 'IP blocked"

I didn't even manage to scrape 1 page, and my Ip is not blocked, I can harvest as many pages as I want using my browser and google doesn't ban me. What can be the problem?


And even when I change Ip (have dynamic IP) the same problem happens(and still I can do whatever I want in google using any browser, problem happens just using scrapebox)
 
Last edited:
Status
Not open for further replies.
Back
Top