Permanently Closed Marketplace Sales Thread

Status
Not open for further replies.
Hey guys

I'm sorry if it was asked before, but I couldn't find a direct question like this.
I want to find affiliates of the specific company using Scrapebox. I have the affiliate link footprint, and I can scrape necessary volume of web sites. Then, I can check if these pages have this footprint in their source code, that's obvious and can be easily done using SP basic tools.
The task becomes tricky if I want to find those web sites who have hidden their affiliate links under internal (or other type) of redirect. These pages won't have the affiliate footprint in their source code, therefore they won't be detected with standard SP tools.
As an option, I can scrape all the internal (or/and external) links from each scraped page - there's a a plugin for that; and then load them to redirect checker - there's plugin for that; and then somehow filter and group all the results, but it's going to be a mess. I really don't want to dig into all backlinks/redirect chains from 50k+ pages.

Is there a way to check it in an easier way? Something that could bulk check all the links on the pages INCLUDING all their redirects, and show which pages have this pattern in their links+redirect chains?


Thanks!
 
My guys have huge problems with scraping. Since the last update came out, they all had it (with old and new version of software).

They are all using Storm Proxies (and one has his own batch as well), they all used to work, but now, the scraping is slow, and way too often, not working properly. Most get minuscule set of results (2-5 hundred for the simplest, most generic query you can imagine), one gets nice numbers, but it works extremely slowly for him. Regardless of the proxies (both Storm and the other set used to work well, now neither are). Has there been any changes recently? Or do you need some further info to figure out what's going on?
 
thanks for your reply :)

I see, yes those are public proxies, hmm I am not scraping from one domain so not sure if lessening the connections would make any diff other than decreasing my scraping speed?

Also as all the domains are unique I guess there should be no prob in using my VPS's ip for email scraping at 100-200 threads? or will there by any issue?

And don't the proxies get switched in general? I mean if x doesnt work scrapebox moves to next proxy and keep doing it until it finds a working proxy... or the algo is different?


thanks again

Depends, if your only going 1 level deep then lowering the connections won't matter, if your crawling thru the site then internal pages could experience the full load of all the connections at which point lowering them would be ideal.

It has a few proxy retries, I think 3, but then it will skip the domain. With public proxies you might have 90% (or more) of them dead, slow, timeout etc.. so it could easily skip a majority of them.

Thanks, updated and changed from exact to broad and it worked.

However, now I'm getting a 'HTTP/1.1 503 service' error. Google says its a banned proxy thing but I'm not using a proxy and I can still use google on my IP.

I left it a day to see if that fixed it but it only got through a few querys before doing the same as shown in the image.

http://imgur.com/a/RoyYg

503 is an ip banned, thats regardless of if its a proxy or not.

Google has all sorts of ip bans. For starters the competition finder does Not use javascript, and your browser does. So google has way more control when javascript is used, so they will let you get more done.

Next there are just literally all sorts of ip bans, I mean some queries could work while others done, operators could be banned, while regular search works, some google services could be banned while others work. You can produce this in a browser even with javascript on, if you start using complex queries. I was just literally doing research and hit the captcha just yesterday.

I have a video on it actually

Hey guys

I'm sorry if it was asked before, but I couldn't find a direct question like this.
I want to find affiliates of the specific company using Scrapebox. I have the affiliate link footprint, and I can scrape necessary volume of web sites. Then, I can check if these pages have this footprint in their source code, that's obvious and can be easily done using SP basic tools.
The task becomes tricky if I want to find those web sites who have hidden their affiliate links under internal (or other type) of redirect. These pages won't have the affiliate footprint in their source code, therefore they won't be detected with standard SP tools.
As an option, I can scrape all the internal (or/and external) links from each scraped page - there's a a plugin for that; and then load them to redirect checker - there's plugin for that; and then somehow filter and group all the results, but it's going to be a mess. I really don't want to dig into all backlinks/redirect chains from 50k+ pages.

Is there a way to check it in an easier way? Something that could bulk check all the links on the pages INCLUDING all their redirects, and show which pages have this pattern in their links+redirect chains?


Thanks!

In Scrapebox your going to need to do it in steps, and the redirect checker is ultimately the only way to resolve the end destination url. You could then do some excel editing and get what you want. But it won't be point and click to be sure.

My guys have huge problems with scraping. Since the last update came out, they all had it (with old and new version of software).

They are all using Storm Proxies (and one has his own batch as well), they all used to work, but now, the scraping is slow, and way too often, not working properly. Most get minuscule set of results (2-5 hundred for the simplest, most generic query you can imagine), one gets nice numbers, but it works extremely slowly for him. Regardless of the proxies (both Storm and the other set used to work well, now neither are). Has there been any changes recently? Or do you need some further info to figure out what's going on?

ITs working for me. There were no changes to the harvester in the most 2 recent versions. You could go to settings >> harvester engine configuration >> import >> download default engines - and make sure you have the latest engines downloaded (be sure to back up custom engine first).

Else you could try a sample harvest of the same terms with no proxies, does it work?
 
Depends, if your only going 1 level deep then lowering the connections won't matter, if your crawling thru the site then internal pages could experience the full load of all the connections at which point lowering them would be ideal.

It has a few proxy retries, I think 3, but then it will skip the domain. With public proxies you might have 90% (or more) of them dead, slow, timeout etc.. so it could easily skip a majority of them.
thankss i think you missed one question... if you dont mind


"Also as all the domains are unique I guess there should be no prob in using my VPS's ip for email scraping at 100-200 threads? or will there by any issue?"
 
Google has all sorts of ip bans. For starters the competition finder does Not use javascript, and your browser does. So google has way more control when javascript is used, so they will let you get more done.

Next there are just literally all sorts of ip bans, I mean some queries could work while others done, operators could be banned, while regular search works, some google services could be banned while others work. You can produce this in a browser even with javascript on, if you start using complex queries. I was just literally doing research and hit the captcha just yesterday.

I have a video on it actually

Got ya. Last question, are they temporary bans and will work okay if I just increase the delay or is it a permanent thing and time to get some private proxies.
 
@loopline listened to your advice, some things helped, and somethings are changed by google it seems... The advanced queries we were using just don't work anymore, it seems that part's got nothing to do with SB... Thanks.
 
Got ya. Last question, are they temporary bans and will work okay if I just increase the delay or is it a permanent thing and time to get some private proxies.

My delay at the moment is set to 60 seconds by the way.
 
thankss i think you missed one question... if you dont mind


"Also as all the domains are unique I guess there should be no prob in using my VPS's ip for email scraping at 100-200 threads? or will there by any issue?"

No thats fine, so long as its a mix of domains and your doing only level 1. If you are crawling deeper then level 1 then you will want to turn the connections way down so that 1 domain doesn't experience all your connections at once.

Nice! Will it be a paid addon or will it be released like the 64 bit version to people who already have a licece?

ITs not an addon its an entire duplicate copy of the core scrapebox program, every addon and every plugin. That said I don't know how they will license it.

Got ya. Last question, are they temporary bans and will work okay if I just increase the delay or is it a permanent thing and time to get some private proxies.

Temp bans, in 48 hours or less they should be unbanned. You can just increase the delay and be fine probably. But wth no proxies you will need more then 60 delay.

@loopline listened to your advice, some things helped, and somethings are changed by google it seems... The advanced queries we were using just don't work anymore, it seems that part's got nothing to do with SB... Thanks.

Sounds good. Yes google changes how operators work sometimes, and discontinues some etc..
 
ITs not an addon its an entire duplicate copy of the core scrapebox program, every addon and every plugin. That said I don't know how they will license it.

Looking forward to checking it out :).
 
Yeah I need to install max on a virtual machine on my laptop and check it out at some point. I haven't even tried it yet.

A lot of people know I live in detroit, and there was a MASSIVE storm yesterday and I have been without power for 36 hours and there are still 630K people without power at this point. Point being I have a generator and my cell phone is attached to my laptop, but my response times might be slowed until this is over.

They still have no estimate for my area, and say that 90% of people will have power by Sunday and its Thursday now. So if Im slow for the next few days thats why.
 
I just upgraded to 2.0.0.85 and once I start scrapebox, I get the error: "no engines file found". when I go back to previous version, I get the same error. can you please shed some light on this? did i delete a folder by mistake?

thanks in advance
 
I just upgraded to 2.0.0.85 and once I start scrapebox, I get the error: "no engines file found". when I go back to previous version, I get the same error. can you please shed some light on this? did i delete a folder by mistake?

thanks in advance

Do you see the file engines.dat in the scrapebox configuration folder? If you do, rename it to something else and restart ScrapeBox.
 
Hi, Loopline!
I buy scrapebox amd want to buy premium plugin for article scraping and synonimizing.
The question is - will Scrapebox work with russian articles, sites and keywords? will it spin it and give me good quality of content with spintax syntax like {good|best|best ever} ?

I want to use Scrapebox for content creation for my linkbuilding software and for me needed to get from Scrapebox already spinned article in russian in spintax syntax.

thanks in advance for you reply!
 
Page scanner results are saved for all platforms everytime and not just for those selected like it previously was. Is this done on purpose or is it a bug?
 
Status
Not open for further replies.
Back
Top