Permanently Closed Marketplace Sales Thread

Status
Not open for further replies.
What does "Scrapebox can not run out of C:/path/to-my-scrapebox/scrapebox.exe" error mean ? I have this for few days and can't run my SB
 
What does "Scrapebox can not run out of C:/path/to-my-scrapebox/scrapebox.exe" error mean ? I have this for few days and can't run my SB
Because it can't run from there. There are certain locations that Scrapebox can't run from due to potential crashes or data loss. So you need to run Scrapebox from a folder on your desktop or a folder in your documents folder. Move your scrapebox folder to the desktop or documents and it will run fine.
 
I've just bought premium plugin "Expired Domain Finder" but it does not show in my Scrapebox yet. How long should I wait?
Regards.
 
Scrapebox won't do this. Scrapebox does not support authentication and generally just isn't going to work with this.

GSA SER should do what you want, "probably". I mean the forum platform would have to be compatible with GSA SER or you would have to build it in. But you could stagger account creation and posts etc... GSA SER is sold here on BHW just check the BST section S4nt0s is the guy to talk to about it.

Scrapebox could probably help you with a lot of other things and in general is an indefensible tool in internet marketing, but it won't do this specific application.
Thank you for the reply and suggestion, I'll look into SER. I do still have uses for Scrapebox as well, in time I'm sure I'll purchase a copy of that too. Thanks again!
 
I've just bought premium plugin "Expired Domain Finder" but it does not show in my Scrapebox yet. How long should I wait?
Regards.

It takes up to 12 hours. Then go to the premium plugins >> show available plugins and you should be able to download it.

Else mail support and give them your license email address and the email address you purchased the plugin from (Because if you purchased it from a different mail then the one your license is under, they may not be able to link the two together).

support (at] scrapebox [dot) com

Thank you for the reply and suggestion, I'll look into SER. I do still have uses for Scrapebox as well, in time I'm sure I'll purchase a copy of that too. Thanks again!

Your welcome, cheers!
 
I was wondering what does the TCY-D mean for yandex metric lookup?
 
Hey, I want to add my own engine for scraping. But the thing is that I can get so many results from 1 keyword, but Scrapebox seems to be limited when it comes to results/pages

It seems as if I cannot add more than 100 max pages (if that's what "Page inc" means?), and no more than 1000 in the harvester settings
 
I was wondering what does the TCY-D mean for yandex metric lookup?
In a purely simple term its yandex version of Page Rank.

The D on the end means domains. If you change it to get the metrics per url, it will have a U for Url on the end. Here is more info
http://searchengineland.com/yandex-...rank-named-thematic-index-citation-tic-251528

Hey, I want to add my own engine for scraping. But the thing is that I can get so many results from 1 keyword, but Scrapebox seems to be limited when it comes to results/pages

It seems as if I cannot add more than 100 max pages (if that's what "Page inc" means?), and no more than 1000 in the harvester settings

the inc is the increment. Like if page 2 starts at 1 or 100 or 10 etc... Different engines use different numbers to indicate the pages and also can depend on how many results per page your pulling.

Ive not seen an engine that will allow more then 1000 results. You can set the results box to 99,999 if you want, and scrapebox will attempt to crawl as deep as the engine will let it - however what the engine allows is the only limitation. If the engine allowed 50,000 results scrapebox could easily scrape that many.
 
Thanks, got it working

However, for some reason it's only using one proxy

I have 15 proxies loaded, and they're all working fine

But when I run the Detailed Harvester it's only using the proxy on the first line

Edit:

Also when you hover the results tab (where you suggested 99,999) it says "The valid range is 1 to 1000)", can you go higher than 1000?
 
Last edited:
Thanks, got it working

However, for some reason it's only using one proxy

I have 15 proxies loaded, and they're all working fine

But when I run the Detailed Harvester it's only using the proxy on the first line

Edit:

Also when you hover the results tab (where you suggested 99,999) it says "The valid range is 1 to 1000)", can you go higher than 1000?
The detailed harvester is single threaded so its only going to use 1 proxy at a time. You can see it change proxies, for example when one is blocked.

Let me rephrase, in the latest versions you can enter a number greater then 1000. For example, I mailed support about an "engien" that I build that doens't scrape a search engine. I just repurposed the custom harvester to scrape a specific big site for specific things. It works well but on this big site I may have more results for a "query" (again this is hard to explain because Its not a standard use case) then 1000 results, so I can get say 5000 etc...

I think the tool tip data is left over from 6 previous years of that box only supporting a max of 1000 results.

But all of that is mute if your using a search engine, because Ive never seen (doens't mean it doens't exist, but not sure why it would) one that lets you get more then 1000 results, many don't even let you get close to that.
 
The detailed harvester is single threaded so its only going to use 1 proxy at a time. You can see it change proxies, for example when one is blocked.

Let me rephrase, in the latest versions you can enter a number greater then 1000. For example, I mailed support about an "engien" that I build that doens't scrape a search engine. I just repurposed the custom harvester to scrape a specific big site for specific things. It works well but on this big site I may have more results for a "query" (again this is hard to explain because Its not a standard use case) then 1000 results, so I can get say 5000 etc...

I think the tool tip data is left over from 6 previous years of that box only supporting a max of 1000 results.

But all of that is mute if your using a search engine, because Ive never seen (doens't mean it doens't exist, but not sure why it would) one that lets you get more then 1000 results, many don't even let you get close to that.

Alright

Well, not exactly a search engine. More like the example you mentioned above.

Is there a way to make it work with sites that doesn't use the traditional &page= in the url, instead loading the "next page" through ajax, showing the next results below the previous ones?
 
360.jpg


It does't deal with ajax / javascript. I don't how many times Matt mentioned this over and over again.
However at least there's an alternative.
The step above which black arrow pointed is "scroll bar settings" , the right bottom has:
scroll to the top;
scroll to the bottom;
scroll line(s), how many times to scrolling(loop); the scrolling time interval;
scroll to a special element (xpath)
scroll to designated coordinate.
test button

There're also tons of methods doing waterfall (I call it waterfall) loading.
Ajax next page works in a similar way.
I don't know why they don't "do" ajax, maybe ajax is "such a bitch" in their eyes. (Guessing, perhaps not right.)
 
Last edited:
There're also tons of methods doing waterfall (I call it waterfall) loading.
Ajax next page works in a similar way.
I don't know why they don't "do" ajax, maybe ajax is "such a bitch" in their eyes. (Guessing, perhaps not right.)

Been doing it with ZP, working pretty good

The detailed harvester is single threaded so its only going to use 1 proxy at a time. You can see it change proxies, for example when one is blocked.

Isn't there a way to change proxy for every page? Isn't that the best way to go if you wanna avoid getting your proxies banned?
 
Matt please give him a shout / tell him again that scrapebox doesn't involve javascript that much,
I am fed up with it.
 
Alright

Well, not exactly a search engine. More like the example you mentioned above.

Is there a way to make it work with sites that doesn't use the traditional &page= in the url, instead loading the "next page" through ajax, showing the next results below the previous ones?

Basically no. Scrapebox uses raw sockets and threads, which don't support javascript, flash and other scripts. So if a page loads via ajax for example Scrapebox can only see what you would see in a browser if you turn scripts off.

Been doing it with ZP, working pretty good



Isn't there a way to change proxy for every page? Isn't that the best way to go if you wanna avoid getting your proxies banned?

In the custom harvester you could set the proxy change interval to 1 and it would change it after every request/page. Settings >> connections timeouts and other settings >> more harvester settings >> proxy change interval.
 
I love Scrapebox and been using it for years, but the one thing that i really dont like is the saving of scraped links.
I normally scrape very large list which takes weeks to complete ---- the problem is that while Scrapebox is scraping i cannot access the file where it is saving the links, so i have to wait untill it is finish.

I wish Scrapebox had the same option that is available in Gscraper where one can specify how many lines to save and when it reach the number, and new file is created and so one. That way one dont have to wait forever until the scrape is complete, but instead i can start working on the completed files while Scrapebox is working.

If it is allready possible to split the file being saved, then PLEASE explain how it is done
 
Status
Not open for further replies.
Back
Top