Permanently Closed Marketplace Sales Thread

Status
Not open for further replies.
Your transaction ID for this payment is: xxxxxxxxxxxx.

I would recommend you edit that and remove your transaction Id, its part of your license info.


Its hard to make large harvests with Scrapebox...

I started yesterday 4 instances of Scrapebox, each instance 1k threads, each instance 8k urls per second.
After ~5 hours in harvester session folder i had 4 files, 20gb each, 80gb in total. This give 380gb results per day, to finish harvesting i need 5 days so 1.9TB space needed... but hdd have only 1TB.


Problem:
No option in SB to split results when harvesting in multiple files.


How to fix problem:
Add option in SB to create new file if actual file with results have more lines than X. Gscraper already have this function and im really missing this function in SB.

With above we can use external program (for example GSA PI) to remove duplicates while harvesting. Right now its not possible because GSA PI cannot acces txt file when SB is writing results to this file.

I fail to understand your problem. Your complaining that your hard drive isn't big enough to store the files and your solution is that you want the files to be split into smaller files? That won't make your hard drive bigger.


Now if your saying that the issue is that all runs will complete in 5 days and that you aren't stopping them in the mean time, I guess I see your point, but are you really going to keep that going all month long? Because its not going to solve your problem long term.

You could just as easily use the automator and break up your lists of keywords into 5 parts, that way it will finish each day, then you can have the automator remove duplicate urls and save that days files. Then start over, all automated without having to bring a 2nd program into it.

Or you could just do 1 or 2 harvest runs and when they are done remove dupes, and then start the next 2 runs and run them for 5 days. Because unless your running this non stop all month long it isn't going to matter right?

Anyway, there are ways you could work around it as it stands right now.
 
I'm having a weird issue with the email grabber feature in V2. When I have a bigger list of URLs, let's say 100,000 or more and I run the email grabber, it runs fine for the first couple of thousand URLs but then the threads seem to get stuck on 'processing' some random URLs (without yielding any emails, just says processing). I tried using less connections for the grabber, with and without proxies, it's still like that. Any ideas what I could try to resolve this?
 
Delay|Filter|Agents:
R-M-Stuck.PNG
 
I'm having a weird issue with the email grabber feature in V2. When I have a bigger list of URLs, let's say 100,000 or more and I run the email grabber, it runs fine for the first couple of thousand URLs but then the threads seem to get stuck on 'processing' some random URLs (without yielding any emails, just says processing). I tried using less connections for the grabber, with and without proxies, it's still like that. Any ideas what I could try to resolve this?

Are you extracting urls all from the same domain?

Make sure you have Scrapebox set as trusted/allowed in all security software, such as antivirus etc... As it could be locking the threads if it doesn't like the content on the page or the url its self. Also you can try closing down any unneeded programs, to make sure they are not inteferring.

Another thing you can try is if you can see the urls its getting locked up on, pull those urls out of your list and put a few of them in a list by its self and try it, do they get locked up or do they process fine?


~~~~~~~~~~~~~~~~

New Video


Scrape Links By Crawling A Site
 
Last edited by a moderator:
The skip for identification was most likely the culprit. I mean SER does visit all urls that it is fed. I actually have found that SER works "better" (faster) if you leave that unchecked (which seems backwards), but if its checked then its using your ip to hammer away on all urls.

Is this from your ISP or from your server company? I mean did the complaint come to your ip or you server?

If its your IP you can play the unsecured wifi card, or malware or whatever.

If its your server company, don't sweat it, keep moving, you can always just ask them to change your IP or worst case switch to a new company. I can say I have not ever received such a complaint on any server and I run GSA and Scrapebox on a LOT of servers 24.7

But I did run Xrumer wide open for a few months on a server and they eventually told me that they were going to cancel it, so complains are possible of course but server companies are a dime a dozen. It also may take up some of their time, but as long as they say they cancelled the customer it doesn't hurt them.

It probably was the identification that did this. I unchecked it and never received a complaint again.

It's my server company that sent the complaint, they are France based. They basically told me: there is a complaint for your ip, here is what the complaint says, talk to them to work it out, and that was it. But I'm sure that they won't be comfortable if they receive any more complaints, so if there is any way to stop them it would be great.

=======================

I have a suggestion for the "Google Competition Finder" addon for SC v2. After the check is completed and the results are displayed, normally there will be some keywords that won't pass a result(for whatever reason). They are usually marked with a ?. So if I want to do a re-check just to those keywords without losing the data from the previous keywords I can't. If I click Re-Check it will start all over again checking everything and then another set of keywords will be marked with ?. So I can't get precise results with 2-3 runs if I want to.

So can you please add the option to recheck ONLY the ones that didn't return a proper result?
 
Export and sort? So do I recheck the ones that returned error or question mark?

Well you could recheck the error ones of course as that was probably a blocked proxy when it tried.

The ? ones could be no results, so rechecking isn't going to help, or it could be another issue.

If you take a dozen of the question mark ones and recheck them a few times do they consistently yield just question marks or do they work?
 
ScrapeBox v2.0.0.42 Final Released

  • Added 2 more keyword scraper sources (Wikipedia.org and Alibaba.com)
  • Fixed a bug in "Export And Split" function
  • Fixed issues with stalling proxies in Harvester
  • Added auto abort harvester when threads go below a set minimum
  • Fixed a bug in Proxy Manager related to the autosave proxies feature
  • Fixed an issue with non-ip proxies
  • Fixed a bug in Export and Split
  • Fixed bug in Country Filter of Proxy manager
  • Changed the Proxy Manager Autosave feature which caused issues
  • Fixed a bug related to footprints and keywords in custom harvester
  • Fixed a bug related to footprints and keywords in detailed harvester
  • Fixed a bug in detailed harvester related to keyword encoding
  • Fixed a bug in meta grabber
  • New http://www.scrapebox.com website now live

With the bug reports low and the features in ScrapeBox v2.0 complete enough to make it more functional and perform better then ScrapeBox v1.x we have decided to drop the beta label and release ScrapeBox v2.0 final. Of course the word "final" doesn't mean much, we have been updating ScrapeBox since 14th October 2009 with hundreds of new features, addons and enhancements and this isn't going to change it will continue to get better and evolve. We have also given the ScrapeBox.com website a refresh as well to coincide with ScrapeBox v2.

Thanks to everyone who reported bugs and helped out during the beta testing it's much appreciated.
 
What do you reckon to be the best method to get page value before commenting? In the past I've used Pagerank but it's rarely updated. Thinking about going the Moz route to get page authority but it will probably take forever with the free API.
 
What do you reckon to be the best method to get page value before commenting? In the past I've used Pagerank but it's rarely updated. Thinking about going the Moz route to get page authority but it will probably take forever with the free API.

Unless you want to pay for moz or majestic data then the free moz is probably the best option. You can create a bunch of free moz accounts and tie each to a dedicated private proxy and it drastically speeds things up.
 
Hey bud, would I be able to scrape Web 2.0s, check DA/PA and also check if they're still available?
 
2.0 is great ........ thanks sweetfunny

ScrapeBox v2.0 is great. Thanks sweetfunny.

Thanks yes ScrapeBox v2.0 is way better then v1. Also a big thanks to Softtouch and Loopline too.

Hey bud, would I be able to scrape Web 2.0s, check DA/PA and also check if they're still available?

Sure can, the Vanity Checker is fully trainable to work with any Web 2.0 platform and check if the profiles or pages are available or not. http://www.scrapebox.com/vanity-name-checker

And the Page Authority Addon can give you the DA and PA of the URL's http://www.scrapebox.com/page-authority-addon
 
Hello :)
I have a problem guys, i have a bunch of proxies when i scrape google everything is good after 100 000 link it gives me just errors. Proxies not getting ban but i think it requires chacpa enterence. How can i solve that ? Can i enter manually or can i use CB ? Many thanks.
 
Hello :)
I have a problem guys, i have a bunch of proxies when i scrape google everything is good after 100 000 link it gives me just errors. Proxies not getting ban but i think it requires chacpa enterence. How can i solve that ? Can i enter manually or can i use CB ? Many thanks.
Maintain the number of threads/proxies ≈ 1/5~1/4 (1/3 ?)
It's a "slow" process, but not easily got banned, etc.
Ensure lots of stuff could be scraped.
 
Last edited:
Status
Not open for further replies.
Back
Top