- Jan 25, 2009
- 6,378
- 3,774
Your transaction ID for this payment is: xxxxxxxxxxxx.
I would recommend you edit that and remove your transaction Id, its part of your license info.
Its hard to make large harvests with Scrapebox...
I started yesterday 4 instances of Scrapebox, each instance 1k threads, each instance 8k urls per second.
After ~5 hours in harvester session folder i had 4 files, 20gb each, 80gb in total. This give 380gb results per day, to finish harvesting i need 5 days so 1.9TB space needed... but hdd have only 1TB.
Problem:
No option in SB to split results when harvesting in multiple files.
How to fix problem:
Add option in SB to create new file if actual file with results have more lines than X. Gscraper already have this function and im really missing this function in SB.
With above we can use external program (for example GSA PI) to remove duplicates while harvesting. Right now its not possible because GSA PI cannot acces txt file when SB is writing results to this file.
I fail to understand your problem. Your complaining that your hard drive isn't big enough to store the files and your solution is that you want the files to be split into smaller files? That won't make your hard drive bigger.
Now if your saying that the issue is that all runs will complete in 5 days and that you aren't stopping them in the mean time, I guess I see your point, but are you really going to keep that going all month long? Because its not going to solve your problem long term.
You could just as easily use the automator and break up your lists of keywords into 5 parts, that way it will finish each day, then you can have the automator remove duplicate urls and save that days files. Then start over, all automated without having to bring a 2nd program into it.
Or you could just do 1 or 2 harvest runs and when they are done remove dupes, and then start the next 2 runs and run them for 5 days. Because unless your running this non stop all month long it isn't going to matter right?
Anyway, there are ways you could work around it as it stands right now.
