[/CENTER]
That is basically it, it crashes when the harvest stops when i get to those numbers (thought it was the norm), and i have to go look for the batch files like you correctly state. I'll just contact support with some screenshots next time it happens.
Either way i dunno if my sugestion came out clear, if the program automatically (there really is no need to have it as a check options) removed duplicate URLs (that are just trash really) immediatly after finishing the harvest (before populating the grid, and writing the multiple files) it would save up some resources, eventually put more "important" lines on the limited grid in the interface and avoid the mandatory step to remove duplicate url by hand after each harvest.
Sure, I get what your saying. It makes sense that that would be convenient, yes. However even the remove duplicate domains writes the urls to the grid, THEN removes duplicate domains. Its just auto calling the remove dupe domains, its not actually removing the dupe domains prior to writing them to the grid. So my guess would be that if they would have to do some serious reworking of the harvester to get it to remove dupe urls in real time, prior to adding them to the grid.
However, what I do as a work around with the automator is this. I figure up about "loosely" how many urls it would take to hit 1 million. For instance, I loaded in a chunk of 5000 keywords. In "theory" you can get up to 1000 results for each keyword. So 1000 x 5000 = 5 million. Now as a general rule, I don't often see 1000 results for all keywords, especially not for 5000 keywords/footprints. So as a rule dealing with 5000 keywords will net you less then the 1 million urls, then you have less then 1 million and then you can save or post and save successful or whtaever you are doing.
Then I just merge in the same job, so you get all teh same results, but then I go change the file that the keywords section is linked to. Then just let the automator run and your working with chunks of keywords that are less then the 1 million grid and its all good.
In fact I stage it in a fixed way. Meaning:
I make a folder for my keywords and if I do 5 sets of jobs, like sweetfunny mentioned above, then I would need 5 keyword files.
1
2
3
4
5
So I just make those 5 files and name them something like above 1-5 or whatever.
Then when Im done with a run. I just replace those 5 files with 5 new keyword files and name those file the EXACT same thing, like 1-5
Then just run the job again. no need to even change anything in the job, because it runs off of a mapped file and you just replaced those files. I talk about it in the automator vid.
Of course if you are like posting and exporting successful, you will want to grab those successful files and save them off elsewhere, cause they will be overwritten with running the new job.
but I have a nice folder tree of files for the jobs I run and I just copy out successful and found files and dump them in a different folder after each run and then replace keyword files and keep moving, makes easy work of it.