@luccha, but it won't go through the whole site and list all the URLs for all the sites. If you load a list of URLs, it just gets data for one page.
@fred, And SB link checker? It does have its uses but I suppose you are suggesting checking for links (internal only) on the page, but the problem with that is that it won't spider the whole site. Unless EVERY page is linked from the home page, it won't return much. I tried this and I wasn't getting many URLs at all. It's a nice idea, but in practise, it doesn't work.
@claymc, yes, I actually have 2 lists of 100,000 URLs. So 200k root domains total. I need to spider all these, extract all the URLs, and then manually filter in notepad++ to hopefully filter out a lot of unpostable pages. I would probably end up with around 200,000,000 URLs in one text file, so I would have no choice but to split it into about 50 parts and edit each one individually. (ouch, but can be done in a couple days)
Since there seems to be no solution for what at first to me appeared to be a simple problem, I have just paid someone on vworker to build me a crawler. More than I was hoping to pay but I need this.