Tool to get all URLs from root domain (large list)

alternateswordDo you have a working version? Can you check and post a screenshot if it really works?
 
Hy.

First of all I will also be interested in a tool like that.

Meanwhile I'm using link extractor from SB.

I'm talking about this:

1. Run a harvest using the "site:"on G
2. Load the resulting list link extractor, select "internal" so I can get only the internal links.
3. I load the resulting list in SB and remove all the links from others domains beside the one I need
4. I load the resulting list back to link extractor and I run it again.

Do this several times and you will get "many links" I say many links since this will not guarantee that you will get all the url from the domain but it can get you some nice results.

Thank you and waiting for that tool :P
 
Hy.

First of all I will also be interested in a tool like that.

Meanwhile I'm using link extractor from SB.

I'm talking about this:

1. Run a harvest using the "site:"on G
2. Load the resulting list link extractor, select "internal" so I can get only the internal links.
3. I load the resulting list in SB and remove all the links from others domains beside the one I need
4. I load the resulting list back to link extractor and I run it again.

Do this several times and you will get "many links" I say many links since this will not guarantee that you will get all the url from the domain but it can get you some nice results.

Thank you and waiting for that tool :P

Sorry that I'm asking. But what is SB and where can I find it?
...and in most cases you find these intresting websites behind the checkout (clickbank and others).
They are often not protected. So if you know the correct link you can visit this page directly.
 
Last edited:
Been using successfully XML - Sitemap (the site) ... just make sure you add a sleep or some kind of delay when crawling the site as it might have a negative impact on your server, depending on the number of threads that are used to crawl the site.

Thomas
 
It is strange that we cant give the name of a free or cheap client-side tool for this purpose. Did i miss any?
If your solution works on the client side and cheaper than $30, let me know. I just want to give the root url and want to get list of all pages crawled starting from the root url.
 
I would pay even more than $30 for that. But obviously, there is no such tool. Why it is so difficult? I got some results with DirBuster but it takes really long. But how does the G bot crawls a website? It finds everything, doesn't it?
 
I would pay even more than $30 for that. But obviously, there is no such tool. Why it is so difficult? I got some results with DirBuster but it takes really long. But how does the G bot crawls a website? It finds everything, doesn't it?

It depends on the product. I pay more. But only for this feature, it is enough.

There are crawling libraries for widely used languages. You can program your own crawler. If i cant find a product, i will program my own crawler using crawler4j. A console version will not take more than one day..

http://code.google.com/p/crawler4j/
http://java-source.net/open-source/crawlers
 
What trouble are you having with Gsitecrawler? Google allows 50k urls in their sitemaps. Gsitecrawler does 50k, splits the sitemaps automatically, and throws the rest in the aborted url list. You should be able to pick up all urls with it.
 
haven't read the previous replies cause its so long :P Only read the first few post and to answer the OP's question SB can do it. You don't need any addons to it the main scrapebox gui can do this for you.

I know cause I'm doing this to harvest few info's on yellow pages before. so I need to gather all the pages of yellowpage via scrapebox before I can harvest the telephone numbers of the companies :p

Thanks
 
Sorry that I'm asking. But what is SB and where can I find it?
...and in most cases you find these intresting websites behind the checkout (clickbank and others).
They are often not protected. So if you know the correct link you can visit this page directly.
He mention Scrapebox as SB. There is a free addon available for scrapebox called link extractor.
 
He mention Scrapebox as SB. There is a free addon available for scrapebox called link extractor.

Thank you.

There are so many posts now. But I still have no clue how to find all websites, expecially those behind the checkout site.
For example: w!w.fastforexresult.c!m The member's site is not password protected. If you know the correct link you can visit this site without passing the checkout. You can advise me how to use a crawler/spider/tool to find this page?
 
Last edited:
Back
Top