I have a collection of PDF files, that are essentially how-to guides and workshop manuals that have been scanned. I'm thinking these could provide a nice little income if placed online and monertized correctly.
My plan is to convert each one into text, and then use these for blog posts, unless anybody can suggest a better system for publishing the resulting data??
However my main problem is some of the files are getting on for 10Mb, which most online OCR I've tried them on seem to baulk at. Can anybody recommend a decent OCR package (preferably online as I run Linux) that can cope with images and diagrams within the text along with large files with 100s of pages. ??
Thanks
My plan is to convert each one into text, and then use these for blog posts, unless anybody can suggest a better system for publishing the resulting data??
However my main problem is some of the files are getting on for 10Mb, which most online OCR I've tried them on seem to baulk at. Can anybody recommend a decent OCR package (preferably online as I run Linux) that can cope with images and diagrams within the text along with large files with 100s of pages. ??
Thanks