cottonwolf
Regular Member
- Jan 20, 2015
- 469
- 242
I've scraped about 100gb of urls the last week or two and I've been two lazy to clean them and work with them in gsa and now it's a pain the cave to do so.
What I've done is that I've used the scrapebox harvester session folder files, I didn't bother to wait for the harvester to load gigs of urls after a scrape, and put these large files into dupe remove addon.
However, the duperemove addon, according to me, often messes up the files encoding somehow. It's probably me being stupid and inexperienced.
edit:I used the addon to separate an 8gb file to files of 5 million urls. And I've got 16 parts back.
I don't think that the dup url and dup domain remover addon would work with these messed up part files. /edit
A file containing these lines:
http://website.com
http://domain.com
often becomes
h t t p : / / w e b s i t e . c o m
h t t p : / / d o m a i n . c o m
What can I do to sort these kind of files? GSA SER can't process such a url, obviously.
I've got no idea what file encoding I should use. I think scrapebox saves harvest as unicode. These files take up a huge amount of space. Then when I'm done cleaning I export my files either as utf8 or lately as ANSI and load them to ser. These encodings don't mean much to me.
Thanks!
Edit: I don't even SEO
What I've done is that I've used the scrapebox harvester session folder files, I didn't bother to wait for the harvester to load gigs of urls after a scrape, and put these large files into dupe remove addon.
However, the duperemove addon, according to me, often messes up the files encoding somehow. It's probably me being stupid and inexperienced.
edit:I used the addon to separate an 8gb file to files of 5 million urls. And I've got 16 parts back.
I don't think that the dup url and dup domain remover addon would work with these messed up part files. /edit
A file containing these lines:
http://website.com
http://domain.com
often becomes
h t t p : / / w e b s i t e . c o m
h t t p : / / d o m a i n . c o m
What can I do to sort these kind of files? GSA SER can't process such a url, obviously.
I've got no idea what file encoding I should use. I think scrapebox saves harvest as unicode. These files take up a huge amount of space. Then when I'm done cleaning I export my files either as utf8 or lately as ANSI and load them to ser. These encodings don't mean much to me.
Thanks!
Edit: I don't even SEO
Last edited: