Permanently Closed Marketplace Sales Thread

Status
Not open for further replies.
Sorry if this is already built into the software, I'm on holiday and haven't used it for 2 weeks so can't remember.

Is there anyway for GSA PI to try to sniff out new footprints while it sorts through the links? I have seen other tools that do something similar but it would be handy if it could be implemented in PI.

If it presents an issue with speed then maybe a tick box could be added to enable/disable the feature depending on the users goals. Being able to pull new footprints from already scaled links would help me out a bunch in my scraping operation.

If this is not possible perhaps it could be an idea for a future program?

Cheers

Shaun
 
Can't GSA already do that? How is it different?

It can, but it can only do one sort at a time and it slows SER down slightly. This is perfect for taking your SEO to the next level. Although I previously used two VPS', one to post with my SER setup and one to scrape and identify I now plan to try it one VPS using ScrapeBandit to scrape on its own servers, dump the links into my Dropbox, have PI on my SER VPS checking my Dropbox folder for new links and then identifying and sorting them, once complete it will then put them into a new folder that my SER verify projects use and then SER will post and identify. All the while SER is left to do what it does best when used correctly.....get you rankings.
 
Sorry if this is already built into the software, I'm on holiday and haven't used it for 2 weeks so can't remember.

Is there anyway for GSA PI to try to sniff out new footprints while it sorts through the links? I have seen other tools that do something similar but it would be handy if it could be implemented in PI.

If it presents an issue with speed then maybe a tick box could be added to enable/disable the feature depending on the users goals. Being able to pull new footprints from already scaled links would help me out a bunch in my scraping operation.

If this is not possible perhaps it could be an idea for a future program?

Cheers

Shaun

Well someone else has asked me for that feature before, but it would require quite a bit of work. Currently, Pi is just checking for certain lines when detecting an engine, but if it needed to come up with footprints it would have to grab everything from the page, compare it between all imported URLS, detect the frequency, then you would need another window to manually review the footprints and check the results count against Google, etc. I mean its kinda out of the scope of what this software was designed for, but I think it would be pretty cool to make another stand alone tool for it.

Sven already has part of the code for it in footprint studio, but we're using 2 different versions of delphi so its not a simple transfer of code. A stand alone tool would actually be easier to make for that.

The footprint tool I've been using is called Footprint Factory. It can gather a lot of footrprints, but can be a little buggy at times. It does get the job done though. Maybe we can work on that footprint tool after the next tool that's coming :)
 
so I bought GSA platform identifier some days ago, I'm not fully satisfied with it since its eating lots of CPU and taking huge amount of time to complete a list of 80,000 URLs(five hours).
Although GSA SER is a bit slower than platform identifier but it does not eats up so much amount of RAM and CPU like this does. Yes you can run it independently but the whole point of scraping lists somewhat down sizes with the additional cost and not much additions in the new version.
Yes I know Santos and his team are really good programmers and they will find a solution sooner than later for bugs in this program but as of now it is really taking much time and CPU to process lists.
Now here are my questions: ?
question one ? how can I reduce CPU usage which is showing as 99% or higher on my VPS?
question two ? how can a process lists faster with CPU usage as low as possible?

My settings are here: ?
50 threads, one MBPS per second bandwidth restriction, two projects in one time.
 
so I bought GSA platform identifier some days ago, I'm not fully satisfied with it since its eating lots of CPU and taking huge amount of time to complete a list of 80,000 URLs(five hours).
Although GSA SER is a bit slower than platform identifier but it does not eats up so much amount of RAM and CPU like this does. Yes you can run it independently but the whole point of scraping lists somewhat down sizes with the additional cost and not much additions in the new version.
Yes I know Santos and his team are really good programmers and they will find a solution sooner than later for bugs in this program but as of now it is really taking much time and CPU to process lists.
Now here are my questions: -
question one - how can I reduce CPU usage which is showing as 99% or higher on my VPS?
question two - how can a process lists faster with CPU usage as low as possible?

My settings are here: -
50 threads, one MBPS per second bandwidth restriction, two projects in one time.

I responded to this post on the support forum as well, I see its the exact same post but I'll respond here too.

Are you on a very low end machine? On my end, 50 threads and 1mbps bandwidth limit would barely use any CPU. I'm running 450 threads with 10mbps and its not breaking 50% cpu at the moment:

6A8Us0D


If you want, I'll login and take a look at your VPS to see what's going on.
 
Well someone else has asked me for that feature before, but it would require quite a bit of work. Currently, Pi is just checking for certain lines when detecting an engine, but if it needed to come up with footprints it would have to grab everything from the page, compare it between all imported URLS, detect the frequency, then you would need another window to manually review the footprints and check the results count against Google, etc. I mean its kinda out of the scope of what this software was designed for, but I think it would be pretty cool to make another stand alone tool for it.

Sven already has part of the code for it in footprint studio, but we're using 2 different versions of delphi so its not a simple transfer of code. A stand alone tool would actually be easier to make for that.

The footprint tool I've been using is called Footprint Factory. It can gather a lot of footrprints, but can be a little buggy at times. It does get the job done though. Maybe we can work on that footprint tool after the next tool that's coming :)

Yea you messaged me about the next tool, looking forward to beta testing it for you :).

I have used footprint factory and like you said it is buggy and I would happily pay for a competitor with the GSA logo as I know it will be upto scratch like the rest of your line.
 
Yea you messaged me about the next tool, looking forward to beta testing it for you :).

I have used footprint factory and like you said it is buggy and I would happily pay for a competitor with the GSA logo as I know it will be upto scratch like the rest of your line.

Haha, that tool I mentioned to you before is a different project, there's another one too lol. But ya, I'll talk to Sven about that footprint tool, a lot of the code is already there, would just need to make it a stand alone tool. I'll see what can be done :)
 
My Review Now:

Nice Fast Tool,If you have a good server you can parse millions of URLS in minutes.
 
Haha, that tool I mentioned to you before is a different project, there's another one too lol. But ya, I'll talk to Sven about that footprint tool, a lot of the code is already there, would just need to make it a stand alone tool. I'll see what can be done :)

Oh my, oh my! I'm still very excited about Platform Identifier and now another tool? Please s4nt0s, spare us from guessing I won't be able to sleep tonight! :) Can you PM me a hint at least?

I was the other guy who suggested footprint harvesting into PI to s4nt0s, but then I changed my mind about this idea as well. That really is better to be a separate software. PI is doing the job it was made for very well, period. What is really important for the upcoming GSA Footprint Studio is to keep it well integrated with PI and/or SER.

About PI eating a lot of CPU, yes that is true, but I don't find that it's any bit more than GSA SER identifying a list and it's completely worth it. In fact if doing it in SER the posting part of the software has to be stopped, and identifying can last for days or weeks without you being able to post links so that certainly approves the existence of PI. I had a 2 core Xeon E3 with 2GB ram VPS, and the max I could get from PI was: 3 projects with total of 160 threads at 10Mbps and even one of the projects was set to Deep matching(which eats a lot more CPU than a regular project) AND on top of that I was running scrapebox in parallel at 5000 threads(but with a slower speed because of a lot of bad public proxies). The CPU was maxed out all the time, but it was working for weeks without any issues. Then I wanted to challenge that machine more so I went and installed Proxy Multiply. I left that scraping for new proxies 24/7, but the poor VPS started going crazy on me so I had to lower the speed on PI to like 3-4Mbps and the threads to less than 100. Still everything was going smoothly. :) Now I'm moving to a new dedicated server because I saw the potential of this scrapebox/PI combo setup and want to scale things up.

I even think that PI and Scrapebox complement each other pretty well resource-wise. PI is taking up the CPU, Scrapebox is taking up the RAM and there you go, you're frying up your VPS to the max! :) Just don't forget to tell your hosting provider to keep the firefighter brigade on fast dial :D
 
Last edited:
Haha, that tool I mentioned to you before is a different project, there's another one too lol. But ya, I'll talk to Sven about that footprint tool, a lot of the code is already there, would just need to make it a stand alone tool. I'll see what can be done :)

Mate, you guys really are cleaning up the market and doing it perfectly!

Back on the subject of PI.....If I have PI writing to a folder and SER posting from it does SER know to only post to the new URLs within the identified folder or is there a way to get SER to remove the url from the identified folder once it has been used?

Cheers

Shaun
 
My Review Now:

Nice Fast Tool,If you have a good server you can parse millions of URLS in minutes.

Mate, you guys really are cleaning up the market and doing it perfectly!

Back on the subject of PI.....If I have PI writing to a folder and SER posting from it does SER know to only post to the new URLs within the identified folder or is there a way to get SER to remove the url from the identified folder once it has been used?

Cheers

Shaun

Hi Shaun, SER doesn't delete the link after it has been used, but it does know which URL's its posted to before so no worries there.
 
When I set PI to save identified lists to a destination folder will each new project I create that uses the same destination folder add to that list? I want to make sure the "list" in that folder will grow and not be overwritten by the next project.

Also what's the best way to import the identified URL's to the global list on SER? Also just want to make sure this adds to the list and doesn't overwrite it.
 
Last edited:
When I set PI to save identified lists to a destination folder will each new project I create that uses the same destination folder add to that list? I want to make sure the "list" in that folder will grow and not be overwritten by the next project.

Also what's the best way to import the identified URL's to the global list on SER? Also just want to make sure this adds to the list and doesn't overwrite it.

Yes, each new project you create will just append URLs to the lists and never overwrite previous URL's. For importing identified URLs into the global site lists you can just right click on the project > export to .SL , then go into SER > options > advanced > tools > import site lists > identified.

Keep in mind you can also use the destination folder you set in Pi as the global site lists folder in SER or you can import the .SL directly into a project by right clicking on it > import target lists > from file.
 
Yes, each new project you create will just append URLs to the lists and never overwrite previous URL's. For importing identified URLs into the global site lists you can just right click on the project > export to .SL , then go into SER > options > advanced > tools > import site lists > identified.

Keep in mind you can also use the destination folder you set in Pi as the global site lists folder in SER or you can import the .SL directly into a project by right clicking on it > import target lists > from file.

Awesome thanks. Didn't think about setting PI destination folder as global in SER.
 
I am happy with it as well. I scanned over 300 thousand sites in day and half on my home computer, so that include stopping, etc

Now I am ready to my vps in 6 hours, have to wait for my home to finish.

The vps has 120gb of sites. I bought gscraper and they let you try their proxies for like 5 days. It was christmas time so i was busy at work, so i did not have chance to sort it it. Then I been slowly hacking away at it, with ser. But with ser you can only do one at time. With pi you can just add all at once.

so if you constantly scrape and do not want to babysit ser for indentified get this tool.
You can add all your list of sites
put deep matching on for better detection ratio

my next question is some projects I get 200-400 or more average. Other projects i get average less than 100. I am thinking because the sites it is scanning. I have the better detection ratio off and same thread as the others.

if you need me beta test something let me know

also interested in the footprint.

may i also suggest to add a virus scanner in tools. I know scrapebox has it, but it would be awesome addition
 
Last edited:
I have SB harvesting and dumping into a folder that PI is monitoring - then I have SER running using the PI identified folder. I now know PI will amend new links to the identified files but will SER recognize new links that have been added to that file? For example SER runs through an entire identified list that PI created. Then PI runs again sending more URL's to that folder, will SER see that and start up again?

In SER under a project if I select "use urls for global site list if enabled" > "identified" and have the correct folder set under global settings do I still need to tick the identified box on global settings so SER pulls URL's from that PI identified folder?
 
The vps has 120gb of sites. I bought gscraper and they let you try their proxies for like 5 days. It was christmas time so i was busy at work, so i did not have chance to sort it it. Then I been slowly hacking away at it, with ser. But with ser you can only do one at time. With pi you can just add all at once.

so if you constantly scrape and do not want to babysit ser for indentified get this tool.
You can add all your list of sites
put deep matching on for better detection ratio

here is an update

the 129 gb file was 828 million links. I remove duplicate domains which = to 619,000.

PI says it would be done in 5 hours. Remember I have been working this off and on. since Christmas.

Very little babysitting
 
I am happy with it as well. I scanned over 300 thousand sites in day and half on my home computer, so that include stopping, etc

Now I am ready to my vps in 6 hours, have to wait for my home to finish.

The vps has 120gb of sites. I bought gscraper and they let you try their proxies for like 5 days. It was christmas time so i was busy at work, so i did not have chance to sort it it. Then I been slowly hacking away at it, with ser. But with ser you can only do one at time. With pi you can just add all at once.

so if you constantly scrape and do not want to babysit ser for indentified get this tool.
You can add all your list of sites
put deep matching on for better detection ratio

my next question is some projects I get 200-400 or more average. Other projects i get average less than 100. I am thinking because the sites it is scanning. I have the better detection ratio off and same thread as the others.

if you need me beta test something let me know

also interested in the footprint.

may i also suggest to add a virus scanner in tools. I know scrapebox has it, but it would be awesome addition

Thanks for the suggestions. I'm not sure what you mean 200-400 or more by averages? 200-400 more what exactly? detected URLS? URLs/min? Can you clarify please? :)

I have SB harvesting and dumping into a folder that PI is monitoring - then I have SER running using the PI identified folder. I now know PI will amend new links to the identified files but will SER recognize new links that have been added to that file? For example SER runs through an entire identified list that PI created. Then PI runs again sending more URL's to that folder, will SER see that and start up again?

In SER under a project if I select "use urls for global site list if enabled" > "identified" and have the correct folder set under global settings do I still need to tick the identified box on global settings so SER pulls URL's from that PI identified folder?

You want to make sure you're using automator plugin to dump SB files into a folder because SB locks the files that its currently writing to so until that file finishes, Pi won't be able to read it. That's why its best to use automator to have SB save x amount of URL's to a .txt, then start on a new one.

Yes, SER will automatically recognize new links. It kinda works the same as PI, it will check that identified folder in intervals to look for newly added URL's.

Actually, you can just have the global site list set in the advanced options. You don't have to have it "checked" in the advanced options, as long as its pointing to the proper folder. If you use check boxes next to the selected folders in the advanced menu, it will automatically add identified URL's into that folder. That means from any project, all the identified URLS will be added there. So if you only want specific URL's to stay in the identified folder, you might want to leave that unchecked if using multiple projects.

To get the project to actually run off the global site lists, in the project options you would select, "use urls from global site lists if enabled" and choose the corresponding folders you want to use.

Also in the project options, you might want to go to the search engine selection box > right click > check none. This way SER isn't going to try and scrape from the search engines, it will only use your identified list.

here is an update

the 129 gb file was 828 million links. I remove duplicate domains which = to 619,000.

PI says it would be done in 5 hours. Remember I have been working this off and on. since Christmas.

Very little babysitting

Wow a 129 GB file to dedup? lol, that's crazy. Hope it finishes :)
 
Last edited:
Is it possiblle to auto remove duplicate domain when identified a platform or create a new project to removed it? Is it good to have many urls with same domain? since I've integrated in my GSA global lists and the lists become bigger but fill in with many duplicate domain so I have manually dedupe it in GSA.
 
Status
Not open for further replies.
Back
Top