Permanently Closed Marketplace Sales Thread

Status
Not open for further replies.
In the keywords part, I want to suggest a new feature. In the option "remove line containing/non containing" , one can add only 1 word at a time. I want to add multiple words . Gscraper has the option but the problem in there is that, it needs to be links only. In scrapebox, which is mostly used for research purpose too, i think this feature will be very useful

Added to keyword area -> Button More -> Remove Keywords Containing and Remove Keywords not Containing. You can enter multiple keywords separated by the vline/pipe character (|). Will be available in the next beta update.
 
I have 6000+ port scanned public proxies which are replaced every 2-3 hours. I set scrapebox to 5000 threads for harvesting. I'm rotating multiple harvesting sessions in a loop with the automator each for different engine.

The harvesting sessions go very well at the start with 5000 threads and I get to 340 urls/s which is ok for public proxies. After a while, somewhere around 40-60% of the session progress the threads start to decrease slowly. The session ends when the last thread has finished. I'm aware that all threads cannot finish at the same time so I can't keep at 5000 at all times, but the decreasing lasts for way too long. For example a harvesting session lasts for 8 hours and only in the first 2 hours it uses all of the 5000 threads. The last 6 hours the threads start to decrease and in the last 2-3 hours it's using below 100 threads. This seems like an awful waste of time to me. I want to loop different and shorter harvesting sessions so I can get fresher links much regularly, but the shorter I make the sessions, the more inefficient they will become, scraping with less threads most of the time. If I make the sessions longer I will have to wait more to get the scraped urls(since scrapebox locks the file it writes to and doesn't let you use the harvested urls until it finishes a session).

Same thing is happening with 100, 200 and 5000 threads. More than half of the time it harvests is spent with less and less than the max threads.

If only there was a way to start the next session while the first session still lasts so it can use the threads as they become available from the first session. This way all 5000 threads will be used to the max at all times.

Is there any way or a workaround currently to solve this?

EDIT: Or if there was an option in the automator to "Stop the harvesting session after the threads go below xxxx" so it can continue to the next harvesting job, this can solve it easily.
 
Last edited:
I have 6000+ port scanned public proxies which are replaced every 2-3 hours. I set scrapebox to 5000 threads for harvesting. I'm rotating multiple harvesting sessions in a loop with the automator each for different engine.

The harvesting sessions go very well at the start with 5000 threads and I get to 340 urls/s which is ok for public proxies. After a while, somewhere around 40-60% of the session progress the threads start to decrease slowly. The session ends when the last thread has finished. I'm aware that all threads cannot finish at the same time so I can't keep at 5000 at all times, but the decreasing lasts for way too long. For example a harvesting session lasts for 8 hours and only in the first 2 hours it uses all of the 5000 threads. The last 6 hours the threads start to decrease and in the last 2-3 hours it's using below 100 threads. This seems like an awful waste of time to me. I want to loop different and shorter harvesting sessions so I can get fresher links much regularly, but the shorter I make the sessions, the more inefficient they will become, scraping with less threads most of the time. If I make the sessions longer I will have to wait more to get the scraped urls(since scrapebox locks the file it writes to and doesn't let you use the harvested urls until it finishes a session).

Same thing is happening with 100, 200 and 5000 threads. More than half of the time it harvests is spent with less and less than the max threads.

If only there was a way to start the next session while the first session still lasts so it can use the threads as they become available from the first session. This way all 5000 threads will be used to the max at all times.

Is there any way or a workaround currently to solve this?

EDIT: Or if there was an option in the automator to "Stop the harvesting session after the threads go below xxxx" so it can continue to the next harvesting job, this can solve it easily.

Well you can run multiple instances, so you could setup a second Scrapebox instance to run and have the first item be a 2 hour delay for example and then when you first job runs have it call the 2nd job to start before it starts harvesting.

But I would take a guess in that your burning the proxies so quickly that Scrapebox is only running 100 or less threads because its spending so long cycling thru bad proxies for the current threads it can't get more threads started.

On the general whole if you have 6000 proxies and your only getting 340 urls/s I would question the quality of the proxies. I mean Ive gotten 2000 urls/s just with less then 1000 public proxies that I scraped, not even port scanned. I would say you proxies are at the root of the issue.

Have you tried shorter runs? If you go a lot shorter it might be fine.
 
Well you can run multiple instances, so you could setup a second Scrapebox instance to run and have the first item be a 2 hour delay for example and then when you first job runs have it call the 2nd job to start before it starts harvesting.

I could do that, but since not every harvesting session lasts the same amount of time they will get out of sync pretty quickly and could do more damage I'm afraid.

But I would take a guess in that your burning the proxies so quickly that Scrapebox is only running 100 or less threads because its spending so long cycling thru bad proxies for the current threads it can't get more threads started.

On the general whole if you have 6000 proxies and your only getting 340 urls/s I would question the quality of the proxies. I mean Ive gotten 2000 urls/s just with less then 1000 public proxies that I scraped, not even port scanned. I would say you proxies are at the root of the issue.

How did you get to 2000 urls/s with 1000 public proxies?? Can you please give me a good set of public proxies just to try and compare to my current proxy service? It's not that I don't believe you, but I would like to make a comparison and prove to my current provider that his service sucks.

The fact that the proxies are replaced every 2-3 hours and I'm still getting around 200-300 urls/s every single time with a new set of proxies led me to believe it's the max I can get out of that amount of public proxies. I mean it would probably be impossible to get a bad set of proxies every single time. I would've seen some improvement at least one time out of 100 right? Or his service is probably abused to the max.

I did a quick test like this: As soon as the threads were decreased at around 2000 I stopped the session and started another one. The second harvesting session began scraping with 5000 threads steadily again. So I'm not sure if the proxies are getting burned.

Have you tried shorter runs? If you go a lot shorter it might be fine.

I've tried the shortest I can go with my setup, 1 keyword + footprints = around 10k keywords. It did the same thing, only the threads started decreasing around at 70-80% this time which is much better.
 
Last edited:
I have 6000+ port scanned public proxies
Nope you dont have 6k public proxies, maybe you have 6k dead public proxies or not working in SB. With such number of google passed proxies (6k) you would be able to harvest with XXk urls per second and you would pay at least high $xxx per month. To get good public proxies service (working 24/7) you would need good provider and at least ~$50 a month. If you are paying few bucks dont expect quality.

Also speed depends on footprints. Google is better and better in banning proxies. Dont use footprints with opertators, "powered by" unless there is no other way to harvest.
 
I could do that, but since not every harvesting session lasts the same amount of time they will get out of sync pretty quickly and could do more damage I'm afraid.



How did you get to 2000 urls/s with 1000 public proxies?? Can you please give me a good set of public proxies just to try and compare to my current proxy service? It's not that I don't believe you, but I would like to make a comparison and prove to my current provider that his service sucks.

The fact that the proxies are replaced every 2-3 hours and I'm still getting around 200-300 urls/s every single time with a new set of proxies led me to believe it's the max I can get out of that amount of public proxies. I mean it would probably be impossible to get a bad set of proxies every single time. I would've seen some improvement at least one time out of 100 right? Or his service is probably abused to the max.

I did a quick test like this: As soon as the threads were decreased at around 2000 I stopped the session and started another one. The second harvesting session began scraping with 5000 threads steadily again. So I'm not sure if the proxies are getting burned.



I've tried the shortest I can go with my setup, 1 keyword + footprints = around 10k keywords. It did the same thing, only the threads started decreasing around at 70-80% this time which is much better.

Well have you tried not using 5000 connections? Try 500 and see what happens, if you set the connections too high you just create so much overhead that you are going slower then if you didn't have that.

I don't actually use public proxies as a general rule, but here is a video where I show it.

https://www.youtube.com/watch?v=-0sggET1vWo

Bearing in mind that it is with just basic keywords. So try a test without footprints, just basic keywords, because thats the benchmark. As noted above, if you have obscure footprints, then 340 urls/s may actually be as fast as you are going to get.
 
Nope you dont have 6k public proxies, maybe you have 6k dead public proxies or not working in SB. With such number of google passed proxies (6k) you would be able to harvest with XXk urls per second and you would pay at least high $xxx per month. To get good public proxies service (working 24/7) you would need good provider and at least ~$50 a month. If you are paying few bucks dont expect quality.

Also speed depends on footprints. Google is better and better in banning proxies. Dont use footprints with opertators, "powered by" unless there is no other way to harvest.


Well have you tried not using 5000 connections? Try 500 and see what happens, if you set the connections too high you just create so much overhead that you are going slower then if you didn't have that.

I don't actually use public proxies as a general rule, but here is a video where I show it.

https://www.youtube.com/watch?v=-0sggET1vWo

Bearing in mind that it is with just basic keywords. So try a test without footprints, just basic keywords, because thats the benchmark. As noted above, if you have obscure footprints, then 340 urls/s may actually be as fast as you are going to get.

I took your advices and got some proxies from proxygo. 3k google passed proxies(loving the speed, awesome service!) However the issue persists. I was getting around 2700 urls/s until 85%. It scraped around 450k urls until that point and it took 10 minutes. After that the threads started slowly declining and it's going slower and slower. The average urls/s is declining from 2500 urls/s to currently 350urls/s. The more it nears towards the end the slower it gets, the less and less threads it uses. I'm guessing it will take an hour to finish it up. Until the the average urls/s will get to 1.

from 0-85% - 10 minutes
from 85-100% - 3+ hours

So as you can see it's not the proxies. Any other opinion on this issue guys?

I've tried shorter runs, but it gets uglier. The slow declining happens again with every session. The more the sessions the more slower threads declining I will get. The longer the session, the longer I have to wait for the links to be exported. Either way it's no good.

Lower thread count will get me significantly lower speeds with these truly premium proxies, and the slow declining towards the end still happens.

Could this be a design flaw?

So even though I'm getting speeds of 2700 urls/s with these truly premium and expensive proxies, it seems like the average speed is less than 100u/s near the end, which means the full potential of these fast proxies won't ever be used.

I, like many, don't care about scraping all of the keywords, so how about an option in the automator to stop the harvesting process (and continue with the automator job) as soon as it detects lowering of the threads, can this be accomplished by any chance?

Almost all of the footprints from GSA contain something like: "powered by" "designed by" etc.. If I remove these I will be left with less than 10 footprints. I shall be getting footprint factory as well, but until then I need to scrape with the default footprints.

And a few questions if I may:

If I tick both: "Use server proxies" and "Load proxies from file" will it get proxies from both places, or just one?
If I also tick "Use Proxies" and "use server proxies and "load proxies from file" remain, will scrapebox use all of the sources and by what priority if yes?
 
Last edited:
hi, wasnt scrapebox 2 supposed to come out? i havent used it for ages but when i just opened it, it was only on V1.16.4 after checking for updates. how do i get the sb2?
 
hi, wasnt scrapebox 2 supposed to come out? i havent used it for ages but when i just opened it, it was only on V1.16.4 after checking for updates. how do i get the sb2?

Scrapebox v2 as beta is out for a few months now, but already seems very solid and it's much much better than the first one. I use it for harvesting for almost a month with the automator plugin and it works very good.

Here give it a go.
 
We all know there is "blacklist" filter. Maybe you guys can add something opossite, everything not matching "whitelist" filter will be filtered out?
 
does this work on the new windows

I took your advices and got some proxies from proxygo. 3k google passed proxies(loving the speed, awesome service!) However the issue persists. I was getting around 2700 urls/s until 85%. It scraped around 450k urls until that point and it took 10 minutes. After that the threads started slowly declining and it's going slower and slower. The average urls/s is declining from 2500 urls/s to currently 350urls/s. The more it nears towards the end the slower it gets, the less and less threads it uses. I'm guessing it will take an hour to finish it up. Until the the average urls/s will get to 1.

from 0-85% - 10 minutes
from 85-100% - 3+ hours

So as you can see it's not the proxies. Any other opinion on this issue guys?

I've tried shorter runs, but it gets uglier. The slow declining happens again with every session. The more the sessions the more slower threads declining I will get. The longer the session, the longer I have to wait for the links to be exported. Either way it's no good.

Lower thread count will get me significantly lower speeds with these truly premium proxies, and the slow declining towards the end still happens.

Could this be a design flaw?

So even though I'm getting speeds of 2700 urls/s with these truly premium and expensive proxies, it seems like the average speed is less than 100u/s near the end, which means the full potential of these fast proxies won't ever be used.

I, like many, don't care about scraping all of the keywords, so how about an option in the automator to stop the harvesting process (and continue with the automator job) as soon as it detects lowering of the threads, can this be accomplished by any chance?

Almost all of the footprints from GSA contain something like: "powered by" "designed by" etc.. If I remove these I will be left with less than 10 footprints. I shall be getting footprint factory as well, but until then I need to scrape with the default footprints.

And a few questions if I may:

If I tick both: "Use server proxies" and "Load proxies from file" will it get proxies from both places, or just one?
If I also tick "Use Proxies" and "use server proxies and "load proxies from file" remain, will scrapebox use all of the sources and by what priority if yes?

If you tick use server proxies and load proxies from file, yes it gets proxies from both.

If you tick all the options it uses them all. As for order I would have to double check. I "think" it will grab them all and mash them all together and use them.

Ill do some testing on the automator issue. I can't say that I have stared at it for hours and hours, but I run many sessions on a loop and its quite fast. But to be more on a similar test, you are using the GSA footprints correct? Are you just using all 2200 ish that come included with GSA?

Also before you buy footprint factory, you know GSA has a footprint tool that is built right in? Just tossing it out there as its already there.

Are you harvesting from google.com or one of the like 24 hour googles or what?




We all know there is "blacklist" filter. Maybe you guys can add something opossite, everything not matching "whitelist" filter will be filtered out?

For what? There are white list filters for some things already, but you would have to be more specific about what function your talking about.
 
For what? There are white list filters for some things already, but you would have to be more specific about what function your talking about.

Sometimes i want to filter out all links not containing let's say "user.php" or "/blackhat-seo", while i can see scrape going - i can check if my keywords are bringing enough results.
 
If you tick use server proxies and load proxies from file, yes it gets proxies from both.

If you tick all the options it uses them all. As for order I would have to double check. I "think" it will grab them all and mash them all together and use them.

It will be great if it mashes all together and use them, I did some tests to find out and it definitely feels like it does exactly that.

Ill do some testing on the automator issue. I can't say that I have stared at it for hours and hours, but I run many sessions on a loop and its quite fast. But to be more on a similar test, you are using the GSA footprints correct? Are you just using all 2200 ish that come included with GSA?

I'm using the Article engine category from GSA for one harvesting session and the Wiki engine category for a second session. They both go in a loop one after the other. I have filtered out the footprints which have less than 1000 results in google from both categories.

Also before you buy footprint factory, you know GSA has a footprint tool that is built right in? Just tossing it out there as its already there.

Yes, I know about footprint studio in GSA. Footprint Factory has much more options and customizations, however I feel like the developer has abandoned it so I guess I will try out GSA's built-in footprint studio first for a while.

Are you harvesting from google.com or one of the like 24 hour googles or what?

I'm harvesting only google.com. I still want to try weekly or monthly google, but until this gets resolved I will stay with google.com. I also want to throw in bing or yahoo in there in the same time with google.

The issue I'm talking about doesn't seem like it's an automator issue. My theory of what's going on is: as the harvesting session is approaching to the end, at a certain point the remaining keywords become less than the max threads set. So scrapebox begins to slowly decrease thread by thread until it gets to the end. However this is very inefficient when using lots of threads & public proxies. The scraping itself becomes slower and slower as it nears the end so the average urls speed gets much much lower than the period when it was using all of the threads. And when using the automator to scrape for smaller portions this inefficiency becomes larger and larger not using all the resources most of the time. The smaller the portion the larger the gap with more proxies unused.

There are few ways to solve this:

Particularly for the automator:
1. Make somehow the automator plugin fire up a second harvesting session parallel to the first one which will take the unused threads from the first session and start increasing them in the second one so the total sum of both sessions are always the max number of threads set.
2. The dirty way: make an option to stop the harvesting session as soon as the threads have started to decrease (or as soon as they went below a certain previously set number)

For scrapebox itself:
3. The most efficient way: make scrapebox use multiple proxies per keyword. If one keyword can be scraped page by page with multiple proxies at once the threads should stay the same max number set right until the end.

This is how I see things and how I understand this issue by what I see scrapebox does. I could be wrong, but I think I'm understanding where it comes from.
 
How long does my SB get activated? Already submitted my activation details. New to the thing, sorry.

UPDATE: I received a "Unknown ScrapeBox License Email!" email. I clearly entered my PayPal email address. Please advice. :)
 
Last edited:
7e7005fcf33dfacaed3a0dc8889b61c9.png


Then stuck, won't harvest more, each time stops at around 500 urls +

Same with detailed harvester - stuck at 600~~
 
It will be great if it mashes all together and use them, I did some tests to find out and it definitely feels like it does exactly that.



I'm using the Article engine category from GSA for one harvesting session and the Wiki engine category for a second session. They both go in a loop one after the other. I have filtered out the footprints which have less than 1000 results in google from both categories.



Yes, I know about footprint studio in GSA. Footprint Factory has much more options and customizations, however I feel like the developer has abandoned it so I guess I will try out GSA's built-in footprint studio first for a while.



I'm harvesting only google.com. I still want to try weekly or monthly google, but until this gets resolved I will stay with google.com. I also want to throw in bing or yahoo in there in the same time with google.

The issue I'm talking about doesn't seem like it's an automator issue. My theory of what's going on is: as the harvesting session is approaching to the end, at a certain point the remaining keywords become less than the max threads set. So scrapebox begins to slowly decrease thread by thread until it gets to the end. However this is very inefficient when using lots of threads & public proxies. The scraping itself becomes slower and slower as it nears the end so the average urls speed gets much much lower than the period when it was using all of the threads. And when using the automator to scrape for smaller portions this inefficiency becomes larger and larger not using all the resources most of the time. The smaller the portion the larger the gap with more proxies unused.

There are few ways to solve this:

Particularly for the automator:
1. Make somehow the automator plugin fire up a second harvesting session parallel to the first one which will take the unused threads from the first session and start increasing them in the second one so the total sum of both sessions are always the max number of threads set.
2. The dirty way: make an option to stop the harvesting session as soon as the threads have started to decrease (or as soon as they went below a certain previously set number)

For scrapebox itself:
3. The most efficient way: make scrapebox use multiple proxies per keyword. If one keyword can be scraped page by page with multiple proxies at once the threads should stay the same max number set right until the end.

This is how I see things and how I understand this issue by what I see scrapebox does. I could be wrong, but I think I'm understanding where it comes from.

Im using back connect proxies, but I pulled all 2200 ish footprints from GSA, added them on 24 hour google (which netted 200K results) and scraped google. 100 connections and also 25 connections. When it gets near the end the threads do decrease but nothing like your saying, it reaches the point where all keywords are active in a thread and decreases from there till its done. Its quite fast for me.

Ill keep testing, but I can't reproduce what your talking about, grant it Im working on lower connections, but as you say it should be worse. Although you are running 5000 connections right? I mean at a point you are going to have 5000 active queries and its going to reach the end of the list of keywords and the last 5000 are going to have to finish, which may take a bit, but I would assume not hours like you are experiencing.

How long does my SB get activated? Already submitted my activation details. New to the thing, sorry.

UPDATE: I received a "Unknown ScrapeBox License Email!" email. I clearly entered my PayPal email address. Please advice. :)

You can have multiple emails in your paypal account, but the email that is set as Primary is the email that money gets sent from and that is the email address you need to use. So go in your paypal account and find out which mail is set as primary.

View attachment 61485


Then stuck, won't harvest more, each time stops at around 500 urls +

Same with detailed harvester - stuck at 600~~

Does it does this for all queries? You could have no results for the rest of your keywords, proxies could be banned, all sorts of things could be happening. Try a half dozen basic keywords with no footprint, like

test
car
green
truck
sky
blue
etc...
 
i bought this before 4 years... nowadays google competition result shows "0" only... anybody has solution..
(tested after disable my virus cleaner and firewall and some application as "loopline" said on his video.. no result..)
 
i bought this before 4 years... nowadays google competition result shows "0" only... anybody has solution..
(tested after disable my virus cleaner and firewall and some application as "loopline" said on his video.. no result..)

Would you be able to try ScrapeBox v2 http://www.scrapebox.com/v2-beta

The install the Google Competition Finder v2.0.0.3 does this also produce no results for you?
 
Status
Not open for further replies.
Back
Top