Permanently Closed Marketplace Sales Thread

Status
Not open for further replies.
Thanks for the rep. ;-)

See, I generally prefer to rake up connections for one search engine to say 15 while I scrape using it ONLY. But if I use multiple engines, I use 10 connections.

And hell yes, thinking that a high no of connections would serve to show better results, that's wrong. If you have an optimum setting, then only you can use SB to maximum.,...

Thank you so much :o
What's your connection speed and url/s?

Weird I'm getting 0 URLs or 100 URLs scraped per keyword where I should be getting 1000 wtf -_-
 
Last edited:
Thank you so much :o
What's your connection speed and url/s?

Weird I'm getting 0 URLs or 100 URLs scraped per keyword where I should be getting 1000 wtf -_-

My DL speed according to SpeedTest is 0.51 Mb/s and I have 10 connections for each SE, and I scrape only Yahoo, normally. I get ~70 URL/s with that.

Maybe you have the limit set to 100? Above the proxies box, make sure you have 1000.
 
Can anyone help me
i wanna buy scrapebox but don't have paypal or cc coz it's not allowed in my country
please i need someone to buy it for me and i will pay him back throught e-bank like moneybookers

any help is welcome

please pm me
 
Hmm this is really weird. I tested both of these settings:
25 connections
10 connections

For 25 connections, I got around 50 url/s and same with 10 connections and stopped when I reached the 6000 mark BUT when I removed duplicates, 2000+ were found for the URLs scraped using 25 connections and only 200 were found using 10 connections. Can anyone explain to me why this is happening since the URL/s is the same?
I used the same keywords in both cases without any proxies whatsover (minor test).
 
Last edited:
I noticed that the harvester can "only" handle 1 000 000 URLs, is this a limit that can/will be removed? I also think that might have been the reason for my latest crash. I had a bit over 20k unique domains I was harvesting "deeplinks" for, with 1000 hits per keyword/site.

Thanks in advance.

//Chamezz.
 
Just receive my Private Proxies from squid proxy
today they work FN amazing
1 question? what i should"t use them For so i dont wast them
i have 25
thanks
 
I noticed that the harvester can "only" handle 1 000 000 URLs, is this a limit that can/will be removed? I also think that might have been the reason for my latest crash. I had a bit over 20k unique domains I was harvesting "deeplinks" for, with 1000 hits per keyword/site.

Thanks in advance.

//Chamezz.

the one million limit is irrevocable.

Almost everyone crashes when harvesting or posting to HUGE lists. Check system resources. How much ram do you have? Can you get more? Does your processor suck ass? Can you get a better one? Ram is actually more important. I had some luck using the thumb drive as ram trick. Try it may work for you.
 
Hmm this is really weird. I tested both of these settings:
25 connections
10 connections

For 25 connections, I got around 50 url/s and same with 10 connections and stopped when I reached the 6000 mark BUT when I removed duplicates, 2000+ were found for the URLs scraped using 25 connections and only 200 were found using 10 connections. Can anyone explain to me why this is happening since the URL/s is the same?
I used the same keywords in both cases without any proxies whatsover (minor test).

That is because of the fact that some sites have both their "www.site.com" and "site.com" pages indexed. For G, they are different, for SB they are same.

Another thing, I would say don't stop scraping till you reach a certain number of URLs. You set a time limit. Like scrape with 10 connections for 30 minutes today. Then with 25 connections, 30 minutes tomorrow. And then compare the results.

Just receive my Private Proxies from squid proxy
today they work FN amazing
1 question? what i should"t use them For so i dont wast them
i have 25
thanks

See this page http://squidproxies.com/appropriate-use-policy/

If you still don't find an answer, contact them.
 
Last edited:
I'm using a Intel Core i7 930 Overclocked to 4GHz, along with 6GB DDR3 RAM.

The only problem isn't the program crashing, that actually only happened once. The other problem is that now I can't combine my lists from when I scraped each search engine separately, to remove the duplicate URLs.

Btw, is there something wrong with the forum filter? I can't quote people, it says I'm trying to post links or sell stuff. -.-
 
Last edited:
I'm using a Intel Core i7 930 Overclocked to 4GHz, along with 6GB DDR3 RAM.

The only problem isn't the program crashing, that actually only happened once. The other problem is that now I can't combine my lists from when I scraped each search engine separately, to remove the duplicate URLs.

Btw, is there something wrong with the forum filter? I can't quote people, it says I'm trying to post links or sell stuff. -.-

I don't get your question. Is the problem that you're facing that when you scrape a SE, and then again scrape using other SE, then the former list in harvester is replaced, and not actually merged with the new scrape? Right?

Then, technically, idk what's happening, but you can always import to add to current list the file from your "Harvester_Sessions" folder.

Also, There's no filter as far as I know. It MIGHT be because when you quote someone, there's a link to the original post as well, and you can't post links in BHW unless you have 15 posts.

For that, I would suggest you use this:

[QUOTE]Copy-paste the portion you want to quote[/QUOTE]

rather than

[quote="Chamezz, post: 2945993"]Don't click the "Quote" button from original post[/QUOTE]
 
That is because of the fact that some sites have both their "www.site.com" and "site.com" pages indexed. For G, they are different, for SB they are same.

Another thing, I would say don't stop scraping till you reach a certain number of URLs. You set a time limit. Like scrape with 10 connections for 30 minutes today. Then with 25 connections, 30 minutes tomorrow. And then compare the results.



See this page http://squidproxies.com/appropriate-use-policy/

If you still don't find an answer, contact them.

That wasn't exactly the answer to my question. Mine was why it showed different duplicate results (2000 compared to 200) when using different harvester connections.
 
I don't get your question. Is the problem that you're facing that when you scrape a SE, and then again scrape using other SE, then the former list in harvester is replaced, and not actually merged with the new scrape? Right?

Then, technically, idk what's happening, but you can always import to add to current list the file from your "Harvester_Sessions" folder.

Also, There's no filter as far as I know. It MIGHT be because when you quote someone, there's a link to the original post as well, and you can't post links in BHW unless you have 15 posts.

For that, I would suggest you use this:



rather than


Ah, thank you for the suggestion.

My problem is that the harvester can't handle over 1 million URLs, so if I scrape one SE first and get 6-700k urls. Then i scrape another SE and get say, 4-500k urls. If I then want to merge those two lists, which would add up to 1.2 million urls, to remove all the duplicate urls. The limit sets in.

My "process":

Scrape keywords
Remove duplicate domains
Replace beginning of all links with site:
Put domains in keyword-list
Harvest links from one SE at a time, due to the limit.
Save the links to a separate list after every scrape.

Then what I would like to do is merge those lists and remove all duplicate links.

Does that clear it up?

I tried quoting without name and just the portion I wanted to quote, still didn't work.

I also found what was causing the "filter" to react. But I can't write it, yet.
 
Ah, thank you for the suggestion.

My problem is that the harvester can't handle over 1 million URLs, so if I scrape one SE first and get 6-700k urls. Then i scrape another SE and get say, 4-500k urls. If I then want to merge those two lists, which would add up to 1.2 million urls, to remove all the duplicate urls. The limit sets in.

My "process":

Scrape keywords
Remove duplicate domains
Replace beginning of all links with site:
Put domains in keyword-list
Harvest links from one SE at a time, due to the limit.
Save the links to a separate list after every scrape.

Then what I would like to do is merge those lists and remove all duplicate links.

Does that clear it up?

I tried quoting without name and just the portion I wanted to quote, still didn't work.

I also found what was causing the "filter" to react. But I can't write it, yet.

Im pretty sure the duperemove addon can handle well over the 1 million limit. Click addons and than launch the duperemove one. I merged 10 million into one file the other day and removed the duplicates. Also Loopline has a free multi tool that does the same thing and has the added feature of being able to keep 3 or x number of urls from each domain. So I merged my 10 million urls into one file and now I am having looplines multi tool sort it and only keep 3 urls from each domain.
 
I jsut went to ask the scrapebox team a question and it says that support is done since the 18th of june because of a flood. anybody know when its gonna be up again. I want to switch my license so I can use scrapebox on my vps.

As the contact form said, support is up and license activations are still being done on a priority basis. I have had contact with support and its all good, just they are having a harder time of it right now.

I have macbook. How to run SB on mac? I try bootcamp but i get error on installation. I really need run scrapebox. Any advice?

Bootcamp should work, what error are you getting? Parallels works.

Use a fair amount, and spin the holy hell out of them. The spintax exponentially increases the variety of your comments by a great deal. Something like the best spinner can help with that alot, though even just taking a few minutes to think of variations on your words can spin like a top.

Other free tip - proof read the various spun versions and make sure they don't sound too bot-generated.

Gregs is back! Good to see you mate!

site:.edu will be more accurate right? I use inurl and get rubbish results lol

probably site:.edu is better as inurl ignore the period so google only sees:

inurl:edu and that could give you a lot of randomness, where as with site:.edu google sees

site:.edu

so its more accurate.

Why not take it to the next level aye?

link:http://www.CompetitorsSite.com site:.edu

Steal those competitors .edu and .gov backlinks. :)



I'm on a 4MBPS connection and I have no idea how to check the optimal settings to set my SB harvester too.
I only use G and Y.

I tried 10 each, 20 each, 30 each, 500 each, and I couldn't find a difference. All of them became 40url\s at the end. How do I properly check this?
I will give a super huge thanks and rep to anyone that is so kind to help me. I feel like a lost fool :confused:

I run dedi servers on 1000Mbit lines, and it doesn't matter if I set connections to 10 or 50 on Yahoo, I still only get about 200urls/sec.

The engines kick it out pretty quick either way. You can use proxies to get a higher urls/sec rate, but the engines have it more or less throteled per IP so you get so much of their bandwidth and thats it. So it looks like for your inet connection speed/engine match up 40 urls/sec is it. Add more proxies and see if that makes a difference, if not then its your connection.

Cause if your connection is too slow and you set it to 500 connections, some will go thru and the rest will timeout and you will still only get what fits in your pipe. 4Mbit is decent, so you should be able to do well when working with the engines.


Hmm generally when you scrape, are you supposed to do both SE's at the same time or one session per keyword for different SE's?
Do you mean that if it's not properly set up, I will get less total URLs scraped and total unique URLs scraped? That's weird, I wonder how that works, but thank you :p rep given

I run all 4 search engines at the same time no problem, but again it depends on your 4mbit conneciton and how many connections you are running to the engines.

In theory depending on your proxies and such you could definitely get less then ideal results. The single threaded harvester is accurate, but the multi threaded harvester, in the words of the developer (if I accurately recall a hundred pages back, lol) "is designed to rip the heart out of google and the other engines" meaning its going for over all volume and if the proxy fails past the retries set under settings it just moves on.

I don't think there is a right and wrong way to do it in regards to how many engines you use at once. Just try 1 and look at the results and then try 2 or all 4. If overall results are worse by using all 4 then 1 at a time for the same searches with the same proxies or no proxies then adapt to using 1 or 2 at a time. If your proxies are always accurate (private proxies) and your connection is fast enough then you should get the same results either way. Also you can try turning the over all connections down for each engine if your are using all 4, that will help accuracy if your internet connection is the issue.

Hmm this is really weird. I tested both of these settings:
25 connections
10 connections

For 25 connections, I got around 50 url/s and same with 10 connections and stopped when I reached the 6000 mark BUT when I removed duplicates, 2000+ were found for the URLs scraped using 25 connections and only 200 were found using 10 connections. Can anyone explain to me why this is happening since the URL/s is the same?
I used the same keywords in both cases without any proxies whatsover (minor test).

You stopped each at 6000 results?

It could be that on 1 you removed duplicate urls and on the other you removed duplicate domains.

It could also be that with the 25 connections more of your connections timed out and therefor moved further down your keyword list thus resulting in more unique urls.

The accurate test would be to let both finish out completely and then compare results. Of course without proxies, your IP might get blocked first, but I have never really noticed anything like this. But when it comes to working with millions of urls, I am more concerned about over all volume.

I noticed that the harvester can "only" handle 1 000 000 URLs, is this a limit that can/will be removed? I also think that might have been the reason for my latest crash. I had a bit over 20k unique domains I was harvesting "deeplinks" for, with 1000 hits per keyword/site.

Thanks in advance.

//Chamezz.

That 1 million limit is the limit set by Microsoft for the grid that scrapebox uses for that element. All urls beyond 1 millon are stored in the harvester_sessions folder in your mail scrapebox folder. I have harvested 100million+ urls in 1 run.

Just receive my Private Proxies from squid proxy
today they work FN amazing
1 question? what i should"t use them For so i dont wast them
i have 25
thanks

Well so long as you don't violate the squid proxies TOS then there isn't anything you shouldn't do with it. I use my Your Private Proxy proxies for everything in scrapebox, bar nothing.

A new feature request:

One click update all installed addons...

Now why do you have to go and make things simple? lol :p

I'm using a Intel Core i7 930 Overclocked to 4GHz, along with 6GB DDR3 RAM.

The only problem isn't the program crashing, that actually only happened once. The other problem is that now I can't combine my lists from when I scraped each search engine separately, to remove the duplicate URLs.

Btw, is there something wrong with the forum filter? I can't quote people, it says I'm trying to post links or sell stuff. -.-

I just multi quoted several people so nothing wrong with quoting on my end.

There is a dupe remove addon that will remove duplicate urls or duplicate domains for up to 180 million urls. You first use it to combine all files you want to remove duplicates from and then it will remove the dupes.

I heard that this was $47, did it go up to $57 now?
 
Thanks loopline for taking the time to answer my questions :] Rep and thanks given!
 
As the contact form said, support is up and license activations are still being done on a priority basis. I have had contact with support and its all good, just they are having a harder time of it right now.



Bootcamp should work, what error are you getting? Parallels works.



Gregs is back! Good to see you mate!



probably site:.edu is better as inurl ignore the period so google only sees:

inurl:edu and that could give you a lot of randomness, where as with site:.edu google sees

site:.edu

so its more accurate.

Why not take it to the next level aye?

link:http://www.CompetitorsSite.com site:.edu

Steal those competitors .edu and .gov backlinks. :)





I run dedi servers on 1000Mbit lines, and it doesn't matter if I set connections to 10 or 50 on Yahoo, I still only get about 200urls/sec.

The engines kick it out pretty quick either way. You can use proxies to get a higher urls/sec rate, but the engines have it more or less throteled per IP so you get so much of their bandwidth and thats it. So it looks like for your inet connection speed/engine match up 40 urls/sec is it. Add more proxies and see if that makes a difference, if not then its your connection.

Cause if your connection is too slow and you set it to 500 connections, some will go thru and the rest will timeout and you will still only get what fits in your pipe. 4Mbit is decent, so you should be able to do well when working with the engines.




I run all 4 search engines at the same time no problem, but again it depends on your 4mbit conneciton and how many connections you are running to the engines.

In theory depending on your proxies and such you could definitely get less then ideal results. The single threaded harvester is accurate, but the multi threaded harvester, in the words of the developer (if I accurately recall a hundred pages back, lol) "is designed to rip the heart out of google and the other engines" meaning its going for over all volume and if the proxy fails past the retries set under settings it just moves on.

I don't think there is a right and wrong way to do it in regards to how many engines you use at once. Just try 1 and look at the results and then try 2 or all 4. If overall results are worse by using all 4 then 1 at a time for the same searches with the same proxies or no proxies then adapt to using 1 or 2 at a time. If your proxies are always accurate (private proxies) and your connection is fast enough then you should get the same results either way. Also you can try turning the over all connections down for each engine if your are using all 4, that will help accuracy if your internet connection is the issue.



You stopped each at 6000 results?

It could be that on 1 you removed duplicate urls and on the other you removed duplicate domains.

It could also be that with the 25 connections more of your connections timed out and therefor moved further down your keyword list thus resulting in more unique urls.

The accurate test would be to let both finish out completely and then compare results. Of course without proxies, your IP might get blocked first, but I have never really noticed anything like this. But when it comes to working with millions of urls, I am more concerned about over all volume.



That 1 million limit is the limit set by Microsoft for the grid that scrapebox uses for that element. All urls beyond 1 millon are stored in the harvester_sessions folder in your mail scrapebox folder. I have harvested 100million+ urls in 1 run.



Well so long as you don't violate the squid proxies TOS then there isn't anything you shouldn't do with it. I use my Your Private Proxy proxies for everything in scrapebox, bar nothing.



Now why do you have to go and make things simple? lol :p



I just multi quoted several people so nothing wrong with quoting on my end.

There is a dupe remove addon that will remove duplicate urls or duplicate domains for up to 180 million urls. You first use it to combine all files you want to remove duplicates from and then it will remove the dupes.

You quoted me but forgot to answer my question :p
 
Status
Not open for further replies.
Back
Top