Permanently Closed Marketplace Sales Thread

Status
Not open for further replies.
Hi @loopline, I'm trying to use the custom harvester to do searches. I'm aware that there's an issue using it that was fixed in 2.0.0.76, trouble is when I go to update my version it only goes updates to 2.0.0.75... I noticed that and tried to update again and still isn't working but now my backup of my previous version is gone and I'm stuck with only 2.0.0.75. I tried the download link you suggested but it also only gives me 2.0.0.75. :(

It's my first time really using ScrapeBox so there are a couple things I'm not sure about...

I'm trying to use the custom harvester and have about 4k keywords using operators like site:blogspot.com keyword. If I use User Proxies option I get a tiny amount of results and then i'm flooded with Socket 10054 errors, after about 800 errors SB crashes and I can't even submit a report using the bug report option (I fill it out but says it couldn't submit). If I try again, all I get are 10054 errors and no results. I'm only using 10 threads in the harvester. If I use my private proxies instead, I get a 503 error and no results.

I'm using AVG (free version without firewall enabled), AVG is known to 'cause issues but I'm not sure if it's because of this I've disabled it and the problem still persists and I'm not using it's firewall because it's free version so there's no rules to change.
I've tried downloading and using the Windows 7 hotfix you linked in a previous post regarding the 10054 socket errors but whenever I tried the hotfix said it wasn't applicable for my system.
I've downloaded and used the addtofwexecutable from the ScrapeBox download page to add exceptions to the Windows Firewall.
I'm trying to add exceptions to my router and it's firewall but I don't know which ports to allow, etc.

I find it weird that I'm getting some results sometimes (albeit very few, like 100-400 url's) and then other times not. Like, it's not being blocked but then it is all of a sudden. I saw a video of someone doing the exact same thing using 500 threads from 8 months ago, so I don't feel like I'm using too many eventhough I've heard Google's cracked down on things somewhat.

Any thoughts/suggestions? Would love to get this working, I've seen so many awesome things that ScrapeBox can do and I'm excited to get started but this is major roadblock. :(

Firstly this video should help some


I think your using server proxies? Your statements were conflicting so Im guessing. Server proxies and public proxies more often return that socket error 10054 where they are killing the conneciton.

I would uninstall your AVG as a test. Disabling it only stops new rules from forming.

On private proxies your probably going to fast and getting them blocked, 503 is an ip block error. if you don't have at least several hundred private proxies 10 connections is too many. Start with 1 connection for every 100 private proxies you have and go from there. The video above will direct you.

The .76 won't likley offer you anything the .75 doesn't, and scrapebox rolled back the version they have on their site because some users were having issues with the .76 anyway. So stick with what you have till .77 comes out.

Try the detailed harvester, I think the custom harvester is too much for what your trying to do. The detailed harvester will likely yield you better results because of how it works.

Your using advanced operators which will get blocked faster then regular search so plan on going extra slow.

Also as a test you can try a couple keywords/footprints etc.. with no proxies, does that work?

The going along great and then hitting a wall is typical for crossing a threshold where a security software thinks scrapebox is a virus because it has done so many connections and its starts hijacking/scanning all the connections or just closes them or blocks scrapebox.

If you need a good free one, try comodo and whitelist Scrapebox and it should work well.

~~~~~~~

NEW VIDEO

Social Account Scraper Video

 
Last edited:
@loopline Thanks, I'll give that a shot. I uninstalled AVG but it is still throwing up errors. I've had slightly more success now using the detailed harvester and a longer delay but I guess it's just a matter of finding that sweet spot. It went through all the pages from 2 of the keywords and inbetween it added the delay and then started on the next keyword again but I didn't notice it finding anything after the delay the second time. Used 20 second delay with 10 private proxies for "site: blogspot.com mykeyword" type search.

Trying with yahoo and bing using 20 second delay and the ScrapeBox server proxies and the something similar is happening. I'm getting very few initial results and then getting different errors.

9/24/2016 12:38:47 PM: HTTP: 500 , SOCKET: Connection timed out
9/24/2016 12:38:47 PM: HTTP: 0 , SOCKET: Connection reset by peer

I'll try with higher delays and report back but it could take some time because of potential bans. Is there anyway to whitelist scrapebox in my router in case that is doing it?

EDIT: Just tried again with just basic keywords with no operators, 20 second delay using the private proxies and again using ScrapeBox server proxies searching with Google.
ScrapeBox server proxies give 10054 error instantly.
Private proxies (10 of them) give 503 error instantly.

Here's the video I saw of somebody doing the same thing using the ScrapeBox server proxies and 500 threads no problem from January. (2 mins in)

 
Last edited:
So your saying its harvesting links that match whats in the must not be in link field?

Can you give some example urls that are harvested, and show what you have in the must not be removed section? Also if you happen to have a keyword that reproduces this that would be great.
I am NOT configuring a long / complex must be or must not be. Instead I deleted them.
Only rewrite "Just before the url" and "Right after the url"; thus every keyword harevested
per page is near 50 result. "Noice" are elimated, they are ads etc.
But the saved "urls" needs to be processed. The saved file's size is about 2 times than if were havested by default settings.
Then a "Excute exteral program" is followed and Done!!
 
Aliens are controlling your machine and spamming the inter-universe with the custom harvester.

Sorry, in a rare mood, lol

Actually there is a bug in the current version where proxies that use user/pass aren't being dealt with correctly. Proxies that are ip authenticated or when not using proxies works fine. It only affects the custom harvester, the detailed harvester is fine with all proxy types.

The current version on the website is the .74 which should be fine, but you can use the detailed harvester as its already fixed in the .77 which isn't yet released.

http://www.scrapebox.com/payment-received


More then likely the detailed harvester is better for most people using user/pass private proxies anyway, but thats why, or aliens you can pick.

Ok that;s good to know.

Could you link me to the last known version where the custom harvester worked with user/pass proxies properly?

It's a lot faster than the detailed, so I would prefer to use it.

Thanks loopline.
 
Hey :) i recently bought the expired domain finder plugin and i really like it. The only thing that bothers me is why it is limited to 200 threads? I have a very good machina/connection and i´m sure i could use more threads. Could you maybe change the maximum threads to a higher amount in a future update?
 
@loopline Thanks, I'll give that a shot. I uninstalled AVG but it is still throwing up errors. I've had slightly more success now using the detailed harvester and a longer delay but I guess it's just a matter of finding that sweet spot. It went through all the pages from 2 of the keywords and inbetween it added the delay and then started on the next keyword again but I didn't notice it finding anything after the delay the second time. Used 20 second delay with 10 private proxies for "site: blogspot.com mykeyword" type search.

Trying with yahoo and bing using 20 second delay and the ScrapeBox server proxies and the something similar is happening. I'm getting very few initial results and then getting different errors.

9/24/2016 12:38:47 PM: HTTP: 500 , SOCKET: Connection timed out
9/24/2016 12:38:47 PM: HTTP: 0 , SOCKET: Connection reset by peer

I'll try with higher delays and report back but it could take some time because of potential bans. Is there anyway to whitelist scrapebox in my router in case that is doing it?

EDIT: Just tried again with just basic keywords with no operators, 20 second delay using the private proxies and again using ScrapeBox server proxies searching with Google.
ScrapeBox server proxies give 10054 error instantly.
Private proxies (10 of them) give 503 error instantly.

Here's the video I saw of somebody doing the same thing using the ScrapeBox server proxies and 500 threads no problem from January. (2 mins in)


The 10054 for the server proxeis is where a connection is killed. This is semi common for public proxies. They aren't likley to work with google as they get banned quickly and further using a delay with server proxies makes no sense as they are going to die even if you do nothing. So reserve the delay for the private proxies and use no delay for server proxies. I would also go to settings >> connections timeouts and other settings >> more harvester tab - and run the proxy retries all the way up for server proxies.

503 and 302 from google are ip bans. but if you want 12-48 hours google should unban your private proxies and then you can try a higher delay perhaps. Also you can try deeperweb and google api, which are google powered, but have their own ip bans. They also tend to work ok with server proxies, but pretty decent indeed with private proxies.

I am NOT configuring a long / complex must be or must not be. Instead I deleted them.
Only rewrite "Just before the url" and "Right after the url"; thus every keyword harevested
per page is near 50 result. "Noice" are elimated, they are ads etc.
But the saved "urls" needs to be processed. The saved file's size is about 2 times than if were havested by default settings.
Then a "Excute exteral program" is followed and Done!!

Good deal, if that works for you, great! I love the flexibility of Scrapebox. :)

Ok that;s good to know.

Could you link me to the last known version where the custom harvester worked with user/pass proxies properly?

It's a lot faster than the detailed, so I would prefer to use it.

Thanks loopline.

There isn't one, they don't archive version 2 on the server and the current version is .76. However you can go to help >> restore previous version - and roll back to whatever you had before.

you can also run multiple isntances of scrapebox and harvest in each of them. Or if your using private proxies with user/pass and you can have your provider set them up as ip authenticated proxies then that would solve the issue as well.. The custom harvester works fine with ip auth proxies or public proxies. Just not user/pass proxies.

Hey :) i recently bought the expired domain finder plugin and i really like it. The only thing that bothers me is why it is limited to 200 threads? I have a very good machina/connection and i´m sure i could use more threads. Could you maybe change the maximum threads to a higher amount in a future update?

Windows gets "grumpy" when you start running more then 200 threads. I can speak from experience, I don't run over 100 on average because you get diminishing returns. Same thing happens for other programs, its not just scrapebox.

But that said you can run an unlimited number of instances for scrapebox. So long as your scraping more then 1 domain to start you can get multiple copies of scrapebox expired domain finder cranking at once.
 
Is your advice to only run 100 threads of the domain finder at once? I thought that more threads would not be a problem because only ~ 20% of my bandwidth and cpu/ram resources are used at average. I often use like my proxyharvester with 200 threads, my proxy harvester with 200 threads and 2 or 3 instances of the domain finder with 200 threads at the same time and i find a huge amount of expired domains this way. Its more of "brute force" way. Is the domain finder "skipping" domains or whats the problem with having so much threads running at the same time? Can i change some windows settings to maybe optimize my performance? I run windows 10.
 
I ran 50 threads on vps but you are free to increase/update it a little higher in real time, I am not in a hurry slapping/pushing it's(vps) arse.

There are many debates on how many connections should openned over and over again, if has enough bandwidth, try increase
it a little 'higher'. On my machine, >50 connections makes no / little sense, the total bandwidth is shared by the #literal number of connections. More than this there's no more free "bandwidth funnel" allowing data transferring.

Threads or connections,---I have no clear idea what's the differece between them, but I guess sure you know what it means.
 
I haven't search or read the whole thread here but wondering if I can "penetrate" cloudflare protected website when posting? I don't know how to record all the html pages in which there might be several cloudflare script on them.
 
Looking for some advice to deal with Socket Error 10054 on Windows 7. I downloaded the Hotfix that was mentioned by Loopline earlier and tried to install but it said no applicable to my machine.

This happens on the Social Account Scraper with as little as 5 connections and after about 10 minutes of scraping Google, but it doesnt for Bing. Really weird.
 
Thanks Loopline, nothing seems to be helping though. I've tried again with the delay up to 180 seconds (3 minutes), uninstalled AVG completely, installed Comodo and whitelisted ScrapeBox, but I'm still getting the same behavior from ScrapeBox. Comodo was helpful in determining what ports I needed to open up in my router but still no luck. It goes through 1-2 keywords and then throws up errors. It doesn't even recognize the other proxies after that. Only thing I can think of is that they're bad proxies that aren't completely private and being banned regularly or my ISP/router is blocking it somehow (I'm using Telus as my ISP in Canada). I was planning on using a VPS at some point in the future anyway so I think I'll just try that since I've already dumped countless hours into getting this working.
 
Is your advice to only run 100 threads of the domain finder at once? I thought that more threads would not be a problem because only ~ 20% of my bandwidth and cpu/ram resources are used at average. I often use like my proxyharvester with 200 threads, my proxy harvester with 200 threads and 2 or 3 instances of the domain finder with 200 threads at the same time and i find a huge amount of expired domains this way. Its more of "brute force" way. Is the domain finder "skipping" domains or whats the problem with having so much threads running at the same time? Can i change some windows settings to maybe optimize my performance? I run windows 10.

Windows wasn't "built" with running thousands of connections in mind.

What I was saying is its better to run 5 instances of the expired finder, each at 200 connections then it is to run 1 instance at 1000 connections. Inside of 1 instance there is a point of diminishing returns when upping the connections, its not a linear progression. If that makes sense. On my personal setup, which as all sorts of custom scripts and instances of scrapebox doing all sorts of things, I find that running 100 thread per instance is ideal and my success rate drops off when I exceed that.

your setup may be different, but all setups are subject to windows overall limitations in that hundreds and hundreds of connections on the same instance is not ideal. So try and see is all Im saying.

I haven't search or read the whole thread here but wondering if I can "penetrate" cloudflare protected website when posting? I don't know how to record all the html pages in which there might be several cloudflare script on them.

Cloudfare is hard. There is no way that I am aware of to get around it, unless a different IP will solve the issue.

Looking for some advice to deal with Socket Error 10054 on Windows 7. I downloaded the Hotfix that was mentioned by Loopline earlier and tried to install but it said no applicable to my machine.

This happens on the Social Account Scraper with as little as 5 connections and after about 10 minutes of scraping Google, but it doesnt for Bing. Really weird.

This is with private proxies or public proxies or no proxies?

Did you whitelist the social account scraper in your security software?

Thanks Loopline, nothing seems to be helping though. I've tried again with the delay up to 180 seconds (3 minutes), uninstalled AVG completely, installed Comodo and whitelisted ScrapeBox, but I'm still getting the same behavior from ScrapeBox. Comodo was helpful in determining what ports I needed to open up in my router but still no luck. It goes through 1-2 keywords and then throws up errors. It doesn't even recognize the other proxies after that. Only thing I can think of is that they're bad proxies that aren't completely private and being banned regularly or my ISP/router is blocking it somehow (I'm using Telus as my ISP in Canada). I was planning on using a VPS at some point in the future anyway so I think I'll just try that since I've already dumped countless hours into getting this working.

To be honest Id agree that at this point the VPS makes more sense. Scrapebox only uses port 80 in general except the whois checker uses port 43 and then any proxy ports for private proxies. Else you don't need to open any ports. But a router can be picky.

You could try and bypass the router and connect direct or try a mobile dongle, friends network etc.. And try and eliminate the router as the issue or if all works maybe it is the router (or ISP). But if your leaning towards a VPS that will probably be simpler, give you dedicated resources, lower latency since its in a data center, which means faster progress etc..
 
Windows wasn't "built" with running thousands of connections in mind.

What I was saying is its better to run 5 instances of the expired finder, each at 200 connections then it is to run 1 instance at 1000 connections. Inside of 1 instance there is a point of diminishing returns when upping the connections, its not a linear progression. If that makes sense. On my personal setup, which as all sorts of custom scripts and instances of scrapebox doing all sorts of things, I find that running 100 thread per instance is ideal and my success rate drops off when I exceed that.

your setup may be different, but all setups are subject to windows overall limitations in that hundreds and hundreds of connections on the same instance is not ideal. So try and see is all Im saying.



Cloudfare is hard. There is no way that I am aware of to get around it, unless a different IP will solve the issue.



This is with private proxies or public proxies or no proxies?

Did you whitelist the social account scraper in your security software?



To be honest Id agree that at this point the VPS makes more sense. Scrapebox only uses port 80 in general except the whois checker uses port 43 and then any proxy ports for private proxies. Else you don't need to open any ports. But a router can be picky.

You could try and bypass the router and connect direct or try a mobile dongle, friends network etc.. And try and eliminate the router as the issue or if all works maybe it is the router (or ISP). But if your leaning towards a VPS that will probably be simpler, give you dedicated resources, lower latency since its in a data center, which means faster progress etc..
I forget that I can use page scanner to spot them with or without change IPs.
 
There isn't one, they don't archive version 2 on the server and the current version is .76. However you can go to help >> restore previous version - and roll back to whatever you had before.

you can also run multiple isntances of scrapebox and harvest in each of them. Or if your using private proxies with user/pass and you can have your provider set them up as ip authenticated proxies then that would solve the issue as well.. The custom harvester works fine with ip auth proxies or public proxies. Just not user/pass proxies.

Yeah I made the mistake of updating twice in a row and hence losing my working old version.

Would you by any chance dig up an older version in exchange for a beer? :)

My proxies have to stay as they are unfortunately so that's not an option.
 
This is with private proxies or public proxies or no proxies?

Did you whitelist the social account scraper in your security software?

.
Private and No Proxies. Whats weird, is i created a custom grabber with Regex to grab the social links instead and it runs fine all night long. Its just not as tidy and requires a good bit of clean up.

I'm not running any security software. Maybe its a modem thing, but it doenst make sense why it works on the custom grabber and not the social scraper. Same amount of threads.
 
Yeah I made the mistake of updating twice in a row and hence losing my working old version.

Would you by any chance dig up an older version in exchange for a beer? :)

My proxies have to stay as they are unfortunately so that's not an option.

I pmed you

Private and No Proxies. Whats weird, is i created a custom grabber with Regex to grab the social links instead and it runs fine all night long. Its just not as tidy and requires a good bit of clean up.

I'm not running any security software. Maybe its a modem thing, but it doenst make sense why it works on the custom grabber and not the social scraper. Same amount of threads.

Well that doesn't really make sense. The social addon is its own application completely where as the custom data scraper is part of the main scrapebox application.

However you said it happens with google, but never with bing?

What anti virus and malware checker are you using?
 
Last edited:
Status
Not open for further replies.
Back
Top