Permanently Closed Marketplace Sales Thread

Status
Not open for further replies.
Sent an email about a bug, wanted to make sure you got it because it's killing my team's workflow.

When scraping emails by crawling sites and you select "save URL with email" it's no longer saving the URL with the email.

Can you please confirm this is a bug you're able to recreate? Is there any way for me to download a previous version until a patch is released? If so, what is the latest version that doesn't have this issue?
 
@loopline

Can you give some recommended UDP/TCP timeout settings for running Scrapebox on DD-WRT?

FYI - I'm using 100 private proxies.

When using the domain resolver, every single domain that is checked is opening a new UDP port/connection. So if I use 50 threads, by the time a list of 100k domains is 20% finish, I have 16,000 ports/connections open on DD-WRT. I have a really good router, but never in the past did running SB at only 50 threads hit 16k connections that quickly.

My current settings are:

TCP Congestion Control:
Westwood

TCP Timeout:
3600 seconds (This is high -- default, but it seems SB is using UDP connections anyways).

UDP Timeout:
120 seconds

Also (important) -- Does SB use "WinHTTP" for applying proxies? This service wasn't started for me -- are my connections not going through the proxies correctly? Can you say which services are required to implement the proxies properly? I disabled a bunch of services to speed up Windows 7 recently and wondering if that's what did this.

Thanks a lot!
 
Last edited:
Some other services that I have disabled that I'm wondering if they have any impact on proxies being used:

SSDP Discovery (There are no homegroups / network groups on my network. No file-sharing. I only use Teamviewer).
UPnP Device Host
IP Helper
Server
Workstation
TCP/IP NetBIOS Helper

Update: It definitely seems as if without a doubt one or more of these services is handling the ability to use proxies. My connections are still filling up very fast, but instead of local IP -> Resolving IP, as in... it seems that it's local IP -> router or local IP -> DNS... and no connections are leaking. The question is -- which service(s) are responsible for this?
 
Last edited:
Final update:

It's really hard to say if any of these services have really any impact. I disabled them all, ran the domain checker again, and don't really see any "leaks" besides a few Google connections (likely from Chrome) that I'm not sure why they are there. I closed all connections / browsers etc while running the test. Still though, are you aware of any services that impact the ability of SB to properly use proxies?

One thing I did notice however, is that when I run GSA, the proxies are going local IP -> proxy server, but when using SB, it seems to be my local IP connecting with the DNS, instead of going through the proxies... If that makes sense.

GSA -
192.168.1.xxx -> proxy

SB -
192.169.1.xxx -> 192.168.1.1
 
Last edited:
Sorry for so many posts but sadly BHW won't let me edit my previous posts. I'm going to assume when GSA posts, the reason it's connecting to the proxies is because it's posting through TCP, whereas in SB using Domain Resolver, is going through UDP.

Okay, I ran an indexing check and SB is using the proxies properly when connecting through TCP. Now I just need to better understand UDP timeouts so I can use the domain resolver and other UDP functions of SB without quickly hitting 16k (I've even tested with 32k) connections. I have a really good router, I think my timeout settings are just off.
 
Are there any BHW, or otherwise, discounts for the Premium Plugins or only for the base software?
 
Ok I understand that now thanks for clearing it up!

Still having the same issue though, the program freezes. I'll up my post count and return when I can upload pictures.

I replied your email about this. I think its about this anyway. if not shoot me another mail, mail is easier then this anyway.

Sent an email about a bug, wanted to make sure you got it because it's killing my team's workflow.

When scraping emails by crawling sites and you select "save URL with email" it's no longer saving the URL with the email.

Can you please confirm this is a bug you're able to recreate? Is there any way for me to download a previous version until a patch is released? If so, what is the latest version that doesn't have this issue?

Check your PM.

@loopline

Can you give some recommended UDP/TCP timeout settings for running Scrapebox on DD-WRT?

FYI - I'm using 100 private proxies.

When using the domain resolver, every single domain that is checked is opening a new UDP port/connection. So if I use 50 threads, by the time a list of 100k domains is 20% finish, I have 16,000 ports/connections open on DD-WRT. I have a really good router, but never in the past did running SB at only 50 threads hit 16k connections that quickly.

My current settings are:

TCP Congestion Control:
Westwood

TCP Timeout:
3600 seconds (This is high -- default, but it seems SB is using UDP connections anyways).

UDP Timeout:
120 seconds

Also (important) -- Does SB use "WinHTTP" for applying proxies? This service wasn't started for me -- are my connections not going through the proxies correctly? Can you say which services are required to implement the proxies properly? I disabled a bunch of services to speed up Windows 7 recently and wondering if that's what did this.

Thanks a lot!

I have no idea which services are needed. You can mail support and ask, but I don't mess with things at this level. I just point and shoot. Meaning I spin up a server, put scrapebox on it and run. I don't even mess with my local machine at this level. Im not saying its bad, Im just saying I just have too much to do and I just need stuff to "work" and I keep moving. Windows works out of the box, so I go with it. lol

Its actually been about 5 or 6 years since I tried running any real production for Scrapebox on a home based machine. At that point I moved to a server, even a $30 VPS can get some serious damage done. Anyway thats me, but I don't know the answer to your question.

Some other services that I have disabled that I'm wondering if they have any impact on proxies being used:

SSDP Discovery (There are no homegroups / network groups on my network. No file-sharing. I only use Teamviewer).
UPnP Device Host
IP Helper
Server
Workstation
TCP/IP NetBIOS Helper

Update: It definitely seems as if without a doubt one or more of these services is handling the ability to use proxies. My connections are still filling up very fast, but instead of local IP -> Resolving IP, as in... it seems that it's local IP -> router or local IP -> DNS... and no connections are leaking. The question is -- which service(s) are responsible for this?

Final update:

It's really hard to say if any of these services have really any impact. I disabled them all, ran the domain checker again, and don't really see any "leaks" besides a few Google connections (likely from Chrome) that I'm not sure why they are there. I closed all connections / browsers etc while running the test. Still though, are you aware of any services that impact the ability of SB to properly use proxies?

One thing I did notice however, is that when I run GSA, the proxies are going local IP -> proxy server, but when using SB, it seems to be my local IP connecting with the DNS, instead of going through the proxies... If that makes sense.

GSA -
192.168.1.xxx -> proxy

SB -
192.169.1.xxx -> 192.168.1.1

Sorry for so many posts but sadly BHW won't let me edit my previous posts. I'm going to assume when GSA posts, the reason it's connecting to the proxies is because it's posting through TCP, whereas in SB using Domain Resolver, is going through UDP.

Okay, I ran an indexing check and SB is using the proxies properly when connecting through TCP. Now I just need to better understand UDP timeouts so I can use the domain resolver and other UDP functions of SB without quickly hitting 16k (I've even tested with 32k) connections. I have a really good router, I think my timeout settings are just off.

So just to sum up my answer to all your above posts in 3 words - I don't know.

I don't operate like that, I don't dig into that sort of stuff. I have dug into things like that in the past, but it didn't help me make any money. These days I have so much going I focus on what needs to happen.

My solution to your issues would be to either

1 - buy a more powerful machine if windows was slow and just leave it alone

2 - push scrapebox to a server.

Im just being honest. I respect your level of work and the desire to dig in at that level. My method is inefficient at utilizing my physical machine resources, I realize that, and your trying to be efficient, which is great, I get it. I applaud you for it.

For me, at this point, trying to figure out what your figuring out would be efficient in using my machine resources but inefficient in making money because it takes time, time I can use to make money. So I throw some dollars at the issue and solve it quickly and then go make money elsewhere. Thats just me, and how I prefer to be honestly.

I used to build my machines from scratch, and I built a bitcoin miner from scratch with a crazy open air frame and all, but that was more along the lines of fun. When it comes to making money Ive got no patience for messing with it, I want t order a laptop on amazon, open it, and start working.

Anyway, Ill pm you some info of the person that probably knows the answer to your question.

Are there any BHW, or otherwise, discounts for the Premium Plugins or only for the base software?

There are no discounts for anyone anywhere for the plugins outside of bundle plugins. So you can purchase a bundle of 5 or all 7 plugins and get a nice discount, but thats it. Info here:
http://www.scrapebox.com/plugins
 
There are no discounts for anyone anywhere for the plugins outside of bundle plugins. So you can purchase a bundle of 5 or all 7 plugins and get a nice discount, but thats it. Info here:
http://www.scrapebox.com/plugins

I figured not but thought I'd give it a shot.

I wish I would've known about the bundle, I already have half the plugins ... oh well.
 

You will need to get in contact with your friend who purchased the license. Obviously we cant discuss someone elses license information with you.




I dropped her email but she is not responding to me and I invite you to come on teamviewer I show you email sent me few years back when she bought this license for me and emailed me all details. Now next excuse????? Should we come to solution ?





WHY YOU ARE NOT ANSWERING ME???????? @Sweetfunny @loopline
 

You will need to get in contact with your friend who purchased the license. Obviously we cant discuss someone elses license information with you.




I dropped her email but she is not responding to me and I invite you to come on teamviewer I show you email sent me few years back when she bought this license for me and emailed me all details. Now next excuse????? Should we come to solution ?





WHY YOU ARE NOT ANSWERING ME???????? @Sweetfunny @loopline

As you have been told already, we cant discuss someone elses license information with you. This is between you and whoever purchased the license from, you are not our customer so there's nothing we can do. You can either purchase a valid license, or keep trying to contact your friend.
 

You will need to get in contact with your friend who purchased the license. Obviously we cant discuss someone elses license information with you.




I dropped her email but she is not responding to me and I invite you to come on teamviewer I show you email sent me few years back when she bought this license for me and emailed me all details. Now next excuse????? Should we come to solution ?





WHY YOU ARE NOT ANSWERING ME???????? @Sweetfunny @loopline
As noted above, but also its just basic. In the world of the internet there is zero way to prove you didn't find this license info somewhere some way or hacked into something to get it and that your not trying to steal the license. They have had so many scrapebox liceses stolen over the years, it never ends well.

I know you may say you have emails or anything else, but anything can be hacked.

I mean if you were the license owner and someone else was trying to steal your license Im sure you would appreciate that scrapebox didn't just hand it over on a silver platter.

Anyway thats my 2 cents. I know people have tried to steal my licenses even. Fortunately they mailed me and asked if it was me. Which I appreciated.
 
Some other services that I have disabled that I'm wondering if they have any impact on proxies being used:
SSDP Discovery (There are no homegroups / network groups on my network. No file-sharing. I only use Teamviewer).
UPnP Device Host
IP Helper
Server
Workstation
TCP/IP NetBIOS Helper

Disabling random services like that will cause you trouble. Win8-10 isn't win98.... tweaking it does more harm than good a lot of the time.

Kinda weird that you've got 16k ports open from 50 threads.

You should try breaking your job into smaller pieces, then using the automator to work it's way through the files there.

That being said, I've noticed recent builds of SB seem to have issues with terminating threads after a job has been paused/stopped.
 
Sent an email about a bug, wanted to make sure you got it because it's killing my team's workflow.

When scraping emails by crawling sites and you select "save URL with email" it's no longer saving the URL with the email.

Can you please confirm this is a bug you're able to recreate? Is there any way for me to download a previous version until a patch is released? If so, what is the latest version that doesn't have this issue?

Support emailed me a dev version with the bug fix - excellent support. Still, the best SEO purchase i ever made.


I've noticed another issue that im not sure about (it's been around for a while tho).

When scraping yellowpages.com with several threads (like 20) and proxies it's as slow as (or slower) than scraping with a single thread with no proxies. It's like only one thread is running - am i missing something here? Any ideas? (the proxies are shared, clean, and fast)
 
Last edited:
Support emailed me a dev version with the bug fix - excellent support. Still, the best SEO purchase i ever made.


I've noticed another issue that im not sure about (it's been around for a while tho).

When scraping yellowpages.com with several threads (like 20) and proxies it's as slow as (or slower) than scraping with a single thread with no proxies. It's like only one thread is running - am i missing something here? Any ideas? (the proxies are shared, clean, and fast)
Well "technically" for part of the process only 1 thread is running.

So there are multiple steps involved with the yellow pages scraper due to how yellow pages works. Yellow pages constantly changes things, probably to break scrapers, and takes various measures to prevent their data from being scraped.

All told scrapebox has to do multiple things to get all the data. So part of the process uses the number of threads you have set and part of the process can only be single threaded due to how it works.

Mostly this is an advantage, because for most people it helps prevent yellow pages from blocking proxies as fast, as they block quite quickly. So in the lions share of cases it winds up letting people get more data then they otherwise would.

All in all though its a limitation of things being what they are.
 
Well "technically" for part of the process only 1 thread is running.

So there are multiple steps involved with the yellow pages scraper due to how yellow pages works. Yellow pages constantly changes things, probably to break scrapers, and takes various measures to prevent their data from being scraped.

All told scrapebox has to do multiple things to get all the data. So part of the process uses the number of threads you have set and part of the process can only be single threaded due to how it works.

Mostly this is an advantage, because for most people it helps prevent yellow pages from blocking proxies as fast, as they block quite quickly. So in the lions share of cases it winds up letting people get more data then they otherwise would.

All in all though its a limitation of things being what they are.

Ok figured there was something like that going on.

So what's optimal for speeding up YP scraping? Are there settings I can change if i use a beasty server to make things quicker for the single threaded part?
 
Ok figured there was something like that going on.

So what's optimal for speeding up YP scraping? Are there settings I can change if i use a beasty server to make things quicker for the single threaded part?
Short answer. No

Long answer, sort of. You can run an unlimited number of instances of Scrapebox (on windows) so you could just fire up more instances and divide your locations and/or keywords.

This also has other advantages like if one instance crashes others keep going.
 
Short answer. No

Long answer, sort of. You can run an unlimited number of instances of Scrapebox (on windows) so you could just fire up more instances and divide your locations and/or keywords.

This also has other advantages like if one instance crashes others keep going.

PM Sent
 

Replied.

hi

I want ask something. I am trying Grab/Check module with before after method.

For example, I want scrape Youtube View (https:// www.youtube.com/watch?v=26J0K26GxCg).

http://prntscr.com/jjcz57
http://prntscr.com/jjd0wr

The source code for this view
http://prntscr.com/jjczii

I always get 500 error...
http://prntscr.com/jjczs9

I have tried to remove user agent setting, but the result is still same.

Do you have any idea how to fix this?

thanks

Thats a server error. Does it give you that when you don't use proxies? usually google services will return a 503 if its an ip block.

But if I may suggest, don't reinvent the wheel here.

Scrapebox already has a youtube downloader addon thats free with scrapebox and it scrapes the views of videos and a host of other information as well. So I would personally just use that.

 
Status
Not open for further replies.
Back
Top