Permanently Closed Marketplace Sales Thread

Status
Not open for further replies.
i was always talking about custom harvester while scraping urls from search engine... pause button is not pausing ... just stays on please wait
also when i exprrted after meta extraction.. exported file dosnt work please check that too
 
No the latest is v2.0.0.26

My mistake, for some reason I thought we were talking about the link extractor addon and not the core Scrapebox. Someone must have switched out my coffee with decaf . :)

Unfortunately, I'm still having problems with Link Extractor 2.0.0.31

It seems to crash on me constantly when trying to run through a big list (1.5 million urls currently) with a large number of connections. I switched back to .27 and it works fine again. This is running with the option to treat subdomains as external links.

Since we are talking about the link extractor in this post, ( :) ) There is a .32 version out, try that.
 
ScrapeBox v2.0.0.27 Beta Released:

  • Fix a bug where the proxy reload time was ignored
  • Removed the "Preparing proxy refresh" dialog which popped up during harvesting
  • Fixed a bug in poster related to captcha settings which prevented the captcha solver to be invoked
  • Fixed a bug in Meta scraper
  • New 6 new search engines added
  • Fix Several search engines enhanced.
  • Fixed a bug in RSS Ping
  • Improved domain lookup (Grab/Check button)
  • Fixed a bug in Email Grabber when called from Automator
  • Fixed Wordpress.com Poster in Article Scraper

i was always talking about custom harvester while scraping urls from search engine... pause button is not pausing ... just stays on please wait
also when i exprrted after meta extraction.. exported file dosnt work please check that too

With the harvester, you may need to leave it longer for all threads to terminate. I just tried it several times, and it worked each time. Same with the meta scraper problem, i just used all 3 export options and all worked fine. What issue are you having exactly?
 
Ok, I figured out the issue I'm having with Link Extractor 2.0.0.32 - I'm running it with a high number of threads and it seems to not take my limit and keeps adding more and more threads until it crashes. For example, I set it to run at 2000 threads normally, but it grows to 8000 and then the whole link extractor crashes. This doesn't happen in 2.0.0.27. I realize running with a high thread count isn't supported but it does work that way usually.


 
When using SB V2 (32 bit) Custom Harvester sometimes SB just closes with no message, warning etc..

When i return the SB icon is in the notification area of the taskbar and when i hover over it the icon just disappears.

I don't sit at the laptop when SB V2 is running so not sure why, it seems to be different times and can be as low as 100k harvested URL's or 2 Million+ so the number of URL's harvested doesn't seem to be a problem.

Also happens if i run at any amount of threads.

Doesn't happen every time but probably around 7/10.

I can retrieve the harvested urls from Harvester_Sessions so not a major problem but a little annoying when i leave it running for 12 hours+ and return to find it happened at say 100k URLs lol

Don't get chance to send a bug/error report so posting here.

Also, the link extractor plugin doesn't seem to collect as many URLs with SB V2 compared to V1 using the same list, proxies and connections.

Thanks
 
Ok, I figured out the issue I'm having with Link Extractor 2.0.0.32 - I'm running it with a high number of threads and it seems to not take my limit and keeps adding more and more threads until it crashes. For example, I set it to run at 2000 threads normally, but it grows to 8000 and then the whole link extractor crashes. This doesn't happen in 2.0.0.27. I realize running with a high thread count isn't supported but it does work that way usually.

Do you have the "Treat subdomains as external links" enabled? If so try turning it off. But you pretty well said it, windows doesn't like that many concurrent threads. If all else fails you could split the list, open multiple instances of the link extractor and then use lower connections on each.

When using SB V2 (32 bit) Custom Harvester sometimes SB just closes with no message, warning etc..

When i return the SB icon is in the notification area of the taskbar and when i hover over it the icon just disappears.

I don't sit at the laptop when SB V2 is running so not sure why, it seems to be different times and can be as low as 100k harvested URL's or 2 Million+ so the number of URL's harvested doesn't seem to be a problem.

Also happens if i run at any amount of threads.

Doesn't happen every time but probably around 7/10.

I can retrieve the harvested urls from Harvester_Sessions so not a major problem but a little annoying when i leave it running for 12 hours+ and return to find it happened at say 100k URLs lol

Don't get chance to send a bug/error report so posting here.

Also, the link extractor plugin doesn't seem to collect as many URLs with SB V2 compared to V1 using the same list, proxies and connections.

Thanks

In your main SCrapebox folder there should be a file called bugreport.txt. If you send that over to support they can shed more light on the subject. Assuming it gets created anyway, which it will in most cases.

scrapeboxhelp[at] gmail (dot} com

or

support (at] scrapebox [dot) com

As for the link extractor, do you have the treat subdomains as external links enabled? Are you extracting external or internal or both?

If you send support a small list of urls that produces different results in the 2.0 vs 1.x versions then they can reproduce the issue and adjust whatever needs fixed.
 
there is no problem with exporting.. but the problem is with opening the exportd file.. it says file is corrupted and wont open .. it gives option to repair so if we repair... there is no data inside
ScrapeBox v2.0.0.27 Beta Released:

  • Fix a bug where the proxy reload time was ignored
  • Removed the "Preparing proxy refresh" dialog which popped up during harvesting
  • Fixed a bug in poster related to captcha settings which prevented the captcha solver to be invoked
  • Fixed a bug in Meta scraper
  • New 6 new search engines added
  • Fix Several search engines enhanced.
  • Fixed a bug in RSS Ping
  • Improved domain lookup (Grab/Check button)
  • Fixed a bug in Email Grabber when called from Automator
  • Fixed Wordpress.com Poster in Article Scraper



With the harvester, you may need to leave it longer for all threads to terminate. I just tried it several times, and it worked each time. Same with the meta scraper problem, i just used all 3 export options and all worked fine. What issue are you having exactly?
 
Do you have the "Treat subdomains as external links" enabled? If so try turning it off. But you pretty well said it, windows doesn't like that many concurrent threads. If all else fails you could split the list, open multiple instances of the link extractor and then use lower connections on each.

In your main SCrapebox folder there should be a file called bugreport.txt. If you send that over to support they can shed more light on the subject. Assuming it gets created anyway, which it will in most cases.

scrapeboxhelp[at] gmail (dot} com

or

support (at] scrapebox [dot) com

As for the link extractor, do you have the treat subdomains as external links enabled? Are you extracting external or internal or both?

If you send support a small list of urls that produces different results in the 2.0 vs 1.x versions then they can reproduce the issue and adjust whatever needs fixed.

Thanks, yes i have a bug report, the last one was created today at 00:28 (UK) and it's now 08:40 and i set it running at around 10:00 last night.

Will include as much info as possible in the email.

As for the link extractor i had ''treat subdomains as external links'' enabled so i will double check the settings are equal for both sbv1 and sbv2.
 
Ok, I figured out the issue I'm having with Link Extractor 2.0.0.32 - I'm running it with a high number of threads and it seems to not take my limit and keeps adding more and more threads until it crashes. For example, I set it to run at 2000 threads normally, but it grows to 8000 and then the whole link extractor crashes. This doesn't happen in 2.0.0.27. I realize running with a high thread count isn't supported but it does work that way usually.



First of all, how do you run it with 8000 threads when it only allow to enter up to 200, by screwing with the .ini file? It is not supposed to run with that amount of threads.
Then, I would be happy when you could provide a video of it running, so I can see in detail what it does. Can you do that?
 
I have a trouble with importing on domain level - in new sb it's not working as in the old version - the issue is that when in the file a domain is with http and in scrapebox it's not, sb does not removing it which is a mistake as it's really the same domain.
 
I have a trouble with importing on domain level - in new sb it's not working as in the old version - the issue is that when in the file a domain is with http and in scrapebox it's not, sb does not removing it which is a mistake as it's really the same domain.

This has been fixed and will be in the next beta release (2.0.0.28). Thanks for reporting the issue.
 
Thanks, yes i have a bug report, the last one was created today at 00:28 (UK) and it's now 08:40 and i set it running at around 10:00 last night.

Will include as much info as possible in the email.

As for the link extractor i had ''treat subdomains as external links'' enabled so i will double check the settings are equal for both sbv1 and sbv2.

That feature isn't in 1.x, so you could try unchecking it and just see what happens, but Im sure support will fix you up.



~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~

2.0 Advanced Proxy Manager Guide

[video=youtube_share;C9XaIHhSyks]http://youtu.be/C9XaIHhSyks[/video]
 
Is it possible to scrape Etsy.com sellers emails with scrapebox?
 
The new proxy manager seems a step forward. Still I guess without automation it's not handy to use. Can premium automator plugin do those actions:

1. Export certain proxy (after filtering) to a file.
2. Do regular proxy check.
3. Remove bad sources and load new from a file?
4. Most important - can scrapebox during harvesting also load such filtered proxy from a file/api?
 
can i buy now a licence and change the licence to my new computer in 20 days?
 
Is it possible to scrape Etsy.com sellers emails with scrapebox?

Can you give some sample urls that have emails on them?

I mean the email scraper can scrape from any page, so long as it meets these critera:

The page does not require authentication
The email is in the html of the page and not wrapped or displayed by a script such as javascript etc...


If it meets those criteria then generally its fine, but Im sure there are exceptions out there with something I missed, but if you can provide a few sample pages, then I can answer that more accurately.

The new proxy manager seems a step forward. Still I guess without automation it's not handy to use. Can premium automator plugin do those actions:

1. Export certain proxy (after filtering) to a file.
2. Do regular proxy check.
3. Remove bad sources and load new from a file?
4. Most important - can scrapebox during harvesting also load such filtered proxy from a file/api?

Yes to #1
Yes to #2
No on number #3. But how would you qualify "bad" sources? Also how would you come up with more to add on the fly?
#4 Yes. You can have scrapebox either load new proxies from a file when all proxies die or load from a file every X minutes.


can i buy now a licence and change the licence to my new computer in 20 days?

I think so. When you purchase you get to activate and then you would still have your 1 free transfer for Feb and then on Mar 1 you get another transfer.
 
Here is a page on etsy https://www.etsy.com/people/NynneRosenvinge
alos i found this footprint to use inurl:https://www.etsy.com/people "@gmail.com"

Can you give some sample urls that have emails on them?

I mean the email scraper can scrape from any page, so long as it meets these critera:

The page does not require authentication
The email is in the html of the page and not wrapped or displayed by a script such as javascript etc...


If it meets those criteria then generally its fine, but Im sure there are exceptions out there with something I missed, but if you can provide a few sample pages, then I can answer that more accurately.
 
No on number #3. But how would you qualify "bad" sources? Also how would you come up with more to add on the fly?

Sb harvester of course for new sources:) As we have there source qualifier we could drop sources that have like under 10% alive over few checks eg.

Still one more thing bothering me - maybe automator can do it - how to run proxy grabber and checker in the bacground - when it's on top I cannot do anything in main sb window.
 
Last edited:
Here is a page on etsy https://www.etsy.com/people/NynneRosenvinge
alos i found this footprint to use inurl:https://www.etsy.com/people "@gmail.com"

Can you give some sample urls that have emails on them?

I mean the email scraper can scrape from any page, so long as it meets these critera:

The page does not require authentication
The email is in the html of the page and not wrapped or displayed by a script such as javascript etc...


If it meets those criteria then generally its fine, but Im sure there are exceptions out there with something I missed, but if you can provide a few sample pages, then I can answer that more accurately.

I pmed you back. But that url you listed won't work with the email grabber, because the email is not a proper email, but is typed the way it is specifically so scrapers won't find it.

You would use the page scanner and some regex to identify pages that had emails on them like that, but it wouldn't scrape them, you would have to get someone on fiverr or elance or whatever and go thru all the pages manually and copy out the emails and save them off.

Sb harvester of course for new sources:) As we have there source qualifier we could drop sources that have like under 10% alive over few checks eg.

Still one more thing bothering me - maybe automator can do it - how to run proxy grabber and checker in the bacground - when it's on top I cannot do anything in main sb window.

No, you can't use the automator to use the source qualifier at this time. But what Im getting at is, where are you going to get your new sources to add in to be tested?
 
Hi, does anyone else experiencing problems with the Key Word scraper in V 2? i had to go back and use V1 kw scraper, my apologies if it has been mentioned before!

Loopline! awesome tutorials my man! keep up the good work!
 
Status
Not open for further replies.
Back
Top