Lets discuss ScrapeBOARD. Questions, comments, thoughts, Videos etc..

Status
Not open for further replies.
Thanks for the reply guy's
I think i will give it a go, for $97 it's not a big investment.
 
Yes, they still help, they may not be what they were a few years ago, but they are still helpful and increase SERP.

What you can do to follow up is blast your forum profiles with some blog comments from scrapebox and that is the best way to get them indexed. At the same time it will give your profiles some link juice which will all flow to your money page.

would be wonderful if this had the decaptcha like xrumer integrated..how can it be that just these guys have automatic decaptcha?
 
question:

When i create accounts i get for ex.:

3 000 accounts activaed
6 000 email activated accounts
70 moderated

When scrapeboard downloads activation emails and click on the links, it sayd i have 6 000 activated accounts. The file with accounts also has 6 000 accounts.

Shouldn't it be 3 000 already activated accounts + 6 000 from email, so around 9 000 acounts in total in the file ??
 
Last edited:
UPDATE : Ok, forget... Without proxies (however fresh) it's works.

For me, ScrapeBox and ScrapeBoard URL hervester dosen't works today...

Where can this problem arise ?
Someone in the same case ?

Beny
 
Last edited:
Just bought this earlier this evening.

I've scraped some keywords and then used that to scrape over 10K urls. After trimming out duplicates I end up with 250.

I then ran these through the Forum analyzer and when I try to export the results filtering out 'non public profiles' the text file I export to is blank.

Am I doing something wrong?
 
I for one love the idea of an integrated Captcha breaker. The problem with Scrapeboard is that it's not possible to train the breaker in an easy way.

I've been looking into how I would combine the AntiReCaptcha (arc.traineddata) code with the current eng.traineddata and just can't find a good answer. I'd love it if:
1) It could be integrated into SBoard, if it isn't already (jDownloader has this library and trained data)
2) The possibility to train with the Tesseract tools available. This could come in handy both for one self and to eventually crowd-source the data, thus building a really neat captcha breaker.

Then comes the request for the text-captcha breaker. Not everything is possible to do here (example: what is this sites name?) from a generic answer, and that's not what I'm really looking for either. But it would be great to be able to edit, add and delete answers. This feature, coupled with the crowd-sourcing already in place, would make Scrapeboard the number one champion.

If anyone knows how to merge, train and fix the tesseract trained data for several captcha types - please reply!
 
 
I don't know how to do what your asking. What I can say is that given the nature of what scrapeboard is and its competitors, that a learning mode is likely something that will come at some point. However I think its a better use of their time to teach it to solve captchas, and then when things are good there and lots of other things in place, then do learning mode.

I can say that the developer is top notch and isn't going to do this project half way, but it is going to take time. That doesn't help you today, but be patient, scrapeboad has gotten nothing but better, and it will continue to do so.

Thanks for the answer. I'm not expecting miracles in the near future, I see SBoard as a fun tool that's actually useful. Have totally massive scraped lists for forums that I just let it burn through.

Regarding the progression - I would actually say that I'd probably crowd-source as much as I can. The benefits are several.

The most obvious is of course that you can let others "work" for you. Let them train the captcha breakers while you make all the fun features that keeps you motivated.

The second reason is for the users themselves. Lots of people bought SBoard instead of Xrumer and several will have done so because of a tight budget. If you just give them a small carrot, they'll put work into captchas and get rewarded for doing so.
This will of course also raise a buzz - if it's locally trained then people will start sharing, buying etc and get even more users to SBoard.
If it's a public crowd-sourced training then the buzz will come from having a really good captcha breaker.

Of course, how does one protect the investment with an open database? I see gazillions of captcha-breaking services rising from a crowdsourced database.

Ah well, just my two crowns!

By the way, does anyone have good trained response files to share? Mine are... well, under average. I think that comes to play with SOME of the missed activation emails.
 
Just bought today, out of 2000, 1 got activated. I manually did the captcha etc,

Most common reason: unknown platform and errors, such as 404, etc. Other reasons were like non public profile (maybe less 30)

Matt: there really needs to be more info in the online help guide. On the video they were helpful dont get me wrong, but certain parts you seem to rush it, and when it is written it lot easier to refer to the part I need, instead having to fast forward, plus you can get more info down.
 
Last edited:
Just bought today, out of 2000, 1 got activated. I manually did the captcha etc,

Most common reason: unknown platform and errors, such as 404, etc. Other reasons were like non public profile (maybe less 30)

Matt: there really needs to be more info in the online help guide. On the video they were helpful dont get me wrong, but certain parts you seem to rush it, and when it is written it lot easier to refer to the part I need, instead having to fast forward, plus you can get more info down.

Well mate I don't write the online help. I just do videos. You are correct, I do tend to talk fast. Its not that I am trying to rush it, its just kind of how I am. :)

I am going to be redoing all the videos, probably a couple of times yet. Since its beta I wanted to get some videos done, but so much changes so fast I could do videos this week and next week do them again.

Right this minute I can't do any videos as I had surgery 2 weeks ago and then again just monday of this week. So when I talk for more then about 3 seconds you can hear the pain in my voice and that would make for bad videos. Plus videos take a fair bit of concentration and preparation and I just don't feel up to it at the moment.

Also bear in mind the developer is writing the help file. They are spending their time on making a great product right now, so they can officially launch the program at some point. So I am sure after they get a lot of that done they will be adding more info to the help file. Its all still a work in progress.

I will commit to slowing down though, in future videos anyway. :)

Thanks,
MAtt
 
Hey Matt,
Can you give any insight into future features? I'm sure there's a punchlist that the dev is building too, and I know there hasn't been an update in a while, so I'd be curious to see what they're working on to make it better.
 
I for one love the idea of an integrated Captcha breaker. The problem with Scrapeboard is that it's not possible to train the breaker in an easy way.

I've been looking into how I would combine the AntiReCaptcha (arc.traineddata) code with the current eng.traineddata and just can't find a good answer. I'd love it if:
1) It could be integrated into SBoard, if it isn't already (jDownloader has this library and trained data)
2) The possibility to train with the Tesseract tools available. This could come in handy both for one self and to eventually crowd-source the data, thus building a really neat captcha breaker.

Then comes the request for the text-captcha breaker. Not everything is possible to do here (example: what is this sites name?) from a generic answer, and that's not what I'm really looking for either. But it would be great to be able to edit, add and delete answers. This feature, coupled with the crowd-sourcing already in place, would make Scrapeboard the number one champion.

If anyone knows how to merge, train and fix the tesseract trained data for several captcha types - please reply!

Cool to have someone who knows what they are talking about :p

Just from my personal understanding, and me using it, Tesseract has severe memory leaks. Whether they are so bad that it is obvious, or just slow leaks over time, I do not know. I have only ever used tesseract in a .net assembly, and I did not notice it that bad, but I was doing only small tests.

Next, you need to understand that tesseract is only ever used for reading PDF's or straight images. Not captchas. You will need to code the cleaning of the image and straighten out the letters etc before passing it to tesseract. You can't just throw a random captcha in there (Either recaptcha or any other) and expect it to read it (Because it won't).

A few more things. Tesseract was a tonne more accurate when splitting letters up in a captcha, and feeding the letters individually. Trying to split letters up in a Recaptcha sense can be difficult, and is not always 100%. Not a deal breaker, but your success rates increase 10 fold.

And lastly. Although this could have been in the implementation I was using. Tesseract seemed "Random". 1/20 solves on the exact same captcha would return a wrong result. If I ran it again, it would come back with the correct result. I have no idea what causes it, but it certainly isn't 100%.
 
ScrapeBoad does cleaning up images, straighten text and split the image into individual captchas before they are feed to tesseract. So, training tesseract might not help at all.

@Pyronaut: About your random results, do you save the image uncompressed or use jpeg? We do not get random results, every time we test the same image, the result is always the same. We use .jpg images, but without compression (quality 100%).
 
Hey Matt,
Can you give any insight into future features? I'm sure there's a punchlist that the dev is building too, and I know there hasn't been an update in a while, so I'd be curious to see what they're working on to make it better.

Hey mate, I don't work for scrapebox/board, I simply maintain a "quality" relationship with the dev team. I make it a point to not go around saying anything i am told. If they want it public, they will share it. Not that I have a checklist anyway, but I am fully confident that they have the application well planned out and will do everything they can to make scrapeboard and excellent application.
 
ScrapeBoad does cleaning up images, straighten text and split the image into individual captchas before they are feed to tesseract. So, training tesseract might not help at all.

@Pyronaut: About your random results, do you save the image uncompressed or use jpeg? We do not get random results, every time we test the same image, the result is always the same. We use .jpg images, but without compression (quality 100%).

They were directly downloaded SMF captchas. Of which, my app then turned them into "bitmaps" in memory (Just the way it is in .net).

If Scrapeboard is already cleaning images, then you don't need to "train" it at all. The english language pack it comes with read SMF captchas just fine after cleaning it up a bit.
 
Ok then the only thing what comes in my mind is to delete scrapeboard folder. Not sure what windows are you using but you could try following (For win 7)

Open up control panel
Open up folder options
Click on tab view
Tick into "Show hidden files and folders"
Click ok

Now navigate to C:\Users\YOURUSERNAME\AppData\Roaming
Delete Scrapeboard folder

Create a folder named Scrapeboard to somewhere on your harddrive
extract files from scrapeboard.rar (which you downloaded from scrapeboard website) to Scrapeboard folder and try to run it again and check are you able to register now.

thanks mate :D
 
Hey mate, I don't work for scrapebox/board, I simply maintain a "quality" relationship with the dev team. I make it a point to not go around saying anything i am told. If they want it public, they will share it. Not that I have a checklist anyway, but I am fully confident that they have the application well planned out and will do everything they can to make scrapeboard and excellent application.

I am waiting to buy but when it is complete. Will you ask developers how long it could take to be full functional.
 
Well nobody answered my question in the thread I created so I'm quoting it here :

Hi,

I purchased SBoard last week and played with it yesterday. I harvested public proxies, tested them all ("Test against Google" checkbox checked), scraped around 10k forums urls and let it run overnight. This morning I had around 200 accepted profiles (no manual capchas, I wanted to see if the internal decapcher was any good). I thought it was not bad for a first try!...

...until I got a mail from my ISP telling me to stop spamming or my connection will be shut and a message showing MY IP attacks. :/

What did I do wrong? I noticed the settings in Sboard say "Use proxies when available" (it's checked of course), does this mean that my proxies died in the night and Sboard switched to my personal IP for posting?

Thanks in advance for your help, I'm a lil confused.

Hope you guys can enlighten me.
 
sbox used your proxies. But, before you connect to any proxy, you will pass through your ISP of course.

YOU -> YOUR ISP/GATEWAY -> PROXY -> TARGET SITE

So, your ISP see that the traffic is coming from you.
 
Status
Not open for further replies.
Back
Top