Permanently Closed Marketplace Sales Thread

Status
Not open for further replies.
where is the op?
is it support windows 7?

the op already said she has been having some difficulties with bhw for the past few days as well as being unwell recently....

yes it supports windows 7

problems?
 
it seems the rapid indexer only supports 1k links at a time ? it there any plans of extending this
 
it seems the rapid indexer only supports 1k links at a time ? it there any plans of extending this

Well earlier today I did a run with over 30k links I wanted to ping, no problem.

Unless of course you mean over 1k of your own website links.
 
I've been experiencing a lot of lockups since the last patch. I'm running scrapebox within parallels 5 on windows xp.

Anyone else experience anything similar (long shot, i know).
 
Man, that sucks...
I thought the Harvester is realiable when he tells you that he got all results. But its not the case.

Im using free proxies for scraping and used them on the keywords:

intitle:"Pligg beta"
inurl:"register.php"++"powered by pligg"

Because they are free proxies it mostly doesnt work. But I did it again and again. So I got results some times.
The thing is the first search is around 100 results in normal google but the second one is around 1k.
But with SB I got for both keywords results of 100 urls and the button in the resultlist was green after that. One time there was even 0 Result and the button was green.
So the check if a keyword is fully scraped is broken. Disappointing.
 
I'm having a problem harvesting forum urls. I'm using a footprint to seek out vbulletin forums for my keywords.

The problem is that it's only returning under 100 results, and the forums it's picking are not remotely related to my keywords. I've tried both targeted and broad keywords (and I've tried just 1 keyword as well as several).

To give you an example it's returning shirley mclean (and overclocking) forums for the keywords "earth" and "green energy".

I've checked in google and there should be 100s of targeted forums. Any ideas what I'm doing wrong (probably something simple I know!)?
 
Man, that sucks...
I thought the Harvester is realiable when he tells you that he got all results. But its not the case.

Im using free proxies for scraping and used them on the keywords:

intitle:"Pligg beta"
inurl:"register.php"++"powered by pligg"

Because they are free proxies it mostly doesnt work. But I did it again and again. So I got results some times.
The thing is the first search is around 100 results in normal google but the second one is around 1k.
But with SB I got for both keywords results of 100 urls and the button in the resultlist was green after that. One time there was even 0 Result and the button was green.
So the check if a keyword is fully scraped is broken. Disappointing.


I'm having a problem harvesting forum urls. I'm using a footprint to seek out vbulletin forums for my keywords.

The problem is that it's only returning under 100 results, and the forums it's picking are not remotely related to my keywords. I've tried both targeted and broad keywords (and I've tried just 1 keyword as well as several).

To give you an example it's returning shirley mclean (and overclocking) forums for the keywords "earth" and "green energy".

I've checked in google and there should be 100s of targeted forums. Any ideas what I'm doing wrong (probably something simple I know!)?

Please try the latest 1.14.18
 
Hi,
Is it possible to tell screpebox to harvest pages in specific language only? There is such option when you use non-english google (a checkbox below input field, see the image
14qn8w.png
).
I think it could be turned on by just one parameter in GET request.
 
Man, that sucks...
I thought the Harvester is realiable when he tells you that he got all results. But its not the case.

Im using free proxies for scraping and used them on the keywords:

intitle:"Pligg beta"
inurl:"register.php"++"powered by pligg"

Because they are free proxies it mostly doesnt work. But I did it again and again. So I got results some times.
The thing is the first search is around 100 results in normal google but the second one is around 1k.
But with SB I got for both keywords results of 100 urls and the button in the resultlist was green after that. One time there was even 0 Result and the button was green.
So the check if a keyword is fully scraped is broken. Disappointing.

I won't go in to all the technical details, but the multi-threaded harvester uses a special low bandwidth Google search page it does not tell you "Results 1 - 10 of about 5,000,000" nor does it have page numbers from 1-10 at the bottom, it doesn't even have descriptions, cache links etc under the URL's.

Example: http://i39.tinypic.com/2v84thz.png

It's hard to tell what's the "end" of the results, especially when your using free proxies constantly failing giving the impression one of the (up to) 500 connections has reached end of results.

The multi-threaded harvester is not, i repeat NOT a precision harvesting tool.. It goal is simple, to tear the heart of out Google and pull down the most URL's in the shortest timespan using the least bandwidth possible. It doesn't care about your free proxy that crapped itself on a pageload, it doesn't care you were expecting x,xxx results for a specific query and you got a few hundred less than that...

So your "problem" is simply a case of wrong tool for the wrong job, if you have free proxies prone to failure and every queries results are critical use the single threaded harvester. If you don't give a damn and simply want a million URL's while you make coffee, multi-thread it.

Hi,
Is it possible to tell screpebox to harvest pages in specific language only? There is such option when you use non-english google (a checkbox below input field, see the image
14qn8w.png
).
I think it could be turned on by just one parameter in GET request.

Yes sure you can do that. In the Custom Googles, you can add the extension plus the hosts language like this:

2uj2p2t.png


To get the letter code for a language, do a search on Google when ticking the language box and look in the URL string for hl=XX

2a0ewep.png


The hl= parameter means "Hosts Language".
 
Man, that sucks...
I thought the Harvester is realiable when he tells you that he got all results. But its not the case.

Im using free proxies for scraping and used them on the keywords:

intitle:"Pligg beta"
inurl:"register.php"++"powered by pligg"

Because they are free proxies it mostly doesnt work. But I did it again and again. So I got results some times.
The thing is the first search is around 100 results in normal google but the second one is around 1k.
But with SB I got for both keywords results of 100 urls and the button in the resultlist was green after that. One time there was even 0 Result and the button was green.
So the check if a keyword is fully scraped is broken. Disappointing.

Try harvesting wit the multi harvester turned off.
 
sweet is there any chance of extending the about of your websites you can do at once using rapid indexer ? i have a 33k list and only supports 1k
 
Sweet.

Any chance of a link do follow checker add on any time soon?
 
Last edited:
Please try the latest 1.14.18

Nope... I used the keyword:

intitle:"Pligg beta"

So i went to google and checked this key directly in the browser. It has 355 Results in google.com and in my countries google that are available.

So I went to SB and tried it with public proxies. After some tries I managed to get results. It got 100 results and stopped. And in the resultpage the keyword was green. That means it could get all results. But thats not the truth.

So i tried to harvest them without a proxy and surprisingly now it found 898 results. I only had google.com in use.

So i dont understand this all. But the resultpage doesnt tell the truth for sure.

It shouldnt be a problem to find out if a resultpage at google is fully downloaded or if it is a resultpage or the proxy delivered something else. It cant be green when it only managed to get one resultspage when there are more than that...
 
No, the GREEN keyword does not mean it got all results you wanted, it only means that it has done using that keyword, because it did not get any more results.

I just did a test using private proxies, with your keyword, I wanted 400 results, and got 400 results...

You should not blame SB when in fact your proxies are the cause.
 
sweet is there any chance of extending the about of your websites you can do at once using rapid indexer ? i have a 33k list and only supports 1k

There is no 1k limitation, it's a little over 1 Million:

195h8i.png


Visiting over 1 Million webpages is quite substantial, even running that on a 100MBit server will take a long time.. I'm not real sure you could call it a limitation.

Sweet.

Any chance of a link do follow checker add on any time soon?

The problem is predicting what a link will be isn't real accurate and when something is correct only most the time people complain about it like trying to predict the Posted and Posted, Moderated status ability. That generated so many emails it's been removed now, so until the do/no follow checker is accurate enough that i don't have to spend all my days explaining why it detected something incorrectly it won't be put in unfortunately.

So I went to SB and tried it with public proxies. After some tries I managed to get results. It got 100 results and stopped. And in the resultpage the keyword was green. That means it could get all results. But thats not the truth.

So i tried to harvest them without a proxy and surprisingly now it found 898 results. I only had google.com in use.

So i dont understand this all.

Did you read what i said?
confused.gif


Do not multi-thread on unreliable free proxies if accuracy of results is important to you. Use the single threaded harvester.

Bulk Fake PageRank Checker Addon

ega9lf.jpg


The addon will detect if a list of domains have Fake Pagerank, this happens a lot with TDNAM domains so this addon is useful to use in conjunction with that addon before you go buying that PR7 you think you scored for $5. :D
 
I think no one thought until now that the green light could state that there are more results. And thats a really important info in my opinion. Its not only that you want to use it for scraping blog-urls but when you have the target to scrape the backlinks of a site or something like that then of course you want all the results because its important. When it only comes to blog commenting its not that important, thats true.

So you say when I use the normal harvester, not the multithreaded one, then I would really get all results and it doesnt use the same googleresultpage where you cant determine the end of the page?

You said you used your private proxies... thats fine, but most users are using private proxies only for commenting. Not even thinking that there could be a problem. And even with private proxies I think when your proxyurl is banned in google in the middle of the working of one keyword, maybe at resultpage 5, the led would be green without having all results. So I dont think that private proxies would be a solution. Especially when you have big lists of keywords... I mean how many private proxies would you need to scrape all these keywords where each maybe results in around 8 requests to google? Google is banning ips really fast.

By the way... will there be a fix for the link checker? I mean it makes no sense to check links when the result isnt trustworthy.
At the moment it seems it only checks if the domains that are checked for are in the link somewhere. Or you can check to check on the rootdomain only.
How about adding a text field where you can set something like
^%domain%$
so that the link really is the link that you are checking for? This way you could take away the $ and you would get all links that are deeplinks too. And so on. But at the moment it finds links that arent real links because they only are hidden in other links.
By using these textlinks forwardslashes needs to be trimmed at the domain to be checked and the checked links. And when the user entered the domain only in the form
domain.com
without http:// the http:// would be in need to add automatically in order to work together with ^.

@SF

>Did you read what i said?

Sorry didnt see that there was another page in the thread. So I will use the normal harvester when there this problem doesnt appear.
 
Status
Not open for further replies.
Back
Top