Permanently Closed Marketplace Sales Thread

Status
Not open for further replies.
Hello :)
I have a problem guys, i have a bunch of proxies when i scrape google everything is good after 100 000 link it gives me just errors. Proxies not getting ban but i think it requires chacpa enterence. How can i solve that ? Can i enter manually or can i use CB ? Many thanks.

Well unless the proxies are just flat out dieing totally, then they are getting banned. If it requires captcha then its blocked, but there is no way to enter the catpcha, Captcha Breaker wouldn't solve them anyway and even if you could they become blocked extremely quickly again. This was tried in the past, but it just doesn't work out as well as it sounds like it might.

Are you uisng public proxies or private?

Whats the backlink extractor like now on this thing? :)

It uses MOZ, so you need a moz account and you can get upto 1000 backlinks. Actually V1 uses this as well as V2.
 
Hey loopline and Sweetfunny, could you please inform if any other user agent works for the Google search options except the default one out there i.e. the Opera 8.50 one? I had tried several user agents but when I was trying the test engine, it was not working and not pulling up any results. When I changed back to the Opera 8.50 one it starts working again. Any idea if and what other user agents also work with the Google search engine configuration?
 
Hey loopline and Sweetfunny, could you please inform if any other user agent works for the Google search options except the default one out there i.e. the Opera 8.50 one? I had tried several user agents but when I was trying the test engine, it was not working and not pulling up any results. When I changed back to the Opera 8.50 one it starts working again. Any idea if and what other user agents also work with the Google search engine configuration?

Have you tried (typing/testing) mobile agents?
If not working / is "dead", discard and [FONT=arial, courier new, courier, 宋体, monospace, Microsoft YaHei]re-type.
[/FONT][FONT=arial, courier new, courier, 宋体, monospace, Microsoft YaHei]I use the default one.[/FONT]
 

Attachments

  • UA.PNG
    UA.PNG
    234.6 KB · Views: 59
Hey loopline and Sweetfunny, could you please inform if any other user agent works for the Google search options except the default one out there i.e. the Opera 8.50 one? I had tried several user agents but when I was trying the test engine, it was not working and not pulling up any results. When I changed back to the Opera 8.50 one it starts working again. Any idea if and what other user agents also work with the Google search engine configuration?

It certainly can use any useragent Google supports, you have control over everything like cookies, header data, referrer etc not with just Google but all ~30 engines.

Go to Settings >> Harvester Engines Configuration.
ZJbgDYB


When you change the useragent, Google search results layout may change. The current one is fine and uses a fairly low bandwidth results page, i just done a test and was scraping Google at 366,408 urls/min using only 70 connections which is extremely fast. http://i.imgur.com/LPQdHEH.png

So just keep in mind if you want to change it for whatever reason, the user agent you use may return a page with tons more javascript/bloat that's going to be slower. Once you changed the useragent click "Test Engine" and see what it scrapes, if nothing save the raw html http://i.imgur.com/SpFgF3q.png

Now you need to open the saved webpage in Notepad, and check the html used to display the results and update the "Just before the URL" and "Right after the URL" in the settings in my first screenshot. In other words if you change the useragent, you may need to also tell ScrapeBox how to find the link in the new results page too.

Hello :)
I have a problem guys, i have a bunch of proxies when i scrape google everything is good after 100 000 link it gives me just errors. Proxies not getting ban but i think it requires chacpa enterence. How can i solve that ? Can i enter manually or can i use CB ? Many thanks.

As Loopline mentioned, when you enter the captcha the proxy will get blocked again very quickly. We had a feature in ScrapeBox before to unblock the proxies by entering a Google captcha but it wasn't useful, nobody used it and it was removed.

But one thing to mention for others as well, ScrapeBox isn't just a "Google Scraper" there's tons of other search engines you can scrape from to get 10 times more use from your proxies. DeeperWeb for example i was just harvesting at 500,000 urls/min and they use Google powered results. So don't limit yourself to just Google and Google passed proxies when there's so many options.
 
It certainly can use any useragent Google supports, you have control over everything like cookies, header data, referrer etc not with just Google but all ~30 engines.

Go to Settings >> Harvester Engines Configuration.
ZJbgDYB


When you change the useragent, Google search results layout may change. The current one is fine and uses a fairly low bandwidth results page, i just done a test and was scraping Google at 366,408 urls/min using only 70 connections which is extremely fast. http://i.imgur.com/LPQdHEH.png

So just keep in mind if you want to change it for whatever reason, the user agent you use may return a page with tons more javascript/bloat that's going to be slower. Once you changed the useragent click "Test Engine" and see what it scrapes, if nothing save the raw html http://i.imgur.com/SpFgF3q.png

Now you need to open the saved webpage in Notepad, and check the html used to display the results and update the "Just before the URL" and "Right after the URL" in the settings in my first screenshot. In other words if you change the useragent, you may need to also tell ScrapeBox how to find the link in the new results page too.


Thank you for the amazing explanation. When I changed the user agent, it was not bringing out any result when I clicked on test the engine and was just showing me option to show raw html. Perhaps I would need to check again and try to use it with other engines again. Will changing it to some other user agent reduce the chances of IP bans in Google or would it be same?
 
Thank you for the amazing explanation. When I changed the user agent, it was not bringing out any result when I clicked on test the engine and was just showing me option to show raw html. Perhaps I would need to check again and try to use it with other engines again. Will changing it to some other user agent reduce the chances of IP bans in Google or would it be same?

Yes exactly, as i just mentioned when you change the useragent the design of Google's search results will probably change. Google has a few designs optimized for different browsers and the underlying HTML of the results pages are different.

So if you change the useragent to one that has a different Google design you may need to also train ScrapeBox to find the links in the search results pages by changing the "Just before the URL" and "Right after the URL" fields also.
 
Well unless the proxies are just flat out dieing totally, then they are getting banned. If it requires captcha then its blocked, but there is no way to enter the catpcha, Captcha Breaker wouldn't solve them anyway and even if you could they become blocked extremely quickly again. This was tried in the past, but it just doesn't work out as well as it sounds like it might.

Are you uisng public proxies or private?



It uses MOZ, so you need a moz account and you can get upto 1000 backlinks. Actually V1 uses this as well as V2.

Private proxies mate 20 of them i have mountly i cant get more than that and when chapcha blocks to proxy i check from proxy manager after i pause to campain
, i see everything is fine, every proxies passes the google test.
 
Private proxies mate 20 of them
Private proxies are not good for scraping. Maybe with 100-200 private proxies you could try. For scraping you need public proxy service.
 
Since the last version upgrade, the RankTracker is not working for me. Giving blank results.

Can you please check it. Thanks
 
Since the last version upgrade, the RankTracker is not working for me. Giving blank results.

Can you please check it. Thanks

The recent update of ScrapeBox has nothing to do with the RankTracker, they are separate pieces of software. The last RankTracker update was 1st May and it's working, i have my own projects which run daily in it. So you may need to check your proxies and settings.
 
The recent update of ScrapeBox has nothing to do with the RankTracker, they are separate pieces of software. The last RankTracker update was 1st May and it's working, i have my own projects which run daily in it. So you may need to check your proxies and settings.

Thank you for the fast reply.
I checked the proxies and even replaced them. When I test the proxies within the proxy manager they are good. In the last two days, the rank tracker just gives me blank results.
My settings were not changed: 10 private proxies, use proxies - checked, 5 connections, Delay 60 seconds, stop when url is found - checked,
I have no idea what else could be.
 
When harvesting:
Russian slu,t intitle:"knee(s) hurt" things little
why google no result. Just pink with a zero. Blank result.
Can you please check it.
Have no idea what is happening.
Should I upload a image? Thanks.
 
When harvesting:
Russian slu,t intitle:"knee(s) hurt" things little
why google no result. Just pink with a zero. Blank result.
Can you please check it.
Have no idea what is happening.
Should I upload a image? Thanks.
No you needn't.
You query string is too long.:)
Shorten it a little bit and i got 7 results.
Also pay attention your syntax.
Not blank anymore.
 
Private proxies mate 20 of them i have mountly i cant get more than that and when chapcha blocks to proxy i check from proxy manager after i pause to campain
, i see everything is fine, every proxies passes the google test.
There are different kinds of ip bans. This video will help
https://www.youtube.com/watch?v=P9CbGhfc1aY

But basically depending on what you are doing you can literally have proxies pass the google test, which checks against a basic keyword, but actually be blocked for your query. Thats 1 potential option, which is why the proxies pass the test when you stop. The better question is when you stop and test them and start again can you then do another 100K results or does it fall off very quickly then and/or immediately?

The other potential option some 3rd party software that is monitoring scrapebox and when you hit X number of requsts its blocking all the requests. This seems less likely but is possible.


Thank you for the fast reply.
I checked the proxies and even replaced them. When I test the proxies within the proxy manager they are good. In the last two days, the rank tracker just gives me blank results.
My settings were not changed: 10 private proxies, use proxies - checked, 5 connections, Delay 60 seconds, stop when url is found - checked,
I have no idea what else could be.

As noted above proxies can pass a test but be blocked for the query you wind up using in rank tracker.

If you put it at 1 connection and 60 seconds and turn off proxies does it work? I mean its working fine for me on my machines, so its something local to your machine and that would lend it to being most likely proxy related somehow.

It also could be security software, such as anti-virus, malware checker, firewall etc... hijacking or blocking the requests so make sure the rank tracker is set as trusted/allowed in all security software.
 
Hi guys, is there a way to slow down the harvesting. Like some time delay for 1 threaded harvesting? :).
 
Hi guys, is there a way to slow down the harvesting. Like some time delay for 1 threaded harvesting? :).

Go to settings --> harvester engines configuration if you want to set a delay for any particular search engine. There is an option to add a delay. Set the delay to some seconds, 5, 10s or whatever and then update the engine. Then from the settings, you can change the required number of threads you want to use while harvesting. Different engines need different delay (some might not need any delay) and so, you can test and adjust it accordingly.
 
Thank you for the fast reply.
I checked the proxies and even replaced them. When I test the proxies within the proxy manager they are good. In the last two days, the rank tracker just gives me blank results.
My settings were not changed: 10 private proxies, use proxies - checked, 5 connections, Delay 60 seconds, stop when url is found - checked,
I have no idea what else could be.
Hi, not giving you blank now?
I don't use proxy.
 
Status
Not open for further replies.
Back
Top