Best Tool For Scraping links from Google. Hrefer? Scrapebox? Or something else?

peterhofmann

Newbie
Joined
Mar 2, 2013
Messages
22
Reaction score
0
Hello,

I have tried scrapebox and hrefer to scrape URLs from Google.

For scrapebox, it is easy to use but can't get enough links for me.

For Hrefer, I think there must be some problems with the engines.ini, I saw the query code is trying to get 100 results each page, but it only got 10 results each page, also it can't scrape enough URLs for me.

Do you know how to fix the Engines.ini problems of Hrefer?

Or do you know which software can do better job than them?

Hope that someone can help me, thanks.
 
Both will work.

If SB isn't getting enough links you need to take a look at how many keywords/footprints you are using.

It's not going to have any limitations as far as amount of links you can scrape.
 
It's usually not what tool you use but how you use it. Scrapebox is enough for 99% of your scraping needs and Hrefer is said to be the best scraper you can find. I seriously doubt you will find something that works for you if these 2 don't.

Did you use enough footprints / keywords / proxies?
"No"? - Use more!
"Yes"? - Use more anyway!
 
With ScrapeBox i scraped about 276,000 URLs in about 3 hours. I pressed the abort button because it was enough for me and ScrapeBox crashed with an error saying something about Not Having Enough Memory

Also, i got the same error when I was merging my footprints with my common words list. If I reduce the keywords and abort scraping before 100k, then it works just fine.
 
Last edited:
With ScrapeBox i scraped about 276,000 URLs in about 3 hours. I pressed the abort button because it was enough for me and ScrapeBox crashed with an error saying something about Not Having Enough Memory

Also, i got the same error when I was merging my footprints with my common words list. If I reduce the keywords and abort scraping before 100k, then it works just fine.

This type of stuff is going to be server dependent.
 
definetely hrefer would be much better since there is also a sieve filter in place. if your engine.ini does not work, i suggest you to learn how to modify it. it's pretty simple
 
Why not to try GScraper? 276,000 URLs in about 3 hours is just a piece of cake.
 
Why not to try GScraper? 276,000 URLs in about 3 hours is just a piece of cake.

Hawke, I have Gscraper. I ended up purchasing Scrapebox. GScraper is not a bad product. But it now has many bugs in it. The latest version following 1.2.1.1) inconsistantly scrapes when using a footprint such as site:.com %KW% and site:%KW%.com. By inconsistent I mean some times it will scrape and create a report and a file and sometimes it wont. I have attempted many different footprints that I use, and this problem is consistent throughout the footprints, including the one that ships with the program.

Then as you have tried to improve the program, the memory footprint has grown from 239k of ram at boot to over 500k at boot.

What Gscraper does better than Scrapebox is skip dead proxies where as a dead proxy will just throw an error in Scrapebox and the search is skipped.

Being honest, I feel at this time that I wasted my money on GScraper. I know you are trying hard, but....
 
I own ScrapeBox since the version 0.1 and I have never tried another tool. Anyway I think ALL depends on proxies (and quality of those proxies too). I have tried several proxy providers and now I have an established subscription with three vendors, giving me the chance to have ~50-60 proxy (fresh every month).

I think you should look here at BHW for proxy vendors, there are good offers always active and this should give you the chance to test and test and test. Also, add tons of keywords during scrape (at least 50 keywords).

Hope this helps :)
 
Hinkys and donaldbeck tell the truth out,the more you need the more you pay;)
If you have no idear,Just try xboter Scrape Sonic,maybe it able to meet your needs,you can get proxy from their server and there are many footprint building in software.The key is that it is free now.

Good luck!
 
where as a dead proxy will just throw an error in Scrapebox and the search is skipped.

No in ScrapeBox when a query fails due to a bad proxy, it will retry the query again with a different proxy 3 times by default. You can raise it up to 20 retries under Settings > Adjust multi-threaded harvester proxy retries.

Also the Custom Harvester will go up to 99 retries for every query, see the "Proxy Retries" setting up the top right.

serp-scraper-settings-750x648.png


With ScrapeBox i scraped about 276,000 URLs in about 3 hours. I pressed the abort button because it was enough for me and ScrapeBox crashed with an error saying something about Not Having Enough Memory.

If that ever happens, all your urls are saved in the /Harvester_Sessions/ folder so nothing is ever lost.
 
Last edited:
No in ScrapeBox when a query fails due to a bad proxy, it will retry the query again with a different proxy 3 times by default. You can raise it up to 20 retries under Settings > Adjust multi-threaded harvester proxy retries.

Also the Custom Harvester will go up to 99 retries for every query, see the "Proxy Retries" setting up the top right.

Allow me to clarify.

GScraper places the queries into a circular que. Using the default setting for Scrapebox; if P1 is dead, and P2 is dead, and P3 is dead, then an error is thrown and query is skipped. In Gscraper, if P1...P3 is dead, then Q is moved to P4...Pn. The only time that a query will be skipped is if all the proxies are dead.

Each method has its pro's and con's. Because of the transitory nature of proxies, I believe that GScrapers method is much better than the stack method that Scrapebox appears to use, especially for long scrapes that may be run over several days.

Both SB and GS suffer from the fault of being unable to pause the program to insert new proxies during the scrape. Scrapebox attempts to resolve this with the custom harvester that scrapes for new proxies every X minutes, where as Hawkes suggested work around is to set the flag to delete the used keywords when the scrapping is stopped and then insert the new proxies. Both methods are cumbersome. With SB this is because the proxy scrape and test is based off the settings and may pull >34K proxies and then test them prior to continuing. With GS this is cumbersome because you may lose the majority of a search for several keywords.

When it comes to posting, at one time GS dusted SB at ~70% to ~40% success. Now due to bugs in GS, the posting success is less than SC.

GS has no learning mode for new platforms, whereas SB does. SB wins here.

GS only searches Google, SB searches many different engines and can be taught new engines (if you have the proper settings for Baidu and Yandex, can you provide them as I can't seem to get them right). Scrapebox wins hands down.

Gscraper allows utf8 encoding so that if I want to search for botnet in Russian and Chinese I can enter ботнет and 僵尸网络 directly into the keywords. Scrapebox has no facility that allows me to do this. This would be especially useful when searching foreign (non US) countries search engines using the native language.

One of my peeves with GS is the use of the registry for the database along with attempting to disable registry tracing. Scrapebox does not do this. However, even though frowned on by MS because of the high probability of corrupting the registry, many commercial programs do this as well. I get around this problem by running Gscrapper in user space rather than administrative space. While I do not run SB in administrative space, I would feel comfortable doing so. I would not run GScraper in administrative space.

Then my biggest peeve with .net programmers (GScraper) is that because Visual Basic (.net) has garbage collection and rudimentary memory management, they think that they do not have to manage the memory. Gee, I don't have to delete an unused object and recover the space because the garbage collector will do it. This leads to excessive memory consumption and disk thrashing. I have had to force shutdown of GScraper more than once behind this. However, GScraper is not as offensive as Proxy Goblin in this regard.

Then there is another issue with GScraper, every search is sent to China. This appears to be validation of the programs authority to run, but I do not really know. This excessively consumes outgoing bandwidth and leads to the familiar no authority for the operation message. While I have not had SB long enough to see if this is the case with SB, on the surface SB appears only to validate at program start up.

Not every program is going to have everything that a person might want in it, for the purpose, at this stage of development, and within the tests I have done and the observations I have made, Scrapebox wins hands down. GScraper has gone from good to bad, and is now working on terrible.

The reason I am bringing this out is that maybe it will motivate Hawke into addressing the problems that GScraper has.
 
Then my biggest peeve with .net programmers (GScraper) is that because Visual Basic (.net) has garbage collection and rudimentary memory management, they think that they do not have to manage the memory. Gee, I don't have to delete an unused object and recover the space because the garbage collector will do it. This leads to excessive memory consumption and disk thrashing. I have had to force shutdown of GScraper more than once behind this. However, GScraper is not as offensive as Proxy Goblin in this regard.

Will have to disagree with you there, .Net doesn't have a "rudimentary memory management" system. The main problem is objects not being disposed once being used (such as a connection), which will lead to the garbage collector seeing that the object is still in use. It's a shitty programming, nothing to do with .Net framework.
 
I'll fully agree with you here, but I'd throw in there is plenty of shitty programming within the .net framework itself, and the nature of the .net framework itself just encourages shitty programming. It's a shame because C# is actually a nice language and they actually made a very nice IDE, but it runs on top of layers of shit.
 
Will have to disagree with you there, .Net doesn't have a "rudimentary memory management" system. The main problem is objects not being disposed once being used (such as a connection), which will lead to the garbage collector seeing that the object is still in use. It's a shitty programming, nothing to do with .Net framework.

Microsoft disagrees with you:
SUMMARY Garbage collection in the Microsoft .NET common language runtime environment completely absolves the developer from tracking memory usage and knowing when to free memory. However, you'll want to understand how it works. Part 1 of this two-part article on .NET garbage collection explains how resources are allocated and managed, then gives a detailed step-by-step description of how the garbage collection algorithm works. Also discussed are the way resources can clean up properly when the garbage collector decides to free a resource's memory and how to force an object to clean up when it is freed.
Source: http://msdn.microsoft.com/en-us/magazine/bb985010.aspx
 
Microsoft disagrees with you:

Source: http://msdn.microsoft.com/en-us/magazine/bb985010.aspx

If you read the whole article you would have realized that in fact, Microsoft agrees with me...

To get a resource to clean up properly, the developer must write code that knows how to properly clean up a resource. In the .NET Framework, the developer writes this code in a Close, Dispose, or Finalize method, which I'll describe later.

I do this for a living mate, and not only that, but the article is dated from the year 2000, more than a decade has passed
 
Last edited:
And those links prove what?

When the garbage collector performs a collection, it checks for objects in the managed heap that are no longer being used by the application and performs the necessary operations to reclaim their memory.

This goes right back to me explaining to you that resources must be disposed of to allow the GC to do its job, once you dispose a resource it doesn't sit around there forever waiting for the GC to collect it.

I am unimpressed by your claims of being a programmer.
lol, my work speaks for its self ;)

TwitterMarketing.jpg
 
And those links prove what?



This goes right back to me explaining to you that resources must be disposed of to allow the GC to do its job, once you dispose a resource it doesn't sit around there forever waiting for the GC to collect it.


lol, my work speaks for its self ;)

First paragraph. first link
The .NET Framework's garbage collector manages the allocation and release of memory for your application. Each time you create a new object, the common language runtime allocates memory for the object from the managed heap. As long as address space is available in the managed heap, the runtime continues to allocate space for new objects. However, memory is not infinite. Eventually the garbage collector must perform a collection in order to free some memory. The garbage collector's optimizing engine determines the best time to perform a collection, based upon the allocations being made. When the garbage collector performs a collection, it checks for objects in the managed heap that are no longer being used by the application and performs the necessary operations to reclaim their memory.

So at the minimum it shows that .net has "Rudimentary memory management," as I stated. You complained about an older link, so I gave you the most current links directly from the Microsoft Developers Network.
 
Back
Top