I can't harvest with scrapebox

Drago05

Junior Member
Joined
Oct 31, 2010
Messages
151
Reaction score
11
If use operators like "inurl", "intitle" it doesn't scrape anything. I have filtered about 200 public proxies with 2000 time out but it doesn't work at all.

Does anyone also have the same problem?
 
Yes, i use costume footprints. The strange thing is i can scrape if i don't use advanced operators, but i can't when i use advanced operators. And i need to use advanced operators.
 
yeah advanced operators like " " are fucked. dunno why
 
Does anyone know another program that i can use to scrape url's which works with advanced operators?
 
I faced a little bit problem.I use costume footprints. The strange thing is i can scrape if i don't use advanced operators, but i can't when i use advanced operators. And i need to use advanced. Has any advance operator.
 
I think I see your problem. You are using "costume footprints" i think you should use custom footprints instead.
 
I see i am not the only one with this problem. I hope someone who knows how to work with this program to give some advice.
 
I think I see your problem. You are using "costume footprints" i think you should use custom footprints instead.

Dude, "custom footprint" is a radio button. It got nothing to do with the spell LOL
 
I read somwhere around here yesterday that inurl kills your proxies fast.
I can vouch for that because i used them on around 100k blogs and at one point they started scraping less and less blogs, then they just didn't work at all.
 
Drago, you're not alone. Special operators harvest slow as hell and burn out your proxies fast. G is far more restrictive on special operators than on regular queries. I was once doing keyword research and needed to enter around 100 allintitle: queries over the space of an hour. I thought it wouldn't be a problem. I was wrong - I was IP banned after around 50 queries. Even if you have hundreds or thousands of proxies, unfortunately they will burn out too quickly. At all costs, I avoid scraping using special operators. With any substantial lists, it's practically impossible to do. I harvest tens or even hundreds of millions of URLs at a time - this would be impossible with special operators. A few thousand URLs would be alright, but once you're pushing past a million assuming you get there as it's so damn slow, you are screwed.

My 'solution' to this is just a simple workaround. Instead of using the inurl: operator, use the string in quotes. For example, rather than inurl:index.php? you will be doing "index.php?" . You will often find that you get just as targeted results as with the inurl operator, i.e. accurate / the URLs you are looking for. But just to check, go into your browser and google the string in quotes and take a look at the URLs that come up. Click on a few and browse them. Are they still what you are looking for? If so, then great.

Often, if the inurl: string is quite unique, then simply putting it in quotation marks suffices, because G searches not only the page content, but also the url content, in fact, url content is one of the main things determining search results, ask any owner of an exact match domain. Putting "sadasdfasnh43???php_content=op&At3232" instead of inurl:sadasdfasnh43???php_content=op&At3232 is hardly going to lead to much difference, because the string is so unique G will have to find it in the URL, not the page text itself. Try it out.
 
Thanks for the info. I will try with quotes instead with advanced operators. But you say that, even if you use advanced operators you can scrape though they are slow and dying fast, but i can't scrape nothing if i use advanced operators - zero, nada. I don't know if i use privet proxies will make a difference.
 
I doubt it. The search engines are not perfect. It's likely the special operator is not being read well by the search engine, have you tried writing that footprint in a regular browser search and seeing what comes up?

My advice is you accept that special operators are often not viable and find ways to work around it. I have not been hindered by quotes.
 
Last edited:
These are some tests i made with advanced operators for scraping. I used different type of proxies: elite, anonymous, transparent, socks proxies - no success. But if i uncheck "use multi-threaded harvester" i can see every proxie that scrapebox use at the moment for scraping and for the most of them i see "error (302) IP blocked". So, i assume google blocked them. But the strange thing is that if i don't use advanced operators and scrape with the SAME proxies google don't block them and scrapebox can harvest url's.

I don't understand why when i use advanced operators the proxies are instantly blocked by google. I thought it will take some time before they are blocked, but if they are blocked instantly maybe i do something wrong. I don't know.
 
That's simple, Drago.

Special operators take very, very small amounts of searches for the IP to get blocked. That's why you can never harvest properly with them. Regular search, on the other hand, takes considerably larger amounts of usage to get IP blocked. Doesn't matter what type of proxy or IP - more than a few special operators => IP ban. Like I said you've got to work around it.
 
I've noticed the last couple weeks I've been getting 5k harvested links after removing dupes opposed to over 100k. I'm using the standard wordpress option and keywords I've used in the past. I'm not a complete noob either. But I am human, and could be making a stupid mistake. More curious if anyone else is having an issue, specifically in the last 2-3 weeks.
 
Back
Top