Joseph Lich
BANNED
- Nov 25, 2015
- 401
- 79
Yeah you are right and sagacious after tested so many suckers software tools online only one alive.
Cheers!
Cheers!
What do you get in a browser at google.co.uk for the same terms?
Different results.
So when I query a keyword, I get the exact same results as a browser, but when I run "business name" -site:bizdomain.com, I get a range of different citations in the browser, but when I use SB I get 9 yell.com results.
I'm getting errors after a few keywords scraping with the article scraper paid plugin. I'm using now 16 proxies with 1 thread only and I still get errors, and the scraper stops then most of the time.
I admit some of my proxies are in asia, but these are private proxies only.
If one column in the plugin throws out "ERROR", it will no longer scrape from that site, despite the other 3 article directories still being scraped.
Can this behavior be altered somehow please?
I know that once it reaches error state that it won't scrape any longer for that source. I believe that it retires until the ips are all blocked. So at that point you would need new ips.
Question, when it does this, if you restart it back up right away, does it scrape a bunch of more articles before going to error or only a few or none?

Coule automator perform this chain works?
First find some seed list, for example a wikipedia page, or a page from a famous directory website; get some links by "grab links by crawling site"; a 3 depth crawl can easily get half million internal urls; get some of them and extract external links.
then "backlink checker" these urls;
link extract all the urls
[ comment grab 30 connection
sitemap 200 connection
phone number 200 connection
custom grab 500 connection
link extractor 1000 connection ]
alive check, custom grab etc -- a rapid filter;
post/publish links to a website, batch post to wordpress for example;
the Finder --- get all the metrics.
Plus the proxy manage and Custom Harvester, it must be a distinguished gathering;
+++++++++++++++++++++++++++++++
I tried automator havest | test | save; but the saved proxies
are
harvested proxies; instead of
tested proxies.
Any idea?
View attachment 73711
connections limit for various functions are set to 3500.
various ---- Proxy Manager + Havester?
Could the automator be designed like this: items can be drag drop to the field and items can be connected with each other so the parameters values can be piped/transferred. And is it possible to introduce a special de-dup especially prepared for the the URL frontier ? Although there is a Text File Tool (dups, sort, split etc) which is very good.
The IP-s aren't getting blocked I think, because when I retry, they work again, it's more likely a timeout issue with the proxies.
Out of 16 proxies:
2 are from Singapore, 1 from Japan. The rest are from EU/USA NYC, Miami and Dallas.
Scrapebox's on a dedi in Paris, France.
These are exclusive for scrapebox and social media, no posting on these.
I think it's likely that the 2 singapore proxies are shit. I know for a fact they are on bad nodes.
These are all on cheap servers where I configured squid3 on.
http://postimg.org/image/3kl2ei7r9/
View attachment 73725
So many days after the new malware and phishing filter addon been updated and published. But no one had jumped out and explain it's new function parameter meaning etc. Isn't it very weired? I googled "AS host" and got no clue. Also some other new words: software list | Social list | Times list
|Redirect Sites |Traffic Source ------ What are these?
In a very recent updated version, now can post to some new platforms.
--- But I don't know what are these platforms, Where are they?
Post domains to a wordpress blog then crawl it, which is not necessary under certain circumstances.I think it could do everything except batch post them to wordpress, which I don't know why you would want to do anyway.
How do you know this?Software list, malware list and social list are different lists they can be listed on as being bad. Times listed is the number of times they have been listed on the whole. Redirect is having to do with if its redirected to another site I think in an attempt to steal info.
First half is give more control of the software to a new level; Sometimes I think it's not absolutely needed.Can you explain this one more?
Guestbook:The list of platforms that can be posted to are there.
View attachment 73748
Custom Harvester: Using the token {keyword} replace the root domain, (google here) is this possible?
3500 potential connections so worth it.