Then y r u here. Teasing others ?can't even imagine such situation when we talking about customers.
And this is why, I am learning from the most trusted source - from Google.
Yeah, your SEO experts are the most successful when it coming about link building:
Can you recommend more experts to follow ? Great entertainment.
Seems, you bet on lame horses
All of them build Expired Web 2.0s same as you, because simply you don't have $$ for solid PBNs.
Sad but true.
Also @FatBee promises in Dec 2015 that "i will invest in PBN soon as i earn some money" - seems, no luck until today.
Because he still using very low effective Web 2.0s.
Cheers, Greg.
@Stas Va
Good that you know the difference between historic and fresh index.
For any new site which is under 3 months in age - historic index is completely useless - I am agree here.
For any older site you don't see in fresh index backlinks which was discovered 4 months ago and not crawled again in last 3 months.
Such situations are the most common, as the result you not see significantly part of backlinks in fresh index.
Because Majestic have 50x larger pages index than Ahrefs, there is not possible to crawl all those pages in 3 months.
For this reason Ahrefs have more backlinks in fresh index, but much less backlinks when we will sum fresh and historic index.
And this is why professionals use historic index and choose Majestic.
Let's we talk now about crawling system.
By poor written scripts you aren't able to effective crawl the web.
I'll explain you how working my system, it will be easier to you understand the point.
So, before any single host/page are crawled - for all the hosts system will get the IP for each one.
As the next step base on IP/urls - queue being created in the way, that it not will overwhelm any single server/host.
In crawling process, information about single host speed / http code - modifying the queue.
For example, when the host replying slowly - the next request is made after longer time, same as doing it Google.
Also when some host trying to block my crawlers - some other algorithms are used (other server location, proxy,etc)
Mentioned recalcitrant host is crawled until successful crawl, then used successful method is stored for future use.
So, there no any chance to hide anything as some amateurs think - in reality they're not able to block nothing
In that simple way my system solving all crawling problems.
However, due to multi server application - it's not easy to create it.
Because it need to modify the queue "in the air" on multi servers in multi locations.
Cheers, Greg.
So that's already says that you can't crawl thousands of backlinks in minute without loosing in data. Even if you will create a smart queue of requests, and will manage to check links from the same domain in distance proportional to each other. That won't resolve your problem with the performance that you declare.@Stas Va
For example, when the host replying slowly - the next request is made after longer time, same as doing it Google.
@Stas Va
By poor written scripts you aren't able to effective crawl the web.
@Stas Va
2) If you know how to create and update in air the quene (what I've explained to you) nothing will be crushes.
Seems, you not understood what I wrote. Because I not wrote about your IP but crawled host IP.
Also scripting language is not the way to write such effective crawlers.
Can you explain what's info about backlink you will get crawling IP ? How your crawler will eve check if the link is really there or ahrefs attributes ?@Stas Va
3) The point is to crawl thousands IP's at the same time, not the thousands urls at the same host and by this way, noting not will be crashed.
I am write this from my expirience not any theories.
Are u sure now ? Because few posts earlier you told me ...@Stas Va
4) In partial true. You are not able to index thousands of urls from the single poor host in short time. (see point 2,3)
I am don't care about poor links from low TF domains, because they have tiny power.
Especially as you wrote, from the same domain.
Cheers, Greg.
You are wrong even double.
Not true, except if you using poor written scripts.
Cheers, Greg.
Hei @GregFromMoonsy , thanks! Tried Majestic (used ahrefs for a long time) today and found a lot of hidden backlinks from my competitors. It's amazing! Thanks again dude!
before any single host/page are crawled - for all the hosts system will get the IP for each one.
As the next step base on IP/urls - queue being created in the way, that it not will overwhelm any single server/host.
what's info about backlink you will get crawling IP ?
Garbage data ... ignoring data relevance...
Tried Majestic (used ahrefs for a long time) today and found a lot of hidden backlinks from my competitors. It's amazing! Thanks again dude!
Can someone explain why Ahrefs and Majestic Rds don't match? Always Ahrefs says like 500 RD but majestic says only 10-20 RDs. Anyone know?
Ahrefs is much faster and gives more professional overview.
At first - you should throw your team away, because they aren't able to use the brain at all.
As the next step, you must post the request in "Hire a Freelancer" section with the request to create mentioned solution.
I am pretty sure that each clever programmer not will be have any hassle with understand what I have on mind with mentioned sentences.
Especially if he have the real experience in web crawling.
Let's we talk about the next thing related to Majestic and Ahrefs index:
Tell me the domain name which have as you said the "garbage data" in Majestic History Index.
I will destroy your fake story in few minutes, because I am pretty sure that you created your account yesturday only for trolling reason with Ahrefs.
Seems @AngelSeo also only trolling in this thread. Maybe your nick it's just his second account on BHW.
He know everything but not understand absolutely nothing same as you.
As I previously said, Majestic crashing Ahrefs many times.
Typical example with the fake theory about bigger Ahrefs fresh index:
I told you, if they (your team) aren't able to understand for which reason they need to check IP when we talking about creating the quene, they don't use the brain at all, because here I haven't on mind "crawling the IP". For this reason - my suggestion to you was: hire the clever programmer who don't will be have any hassle with understand this part.You didn't answer... crawling system get by Host IP
Only what is useless here - it's your (and your "team") very low knowledge about crawling the web.Historical DATA ... 318443 - does not exist ...only people who didn't work enough ...historical index - is useless
Not true, except if you using poor written scripts.
Cheers, Greg.
The point is to crawl thousands IP's at the same time, not the thousands urls at the same host and by this way, noting not will be crashed.
I am write this from my experience not any theories.
I told you, if they (your team) aren't able to understand for which reason they need to check IP when we talking about creating the quene, they don't use the brain at all, because here I haven't on mind "crawling IP". For this reason - my suggestion to you was hire the
Only what is useless here - it's your (and your "team") very low knowledge about crawling the web.
And bellow you'll find why.
Cheers, Greg.
I told you, if they (your team) aren't able to understand for which reason they need to check IP when we talking about creating the quene, they don't use the brain at all, because here I haven't on mind "crawling IP". For this reason - my suggestion to you was hire the
Only what is useless here - it's your (and your "team") very low knowledge about crawling the web.
And bellow you'll find why.
Cheers, Greg.
4) In partial true. You are not able to index thousands of urls from the single poor host in short time. (see point 2,3)
I am don't care about poor links from low TF domains, because they have tiny power.
Especially as you wrote, from the same domain.
Cheers, Greg.
You didn't answer
3 from 4 of your points are not relevant to our discussion. Our disput is related only with the first point. I already showed you that most of backlinks for almost every average site, are from 10 % of referrals domains. You already acknowledged that queue will not resolve the original problem .I told you (but you still can't understand), first you need the IP address for each host to create mentioned queue.
Without the IP address you aren't able to create the effective queue because:
As I told you too - quene is modificated in the air - basing on crawling stats, such http code / page speed.
- You don't know which domains are on the same IP or /29 range - in result you crashing single server or you receiving 403 50x http code.
- You don't know in which country server is located, in result crawling is slow or even impossible (such mentioned China)
- You trying to crawl the domains which even don't have the DNS A,AAAA,CNAME records or domain was deleted, in result crawling have very low effectivity.
- You must have daily updated domains database with current domain status (such active,expired,deleted,redemption etc), because in other case - you crawling suspended pages etc (where in result you're not able to find the link)
If you receiving 403 50x http code - you should run additional algorithms to detect if IP was banned or just server have some issues.
In partial true. You are not able to index thousands of urls from the single poor host in short time. (see point 2,3)
Especially as you wrote, from the same domain.
However, my crawling system is able to check fat thousands urls in few minutes and check each link.
So, for professionals it's not the problem at all.
So, you already learned from me a lot of things without even simple "thanks", which cost you nothing.
Instead of this, you trolling with quoting my sentences, which in reality you aren't able to understand technically.
Ask Ahrefs how many servers they have in China with their "biggest" indexGreat that our other dispute about the senselessness of the statistics of crawled pages, which "majestic" declare and you have so actively used in this thread was closed.
Please show me the screenshot of Ahrefs Historical Index + Google sheet with dead BL for this domain, I guess I'll be laugh strong.You can destroy me any time...
Ask Ahrefs how many servers they have in China with their "biggest" index
I'll tell you - they using only OVH, AFAIK OVH don't have any single server in China and each amateur is able to block Ahrefs crawlers at iptables level in invisible way.
Please show me the screenshot of Ahrefs Historical Index + Google sheet with dead BL for this domain, I guess I'll be laugh strong.
That will be enough to show reality.
Hope that your amazing crawling script not need the whole week for crawl those few urls from Ahrefs.
As it was designed by team of first class engineers who asking on forum about help
Cheers, Greg.
@FatBee
Maybe you not finished any school and for this reason it's very hard to you compare really big numbers such:
8 485 618 151 026 Majestic pages in index
149 000 000 000 Ahrefs pages in index
Cheers, Greg.