Future SEO - Knowledge Graph, Satori, Pregel, and ReduceMap

What I want to focus on, is reverse engineering (or attempting to) helping define those website signals that denotes a site as a top tier family/entity member. Sure, we can all say "Just build an authority site!" but what exactly does that mean now? I'd love to see some feedback from people who are building AND ranking new sites right now. What is working for you and how are your sites laid out?
 
I'm surprised this thread isn't getting as much discussion..did I post this in the wrong area, or are peeps just interested in warez'ing submission tools around here?
 
Excellent read, and i have actually read the linked article. This is where the SEO is heading, and the ones who can play into the format will see thier sites propel .

Just my .02 cents.
 
Dude, don't be surprised the tumbleweeds are blowing through this thread - the sad fact is, not to be disparaging of anyone, but 95% of people round here are lazy and just want the latest link scheme that "works", they don't want to trouble their minds with the future of where SEO is going.

This stuff is really interesting, I agree we should do some more brainstorming about this, and what it might mean for SEO in the future. I have a few questions, for example:

- you talk about an "established ontology" - is that how it works? The algorithm (it's all done algorithmically, right, because Google won't get into anything that's manual or not scalable) has DETERMINED the existence of certain relationships between entities, and those are basically fixed? How rigid are they? Can they be bent? Do they then inform the algorithm further and help it determine what does and doesn't fit into that family of entities? Would future gaming of the SERPS involve manipulating the algorithm to create self-serving relationships between entities, especially in "niches" where these relationships are not so firmly established (relatively lax network of "nodes")?

- does all this ultimately boil down to hyperlinking again? Are entity relationships essentially created by hyperlinking, even if that is not the way the data is stored anymore? Presumably it's not JUST about linking, but also about on-page proximity of entities to other entities, assessed through natural language processing of some sort. I.e. if an authority site is talking about "dog training" and frequently mentions "dog leashes" then we know that dog leashes belong in that network somewhere. Let's face it, Google is already doing this, so it's no big surprise. As far as hyperlinking goes, are the implications of this basically to do with niche relationships, and is this why non-niche contextual backlinking has been hit so badly by Penguin? Does this hugely increase the emphasis on the need for high-quality backlinking from authority sites in the niche (to look at it simplistically?)

- isn't "being at the top of the ontology relationship tree" the same as what we have been doing all along, trying to get to no. 1 by fair means or foul? Or does it have greater weight now? I have noticed that sites of mine in that position have actually taken a HIT to their inner pages, you know, the ones with the content related to this main niche that I am supposed to be a super-authority on now? How do we explain that kind of thing - just tweaks?

- are we saying that this is really what Penguin is all about? The Cuttster said in one interview that the name Penguin was meant as obfuscation ("For Penguin, we thought the codename might give away too much about how it works"). I don't doubt he is telling the truth, that would leave the way open for Penguin to be about quite a tangible algorithmic feature. Could many of the changes we are seeing with Penguin be explained by this notion of graphical relationships? Could an algorithm like this EASILY see fake patterns of linking, content etc? I wouldn't be surprised - all this talk of unnatural linking, non-contextual links, excessive anchor text etc. would be on the right track, but it wouldn't be seeing the whole picture, and the Google search people would be quite happy for it to stay that way, since as long as people continue to see the solution to Penguin as simply finding another backlinking scheme that works, they will not make significant headway in defeating it.
 
What I want to focus on, is reverse engineering (or attempting to) helping define those website signals that denotes a site as a top tier family/entity member. Sure, we can all say "Just build an authority site!" but what exactly does that mean now? I'd love to see some feedback from people who are building AND ranking new sites right now. What is working for you and how are your sites laid out?


P.S. I have been brainstorming this for a few months - recent Penguins put a few holes in some of my ideas, but I will have a think about this and see if I can share any definite thoughts, in line with successes I have been having. I think one key notion is "comprehensiveness" - how comprehensively a particular piece of content or a site covers a topic area. That is not to say your niche site has to cover everything under the sun regarding its topic (well, actually, maybe it does), but more about the "richness" of the content in terms of referencing key entities and perhaps introducing new ones in a way that these new algorithms perceive as "natural" and "quality". Very hard to define, I know... But that's why content LENGTH does have SOME bearing, though not just length for length's sake.
 
I have an eCommerce site that was kicked badly by the penguin.

Kinda related to this; a few days back I was looking for some hard drive recovery programs. Google gave me the run around, giving me loads of info about the nature of the problems and the programs. It also gave me a few developer sites for these types of programs.

I then went to Bing. Here I got instant help, the third result was a list of 5 free programs.

There's certainly a difference there. Google is providing information/knowledge, Bing is providing solutions/actions.

I'm thinking of setting up a new authority (information/knowledge) site with all the media content you mention cody, and then selling my eCommerce products from there.

I'm starting work on the new authority site today. There will probably be links to the ecommerce site so that I can send some traffic that way. But also in order to set up a relationship between the too sites.

I'm surprised this thread isn't getting as much discussion..did I post this in the wrong area, or are peeps just interested in warez'ing submission tools around here?

I think a lot of people just don't know what the hell is going on with Google anymore. And reading what information there is to find (like the article you posted), it seems the game has changed, and maybe gotten a whole lot more complicated.

Link building is easy. Understanding complex relationships between online structures is not.

Then again...how many times have we seen Exec VIPs and other knowledgeable people post about the recent Google changes? They are all very quiet on the matter. Either they have no clue what is going on - or they know exactly what is going on and are taking advantage of it...
 
Then again...how many times have we seen Exec VIPs and other knowledgeable people post about the recent Google changes? They are all very quiet on the matter. Either they have no clue what is going on - or they know exactly what is going on and are taking advantage of it...

I am certain almost NO-ONE really knows what is going on - how can they when there have only been a couple of Penguin refreshes, and that is not enough time to do any significant testing. Of course there is historical data and comparisons you can do, but seems to me there is a lot of people just bluffing out there. All I personally have to go on are theories of mine that I had put into practice, that were meant to anticipate things like Penguin - some of them turned out to be on the money, some not, but I am a long way from any definitive answers.
 
- you talk about an "established ontology" - is that how it works? The algorithm (it's all done algorithmically, right, because Google won't get into anything that's manual or not scalable) has DETERMINED the existence of certain relationships between entities, and those are basically fixed? How rigid are they? Can they be bent? Do they then inform the algorithm further and help it determine what does and doesn't fit into that family of entities? Would future gaming of the SERPS involve manipulating the algorithm to create self-serving relationships between entities, especially in "niches" where these relationships are not so firmly established (relatively lax network of "nodes")?

I'm beginning to think that's how it works. The freebase purchase by Google kind of hints towards that. Everything is still done via algorithmic means. I don't think it's fixed by any means. I think it needs to be flexible in order to provide the most up to date relevant information. Case in point: new movies coming out for actors. Google George Clooney. His entity is already on the right hand side. His general info is pulled from wikipedia, and then the other related searches are also linked to the entity (his girlfriends, other movies, etc). I think future gaming of the serps will be two fold: how to become a known family tree member within that entity so that you could leverage the click through. Once you click on George Clooney, go ahead and click on Stacey Keebler (who wouldn't, right? lol). Now the next family tree within the Stacey Keibler entity is populated within the SERP. Her top 5 listings?

en.wikipedia.org/wiki/Stacy_Keibler

www.imdb.com/name/nm0445001/

twitter.com/StacyKeibler

www.facebook.com/stacykeibler

www.justjared.com/tags/stacy-keibler/

So all in all, the top 4 will be hard to beat, since they play directly to the composition of what her entity should be defined by, you know? However number 5 appears to be a celebrity gossip site. Not terribly hard to beat in my opinion. But the fact remains, there are different ways to get into the family tree of relationships now. Just need to feed your website the right amount of info apparently (in my slow to form opinion so far..)


- does all this ultimately boil down to hyperlinking again? Are entity relationships essentially created by hyperlinking, even if that is not the way the data is stored anymore? Presumably it's not JUST about linking, but also about on-page proximity of entities to other entities, assessed through natural language processing of some sort. I.e. if an authority site is talking about "dog training" and frequently mentions "dog leashes" then we know that dog leashes belong in that network somewhere. Let's face it, Google is already doing this, so it's no big surprise. As far as hyperlinking goes, are the implications of this basically to do with niche relationships, and is this why non-niche contextual backlinking has been hit so badly by Penguin? Does this hugely increase the emphasis on the need for high-quality backlinking from authority sites in the niche (to look at it simplistically?)

I don't think hyperlinking is the subject matter anymore. The link is simply a link, however, where the link is coming from (general or specific authority site), will be the defining factor of how much relevance Google is now going to give the site. Goes back to the relationship tree. I'm going to press on with doing high PR ranking from non-niche specific sites, but i'm already targeting niche specific sites as well. I just fear a penalty if to many links are coming from locations that are deemed outside the family universe. Though you gotta think Google already allows for this with news outlets, PR sites and so forth.

- isn't "being at the top of the ontology relationship tree" the same as what we have been doing all along, trying to get to no. 1 by fair means or foul? Or does it have greater weight now? I have noticed that sites of mine in that position have actually taken a HIT to their inner pages, you know, the ones with the content related to this main niche that I am supposed to be a super-authority on now? How do we explain that kind of thing - just tweaks?
[\quote]
Yes and no. Authority site creation now, in my opinion, will require greater family tree awareness when setting up the content. I think when Google crawls through a site, the first page they hit is the main one, and that thing better be on point with what the site is about, giving off enough signals that it's an authorty (backed by other attributes found off page as well). As for the hit you've taken, not sure. We need to figure out WHY the sites ranked above you are doing as well as they are doing. I've got a thought on that later...

- are we saying that this is really what Penguin is all about? The Cuttster said in one interview that the name Penguin was meant as obfuscation ("For Penguin, we thought the codename might give away too much about how it works"). I don't doubt he is telling the truth, that would leave the way open for Penguin to be about quite a tangible algorithmic feature. Could many of the changes we are seeing with Penguin be explained by this notion of graphical relationships? Could an algorithm like this EASILY see fake patterns of linking, content etc? I wouldn't be surprised - all this talk of unnatural linking, non-contextual links, excessive anchor text etc. would be on the right track, but it wouldn't be seeing the whole picture, and the Google search people would be quite happy for it to stay that way, since as long as people continue to see the solution to Penguin as simply finding another backlinking scheme that works, they will not make significant headway in defeating it.

I think if we had a way to see the graphical representation, a lot of this stuff would make more sense. Bad link neighborhoods, profile links, guestbook links..Google knows where these links are coming from based on the type of site or PAGE the link is coming from. Really, all one has to do is start reading through a lot of the Google patents to see various ways to attribute strengths of a site in web rankings. It's convoluted as hell if your not used to reading programming schema, but they sort of lay it out in the patents. At the end of the day, they look to see at WHO is trying to game the system and honestly its not hard to tell when given the overall composite picture of where links are coming from for any given website profile. So what do we have to do then to get ahead?

I think it first starts with understanding WHY the top 5 sites (forget the the top 10), are where they are in the rankings. Most keyword tools will NOT tell you that. They're all still stuck in the traditional model of keyword analysis...which some of that still plays in understanding stuff, just not as much anymore. So if I had a tool (and I might try to find a coder for this), the tool would sort of do this:

1) Enter in a keyword phrase
2) Show me graphically (sort of like that tool I linked above) what other relationship keyword terms are linked to the keyword I typed in. I'm fairly sure the tools are scraping from the related searches function of google. Now, once we have a tree of connected terms start being built out, per term we can drill into the top 5/top 10 sites per keyword term.
3) Now, give me a breakdown of Page Rank, Domain Age, Indexed Pages, etc for each site for that given keyword. Also, give me a backlink profile for each site as well. Where are these links coming from, is there any input for others to also backlink off that page
4) Maybe a function to scrape the page of the top 5/top 10 site to show H1 Tags, keyword density as well
5) Also a function to show social mentions, how many social mentions a week/month breakdown
6) Also drill down into a graphical representation of how interconnected that site is to other sites within a niche too

That's the sort of tool that I think would help with understanding the current state of rankings? What do you guys think, am I onto something, or does a tool exist like that already the way I described it?
I was really hoping this tool would have funded via kickstarter..that would have gone a long way in this effort to understand..

http://www.gooey-search.com/

Thoughts, feelings?
 
I am certain almost NO-ONE really knows what is going on - how can they when there have only been a couple of Penguin refreshes, and that is not enough time to do any significant testing. Of course there is historical data and comparisons you can do, but seems to me there is a lot of people just bluffing out there. All I personally have to go on are theories of mine that I had put into practice, that were meant to anticipate things like Penguin - some of them turned out to be on the money, some not, but I am a long way from any definitive answers.

Absolutely, we're on the same page. But here's the thing, there are bunch of smart people on this forum, apart from the tool kiddies that only come here to download shit. I don't have access to the vip sections of this site, so i can't say for sure if this sort of thing is being discussed, but I wish more people would take a moment and think this stuff through, at least put conjecture towards it. Hell, I'm thinking things through as I type stuff out :) I have no answers, but, I do have a perspective on this due to how long I've been in this game. There won't be definitive answers for awhile I think just due to the nature of how hardcore this new algo implementation is. All we can do is test for metrics and try to replicate. But first, if we can get some structured analysis into place, that might give us a leg up, hence the sort of tool I mentioned above.
 
Did some more digging, some really really interesting reading here on a Google Patent filed in 2005. Sort of gives some insight as to HOW google attributes a good score to a website based on the number of backlinks and the age of the backlinks pointing to a website (in the patent it describes WEBSITE as DOCUMENT..probably referring to the HTML as a document).

http://appft1.uspto.gov/netacgi/nph-Parser?Sect1=PTO2&Sect2=HITOFF&u=%2Fnetahtml%2FPTO%2Fsearch-adv.html&r=1&p=1&f=G&l=50&d=PG01&S1=20070094255.PGNR.&OS=dn/20070094255&RS=DN/20070094255

This stuff may be antiquated or actually may be changed somehow to match what the google algo is doing .Good reading nonetheless
 
I wonder if this thread would have got more action in the black hat area....
 
I wonder if this thread would have got more action in the black hat area....

Hmm, sorry, wish I had more time to devote to this, but I will try to chime in here whenever I can because I do think this is very important. But probably it's over the heads of MOST of the kiddies on this forum - again, not to be disparaging, it's just reality. Most people around here think "white-hat" SEO is AA comment spam and Web 2.0s, and I am not even kidding.

About that "link age" patent - interesting, that is quite old, but I never came across it before. So link age and link velocity almost certainly ARE important factors (not sure why that has been so hotly debated then...), but I am pretty sure Google has moved on from there to analyse backlinking patterns for telltale signs of manipulation (you can bet that the "unnatural links" warning is mostly or entirely algorithmic). Bing have also claimed that for instance they can clearly distinguish between natural viral social sharing patterns and gamed ones.

But going back to our main topic, it seems to me that part of the game of garnering niche authority has a lot to do with your on-page/on-site treatment of a topic - I mentioned comprehensiveness, but I think it's more than that. Seems to me structures like "siloing" go a long way to creating tight little networks of related, interlinked entities which establish your site as an authority on the subject, which coupled with high authority links from other sites/pages that are considered authorities, makes for a winning formula, now and even more in the future (high-authority, contextual backlinks, who'da thunk it).

This is probably just a tiny part of the picture, but just compare this to "creating a page of keyword-stuffed content and hammering it with links" and I think we are already well ahead of most in this game...
 
Yeah, I'm not sure this would've caught much more attention elsewhere either.

I think it is time that our methods of SERP analysis did some major evolving - both to check competition as well as to determine what is really working across various niches. I'm fascinated by the idea that relationships between websites in some sort of a family tree hierarchy could be a major player in SEO either now or very soon. I have a feeling these relationships are already playing a bigger part than we realize in the search results, and it will probably grow soon. Which means two things - in niches that don't have a well established hierarchy, now is the time to step in perhaps by creating all or part of those relationships ourselves. And in niches that already have an established "family tree" we need to figure out the best way of grafting ourselves into that tree... I think that might be closely related with determining why/how Google decides which search pages get built out with extra info right on the SERPs while other pages don't (yet at least). It sounds like this is all determined algorithmically, so again, with some testing it seems like it should be able to be figured out and eventually gamed.

I also have a feeling the "related:" function of the Google search (i.e. do a search for related:Google.com) will also be useful for digging more into how Google views relationships between websites.
 
Maybe a hard-core SEO discussion forum...

Thanks for the "related:" operator tip - I'd clean forgotten about that. It would be a bit odd, wouldn't it, if that operator used a completely different algorithm to the one we are talking about! I am going to go and play with that on some of my sites and see what Google THINKS I am related to! (I fear I am going to get a nasty surprise :D)
 
I had forgotten about it too to be honest, but Cody41 linked to http://www.touchgraph.com/seo/ in one of his posts up above and that seems to be taking data from Google searches using the "related:" operator and then plugging it into a visual representation. It's an interesting tool to play with.

I've also noticed that some sites return a message that says:

Your search - related:randomdomain.com - did not match any documents.

One site in particular has plenty of backlinks, sometimes quite a few from the same site, so that in itself isn't enough to establish a "relationship". It might be worth digging around to find some documentation on this operator. Though it's been around for quite a while, so I'm sure its algorithm has changed over time like everything else...
 
It's hard to take the thread seriously when the title contains "ReduceMap" and most of what's being discussed is just retreads of ancient seo tropes, for example people have been suggesting linking to authority sites as long as I can recollect, except 6 years ago it was to support what people were calling "latent semantics." Armchair speculation is pretty pointless in any case, practical SEO is an ongoing experiment in which that which results in ranking and traffic is good, all else is bad, what's happening behind the curtain is irrelevant.

Act first, theorize later.
 
It's hard to take the thread seriously when the title contains "ReduceMap" and most of what's being discussed is just retreads of ancient seo tropes, for example people have been suggesting linking to authority sites as long as I can recollect, except 6 years ago it was to support what people were calling "latent semantics." Armchair speculation is pretty pointless in any case, practical SEO is an ongoing experiment in which that which results in ranking and traffic is good, all else is bad, what's happening behind the curtain is irrelevant.

Act first, theorize later.

Sorry for the open discourse, didn't mean to offend your sense of what is important and what is not. Love your contribution to the thread by the way. Clearly I goofed on the thread title, when I should have said Mapreduce, but you got the gist of it. Armchair speculation goes towards helping one to flesh out the thought process in order to do the required testing. Clearly some of us need that initial understanding of HOW to proceed before testing.

What's happening behind the curtain is definitely not irrelevant as if one can attempt to understand why certain occurrences are happening, then one can attempt to aim for the same effect and try to replicate. Based on this thread alone, i'm already fleshing out a requirements document for a new seo intelligence program which can go a long way towards providing a new type of seo understanding.
 
What are forums for if not "armchair speculation"? We might as well not talk about anything then.
 
It's hard to take the thread seriously when the title contains "ReduceMap" and most of what's being discussed is just retreads of ancient seo tropes, for example people have been suggesting linking to authority sites as long as I can recollect, except 6 years ago it was to support what people were calling "latent semantics." Armchair speculation is pretty pointless in any case, practical SEO is an ongoing experiment in which that which results in ranking and traffic is good, all else is bad, what's happening behind the curtain is irrelevant.

Act first, theorize later.

If you were taking a basic math test, which would serve you better:
1) To understand how multiplication and division work so you can come to logical conclusions
2) The ability to take the test over and over using trial and error to eventually figure out the right answers

I'm all for taking action and doing split tests to determine what works - but doing that in conjunction with trying to build an understanding of the "why" seems like the much more logical choice to me. I'd rather know why the answers are right over blindly guessing until I stumble into the right answer any day of the week. There is definitely a point when speculation needs to give way to testing and action, but to say that speculation is completely pointless is ridiculous. Surely you must do some sort of speculation to determine what is worth testing in your ongoing experiment?
 
There are definitely several factors that any internet marketer needs to start taking into consideration. However, as JR stated, we really can't read into it too much at this point. As Google's database continues to grow and the processing of information from Knowledge Graph becomes more intricate, everyone should work on creating a site that doesn't look too spammy, does contain GOOD content (GOOD meaning, information that can be readily used by the visitor within its relevant context, not just to generate sales), thus establishing it's own authority.

There would be several variables to take into consideration when deciding what link/contentbuilding strategy you need to take. In business terms, you just have to think of the top 5 (and at times top 10) as the "larger corporations". There is an almost impossible barrier of entry into those slots as you'll be competing with website that have well aged domains, tons of indexed pages/content, and thousands (if not millions) of links pointing to their site. Google is raising the bar, the "barrier of entry" into the top SERPs. Although there are several factors that the search engines will take into consideration in determining your SERPs, there are some that will be weighted more heavily. So for any "small business" hoping to break into their respective "industries" you will need to focus on those factors that are more heavily weighted, and disregard the rest, and what this whole forum is about.

We will continue to test what strategies work, which look more "natural" to the search engines, and see how we can get to the top. I doubt there will be any sort of epiphany causing the search engines to make a 180 decision in how their algorithms work, overnight. The changes will continue to be gradual, with several more Penguin n.0's to come. So as long as you aren't putting out spam, you shouldn't be too concerned about your SERPs being dramatically affected. If you're in an established database, know that it will take longer to get to the top, so take a look at whether or not there is any money to be made in that particular niche, if you will be able to get your ROI, and eventually make a profit. But unless you want to make this your main project and build a software that will help determine profitable keywords in "G's new landscape", then time may be better spent on continuing to find niches that are still easy to rank for, make some money, and continue to build "authority" sites on the side.
 
Back
Top