Build Great Backlinks has posted a new item, 'Say Hello to Fresh Alerts: New
Mentions and Link Notifications in Your Inbox'
Posted by Cyrus-Shepard
Imagine a product similar to Google Alerts, only much better. It's built
specifically for marketers and SEOs. This product not only finds mentions of
your keywords and brand, but also reports new links to any website or URL you
choose. It comes equipped with advanced search operators to discover new
opportunities, and its exportable metrics are sortable by both date and Feed
Authority.
To top it all off, it now alerts you via email whenever it finds something
new.
Announcing Fresh Alerts from Fresh Web Explorer
For the past few months we've enjoyed using Fresh Web Explorer, which has
quickly become my favorite new marketing tool. Since then, our engineers and
developers have been working to add email alerts to the mix to vastly improve
its value.
Starting today, when you use Fresh Web Explorer, you can now set up alerts for
up to 10 queries of your choice. The emails are sent daily whenever anything new
is discovered. Because Fresh Web Explorer refreshes its index every 8 hours,
this means you can be notified of new links and mentions literally within hours
after they appear on the web.
When you run a query in Fresh Web Explorer, you have the opportunity to create
an alert based on that search.
One key feature is the ability to set your timezone. This helps tailor the
reporting specific to your area of the world, so the alerts are more relevant to
you.
Fresh Alerts for SEO and inbound marketing
I've had the opportunity to beta-test Fresh Alerts for two months, and I can
say without hesitation that it's literally changed the way I do SEO and inbound
marketing. We also tested the product with 1,000 Moz beta users, and the
feedback has showcased the variety of ways folks are using Fresh Alerts.
1. Link building
While we built Fresh Alerts as a mentions tool, it does a remarkably good job
at helping to build links through the process of link reclamation. By using the
built-in search operators, you can set your alerts to find non-linking mentions
of your brand or keywords on the web.
For example, if I want to search for folks who mention MozRank (a Moz branded
term) but don't include a link to Moz, I'd set up my Fresh Alert like this:
mozrank ârd:moz.com (mentions of Mozrank that don't link to the root
domain moz.com)
With this alert set, every day I would get a new Fresh Alert in my inbox with
a list of mentions. Looking at the number of non-linking mentions above, I'd
better get link building!
2. Reputation management
Using Fresh Alerts, you can easily be notified whenever anyone mentions you
name or brand on the web. Hopefully the information is positive which gives you
the opportunity to open a relationship or simply stay on top of the information.
If negative, you can reach out and try to mitigate the damage.
Here's a Fresh Alert email set up for mentions of Rand Fishkin. (In this case,
only included mentions that don't link to moz.com are included.)
You could also use reputation-based alerts to send daily emails to your
clients and monitor the conversation about your brand across the web.
3. Competitive intelligence
You can easily set up Fresh Alerts to notify you when your competition is in
the news. Better yet, use the search operators to notify you when specific media
outlets mention your competition.
In the example below, FEW shows us whenever "Amazon" is mentioned specifically
on TechCrunch.
You can also monitor when and where your competitors earn new links. For
example, if you wanted to set up a link alert for yourcompetition.com, simply
use the Root Domain search operator, like so:
rd:yourcompetition.com (alerts for all new links to the root domain)
By understanding how your competitors earn links and mentions, you may
discover new opportunities that are easy to replicate.
4. Reporting and content performance
This is a tip I don't hear people talk about, so I thought I'd share it.
Whenever we publish a big piece of content here at Moz, I set up a Fresh Alert
to notify me whenever someone mentions it.
For example, we recently published the 2013 Search Engine Ranking Factors.
Because this was an important piece of content for us, I set up 2 different
Fresh Alerts:
One Fresh Alert notified me whenever people mentioned "Search Engine Ranking
Factors" but didn't link to Moz
Another alert to tell me when people linked to the report itself
In the first example, I can reach out to those people who mentioned us without
linking to see if I can start a relationship and possibly earn a link.
In the second example, as seen in the graph below, I can monitor our
link-building efforts.
5. Discover publishing and guest-post opportunities
Fresh Alerts has to be one of the easiest ways to find distribution,
publishing and guest-post opportunities for your content. Yes, high-quality
guest posting, when combined with quality content and smart placement, remains a
powerful tactic when integrated with other marketing opportunities.
For example, let's say your subject is "dragons" and you want to find blogs
that have posted guest posts in the past few days. You can simply create an
alert for "dragons" AND "guest post".
This alert will notify you whenever a new post is published mentioning both
"guest post" and "dragons".
This technique isn't limited to guest posting, either. Getting creative, you
could find other publishing opportunities for your specific niche.
The details
Starting now, we've made Fresh Alerts available to subscribers of Moz
Analytics. If you're not a PRO member, you can sign up for a 30-day trial to
give them a spin if you'd like, which also includes access to our new Moz
Analytics and full suite of inbound marketing tools.
Here's what you need to know about Fresh Alerts:
Activate up to 10 Alerts per Moz Analytics account
When Fresh Web Explorer finds new mentions or links, you receive an email
within 24 hours
Alerts are sorted by Feed Authority, our new metric created specifically for
FWE
All advanced operators used by Fresh Web Explorer are available with Fresh
Alerts
Have you tried Fresh Web Explorer already? If so, let us know your best ideas
for Fresh Alerts in the comments below.
Sign up for The Moz Top 10, a semimonthly mailer updating you on the top ten
hottest pieces of SEO news, tips, and rad links uncovered by the Moz team. Think
of it as your exclusive digest of stuff you don't have time to hunt down but
want to read!
You may view the latest post at
http://feedproxy.google.com/~r/seomoz/~3/5_aeFA5ZaXc/say-hello-to-fresh-alerts
You received this e-mail because you asked to be notified when new updates are
posted.
Best regards,
Build Great Backlinks
peter.clarke@designed-for-success.com
Tuesday, 22 October 2013
[Build Great Backlinks] TITLE
Build Great Backlinks has posted a new item, 'A [Poorly] Illustrated Guide to
Google's Algorithm'
Posted by Dr-Pete
Like all great literature, this post started as a bad joke on Twitter on a
Friday night:
If you know me, then this kind of behavior hardly surprises you (and I
probably owe you an apology or two). What's surprising is that Google's Matt
Cutts replied, and fairly seriously:
Matt's concern that even my painfully stupid joke could be misinterpreted
demonstrates just how confused many people are about the algorithm. This tweet
actually led to a handful of very productive conversations, including one with
Danny Sullivan about the nature of Google's "Hummingbird" update.
These conversations got me thinking about how much we oversimplify what "the
algorithm" really is. This post is a journey in pictures, from the most basic
conception of the algorithm to something that I hope reflects the major concepts
Google is built on as we head into 2014.
The Google algorithm
There's really no such thing as "the" algorithm, but that's how we think about
itâas some kind of monolithic block of code that Google occasionally
tweaks. In our collective SEO consciousness, it looks something like this:
So, naturally, when Google announces an "update", all we see are shades of
blue. We hear about a major algorithm update ever month or two, and yet Google
confirmed 665 updates (technically, they used the word "launches") in
2012âobviously, there's something more going on here than just changing a
few lines of code in some mega-program.
Inputs and outputs
Of course, the algorithm has to do something, so we need inputs and outputs.
In the case of search, the most fundamental input is Google's index of the
worldwide web, and the output is search engine result pages (SERPs):
Simple enough, right? Web pages go in, [something happens], search results
come out. Well, maybe it's not quite that simple. Obviously, the algorithm
itself is incredibly complicated (and we'll get to that in a minute), but even
the inputs aren't as straightforward as you might imagine.
First of all, the index is really roughly a dozen data centers distributed
across the world, and each data center is a miniature city unto itself, linked
by one of the most impressive global fiber optic networks ever built. So, let's
at least add some color and say it looks something more like this:
Each block in that index illustration is a cloud of thousands of machines and
an incredible array of hardware, software and people, but if we dive deep into
that, this post will never end. It's important to realize, though, that the
index isn't the only major input into the algorithm. To oversimplify, the system
probably looks more like this:
The link graph, local and maps data, the social graph (predominantly Google+)
and the Knowledge Graphâessentially, a collection of entity
databasesâall comprise major inputs that exist beyond Google's core index
of the worldwide web. Again, this is just a conceptualization (I don't claim to
know how each of these are actually structured as physical data), but each of
these inputs are unique and important pieces of the search puzzle.
For the purposes of this post, I'm going to leave out personalization, which
has its own inputs (like your search history and location). Personalization is
undoubtedly important, but it impacts many areas of this illustration and is
more of a layer than a single piece of the puzzle.
Relevance, ranking and re-ranking
As SEOs, we're mostly concerned (i.e. obsessed) with ranking, but we forget
that ranking is really only part of the algorithm's job. I think it's useful to
split the process into two steps: (1) relevance, and (2) ranking. For a page to
rank in Google, it first has to make the cut and be included in the list. Let's
draw it something like this:
In other words, first Google has to pick which pages match the search, and
then they pick which order those pages are displayed in. Step (1) relies on
relevanceâa page can have all the links, +1s, and citations in the world,
but if it's not a match to the query, it's not going to rank. The Wikipedia page
for Millard Fillmore is never going to rank for "best iPhone cases," no matter
how much authority Wikipedia has. Once Wikipedia clears the relevance bar,
though, that authority kicks in and the page will often rank well.
Interestingly, this is one reason that our large-scale correlation studies
show fairly low correlations for on-page factors. Our correlation studies only
measure how well a page ranks once it's passed the relevance threshold. In 2013,
it's likely that on-page factors are still necessary for relevance, but they're
not sufficient for top rankings. In other words, your page has to clearly be
about a topic to show up in results, but just being about that topic doesn't
mean that it's going to rank well.
Even ranking isn't a single process. I'm going to try to cover an incredibly
complicated topic in just a few sentences, a topic that I'll call "re-ranking."
Essentially, Google determines a core ranking and what we might call a "pure"
organic result. Then, secondary ranking algorithms kick inâthese include
local results, social results, and vertical results (like news and images).
These secondary algorithms rewrite or re-rank the original results:
To see this in action, check out my post on how Google counts local results.
Using the methodology in that post, you can clearly see how Google determines a
base set of rankings, and then the local algorithm kicks in and not only adds
new features but re-ranks the original results. This diagram is only the tip of
the icebergâBill Slawski has an excellent three-part series on re-ranking
that covers 40 different ways Google may re-rank results.
Special inputs: penalties and disavowals
There are also special inputs (for lack of a better term). For example, if
Google issues a manual penalty against a site, that has to be flagged somewhere
and fed into the system. This may be part of the index, but since this process
is managed manually and tied to Google Webmaster Tools, I think it's useful to
view it as a separate concept.
Likewise, Google's disavow tool is a separate input, in this case one
partially controlled by webmasters. This data must be periodically processed and
then fed back into the algorithm and/or link graph. Presumably, there's a
semi-automated editorial process involved to verify and clean this
user-submitted data. So, that gives us something like this:
Of course, there are many inputs that feed other parts of the system. For
example, XML sitemaps in Google Webmaster Tools help shape the index. My goal it
to give you a flavor for the major concepts. As you can see, even the "simple"
version is quickly getting complicated.
Updates: Panda, Penguin and Hummingbird
Finally, we have the algorithm updates we all know and love. In many cases, an
update really is just a change or addition to some small part of Google's code.
In the past couple of years, though, algorithm updates have gotten a bit more
tricky.
Let's start with Panda, originally launched in February of 2011. The Panda
update was more than just a tweak to the codeâit was (and probably still
is) a sub-algorithm with its own data structures, living outside of the core
algorithm (conceptually speaking). Every month or so, the Panda algorithm would
be re-run, Panda data would be updated, and that data would feed what you might
call a Panda ranking factor back into the core algorithm. It's likely that
Penguin operates similarly, in that it's a sub-algorithm and separate data set.
We'll put them outside of the big, blue oval:
I don't mean to imply that Panda and Penguin are the sameâthey operate
in very different ways. I'm simply suggesting that both of these algorithm
updates rely on their own code and data sources and are only periodically fed
back into the system.
Why didn't Google just re-write the algorithm to account for the Panda and/or
Penguin intent? Part of it is computationalâthe resources required to
process this data are beyond what the real-time infrastructure can probably
handle. As Google gets faster and more powerful, these sub-algorithms may become
fully integrated (and Panda is probably more integrated than it once was). The
other reason may involve testing and mitigating impact. It's likely that Google
only updates Penguin periodically because of the large impact that the first
Penguin update had. This may not be a process that they simply want to let loose
in real-time.
So, what about the recent Hummingbird update? There's still a lot we don't
know, but Google has made it pretty clear that Hummingbird is a fundamental
rewrite of how the core algorithm works. I don't think we've seen the full
impact of Hummingbird yet, personally, and the potential of this new code may be
realized over months or even years, but now we're talking about the core
algorithm(s). That leads us to our final image:
Image credit for hummingbird silhouette: Michele Tobias at Experimental Craft.
The end result surprised even me as I created it. This was the most basic
illustration I could make that didn't feel misleading or simplistic. The reality
of Google today far surpasses this diagramâevery piece is dozens of
smaller pieces. I hope, though, that this gives you a sense for what the
algorithm really is and does.
Additional resources
If you're new to the algorithm and would like to learn more, Google's own "How
Search Works" resource is actually pretty interesting (check out the
sub-sections, not just the scroller). I'd also highly recommend Chapter 1 of our
Beginner's Guide: "How Search Engines Operate." If you just want to know more
about how Google operates, Steven Levy's book "In The Plex" is an amazing read.
Special bonus nonsense!
While writing this post, the team and I kept thinking there must be some way
to make it more dynamic, but all of our attempts ended badly. Finally, I just
gave up and turned the post into an animated GIF. If you like that sort of
thing, then here you go...
Sign up for The Moz Top 10, a semimonthly mailer updating you on the top ten
hottest pieces of SEO news, tips, and rad links uncovered by the Moz team. Think
of it as your exclusive digest of stuff you don't have time to hunt down but
want to read!
You may view the latest post at
http://feedproxy.google.com/~r/seomoz/~3/_rxMm03y4n8/a-poorly-illustrated-guide-to-googles-algorithm
You received this e-mail because you asked to be notified when new updates are
posted.
Best regards,
Build Great Backlinks
peter.clarke@designed-for-success.com
Google's Algorithm'
Posted by Dr-Pete
Like all great literature, this post started as a bad joke on Twitter on a
Friday night:
If you know me, then this kind of behavior hardly surprises you (and I
probably owe you an apology or two). What's surprising is that Google's Matt
Cutts replied, and fairly seriously:
Matt's concern that even my painfully stupid joke could be misinterpreted
demonstrates just how confused many people are about the algorithm. This tweet
actually led to a handful of very productive conversations, including one with
Danny Sullivan about the nature of Google's "Hummingbird" update.
These conversations got me thinking about how much we oversimplify what "the
algorithm" really is. This post is a journey in pictures, from the most basic
conception of the algorithm to something that I hope reflects the major concepts
Google is built on as we head into 2014.
The Google algorithm
There's really no such thing as "the" algorithm, but that's how we think about
itâas some kind of monolithic block of code that Google occasionally
tweaks. In our collective SEO consciousness, it looks something like this:
So, naturally, when Google announces an "update", all we see are shades of
blue. We hear about a major algorithm update ever month or two, and yet Google
confirmed 665 updates (technically, they used the word "launches") in
2012âobviously, there's something more going on here than just changing a
few lines of code in some mega-program.
Inputs and outputs
Of course, the algorithm has to do something, so we need inputs and outputs.
In the case of search, the most fundamental input is Google's index of the
worldwide web, and the output is search engine result pages (SERPs):
Simple enough, right? Web pages go in, [something happens], search results
come out. Well, maybe it's not quite that simple. Obviously, the algorithm
itself is incredibly complicated (and we'll get to that in a minute), but even
the inputs aren't as straightforward as you might imagine.
First of all, the index is really roughly a dozen data centers distributed
across the world, and each data center is a miniature city unto itself, linked
by one of the most impressive global fiber optic networks ever built. So, let's
at least add some color and say it looks something more like this:
Each block in that index illustration is a cloud of thousands of machines and
an incredible array of hardware, software and people, but if we dive deep into
that, this post will never end. It's important to realize, though, that the
index isn't the only major input into the algorithm. To oversimplify, the system
probably looks more like this:
The link graph, local and maps data, the social graph (predominantly Google+)
and the Knowledge Graphâessentially, a collection of entity
databasesâall comprise major inputs that exist beyond Google's core index
of the worldwide web. Again, this is just a conceptualization (I don't claim to
know how each of these are actually structured as physical data), but each of
these inputs are unique and important pieces of the search puzzle.
For the purposes of this post, I'm going to leave out personalization, which
has its own inputs (like your search history and location). Personalization is
undoubtedly important, but it impacts many areas of this illustration and is
more of a layer than a single piece of the puzzle.
Relevance, ranking and re-ranking
As SEOs, we're mostly concerned (i.e. obsessed) with ranking, but we forget
that ranking is really only part of the algorithm's job. I think it's useful to
split the process into two steps: (1) relevance, and (2) ranking. For a page to
rank in Google, it first has to make the cut and be included in the list. Let's
draw it something like this:
In other words, first Google has to pick which pages match the search, and
then they pick which order those pages are displayed in. Step (1) relies on
relevanceâa page can have all the links, +1s, and citations in the world,
but if it's not a match to the query, it's not going to rank. The Wikipedia page
for Millard Fillmore is never going to rank for "best iPhone cases," no matter
how much authority Wikipedia has. Once Wikipedia clears the relevance bar,
though, that authority kicks in and the page will often rank well.
Interestingly, this is one reason that our large-scale correlation studies
show fairly low correlations for on-page factors. Our correlation studies only
measure how well a page ranks once it's passed the relevance threshold. In 2013,
it's likely that on-page factors are still necessary for relevance, but they're
not sufficient for top rankings. In other words, your page has to clearly be
about a topic to show up in results, but just being about that topic doesn't
mean that it's going to rank well.
Even ranking isn't a single process. I'm going to try to cover an incredibly
complicated topic in just a few sentences, a topic that I'll call "re-ranking."
Essentially, Google determines a core ranking and what we might call a "pure"
organic result. Then, secondary ranking algorithms kick inâthese include
local results, social results, and vertical results (like news and images).
These secondary algorithms rewrite or re-rank the original results:
To see this in action, check out my post on how Google counts local results.
Using the methodology in that post, you can clearly see how Google determines a
base set of rankings, and then the local algorithm kicks in and not only adds
new features but re-ranks the original results. This diagram is only the tip of
the icebergâBill Slawski has an excellent three-part series on re-ranking
that covers 40 different ways Google may re-rank results.
Special inputs: penalties and disavowals
There are also special inputs (for lack of a better term). For example, if
Google issues a manual penalty against a site, that has to be flagged somewhere
and fed into the system. This may be part of the index, but since this process
is managed manually and tied to Google Webmaster Tools, I think it's useful to
view it as a separate concept.
Likewise, Google's disavow tool is a separate input, in this case one
partially controlled by webmasters. This data must be periodically processed and
then fed back into the algorithm and/or link graph. Presumably, there's a
semi-automated editorial process involved to verify and clean this
user-submitted data. So, that gives us something like this:
Of course, there are many inputs that feed other parts of the system. For
example, XML sitemaps in Google Webmaster Tools help shape the index. My goal it
to give you a flavor for the major concepts. As you can see, even the "simple"
version is quickly getting complicated.
Updates: Panda, Penguin and Hummingbird
Finally, we have the algorithm updates we all know and love. In many cases, an
update really is just a change or addition to some small part of Google's code.
In the past couple of years, though, algorithm updates have gotten a bit more
tricky.
Let's start with Panda, originally launched in February of 2011. The Panda
update was more than just a tweak to the codeâit was (and probably still
is) a sub-algorithm with its own data structures, living outside of the core
algorithm (conceptually speaking). Every month or so, the Panda algorithm would
be re-run, Panda data would be updated, and that data would feed what you might
call a Panda ranking factor back into the core algorithm. It's likely that
Penguin operates similarly, in that it's a sub-algorithm and separate data set.
We'll put them outside of the big, blue oval:
I don't mean to imply that Panda and Penguin are the sameâthey operate
in very different ways. I'm simply suggesting that both of these algorithm
updates rely on their own code and data sources and are only periodically fed
back into the system.
Why didn't Google just re-write the algorithm to account for the Panda and/or
Penguin intent? Part of it is computationalâthe resources required to
process this data are beyond what the real-time infrastructure can probably
handle. As Google gets faster and more powerful, these sub-algorithms may become
fully integrated (and Panda is probably more integrated than it once was). The
other reason may involve testing and mitigating impact. It's likely that Google
only updates Penguin periodically because of the large impact that the first
Penguin update had. This may not be a process that they simply want to let loose
in real-time.
So, what about the recent Hummingbird update? There's still a lot we don't
know, but Google has made it pretty clear that Hummingbird is a fundamental
rewrite of how the core algorithm works. I don't think we've seen the full
impact of Hummingbird yet, personally, and the potential of this new code may be
realized over months or even years, but now we're talking about the core
algorithm(s). That leads us to our final image:
Image credit for hummingbird silhouette: Michele Tobias at Experimental Craft.
The end result surprised even me as I created it. This was the most basic
illustration I could make that didn't feel misleading or simplistic. The reality
of Google today far surpasses this diagramâevery piece is dozens of
smaller pieces. I hope, though, that this gives you a sense for what the
algorithm really is and does.
Additional resources
If you're new to the algorithm and would like to learn more, Google's own "How
Search Works" resource is actually pretty interesting (check out the
sub-sections, not just the scroller). I'd also highly recommend Chapter 1 of our
Beginner's Guide: "How Search Engines Operate." If you just want to know more
about how Google operates, Steven Levy's book "In The Plex" is an amazing read.
Special bonus nonsense!
While writing this post, the team and I kept thinking there must be some way
to make it more dynamic, but all of our attempts ended badly. Finally, I just
gave up and turned the post into an animated GIF. If you like that sort of
thing, then here you go...
Sign up for The Moz Top 10, a semimonthly mailer updating you on the top ten
hottest pieces of SEO news, tips, and rad links uncovered by the Moz team. Think
of it as your exclusive digest of stuff you don't have time to hunt down but
want to read!
You may view the latest post at
http://feedproxy.google.com/~r/seomoz/~3/_rxMm03y4n8/a-poorly-illustrated-guide-to-googles-algorithm
You received this e-mail because you asked to be notified when new updates are
posted.
Best regards,
Build Great Backlinks
peter.clarke@designed-for-success.com
Monday, 21 October 2013
[Build Great Backlinks] TITLE
Build Great Backlinks has posted a new item, 'It's Penguin-Hunting Season: How
to Be the Predator and Not the Prey'
Posted by russvirante
Penguin changed everything. For most search engine optimizers like myself,
especially those who operate in the gray areas of optimization, we had long
grown comfortable with using "ratios" and "percentages" as simple litmus tests
to protect ourselves against the wrath of Google. I can't tell you how many
times I both participated in and was questioned about what our current "anchor
text ratio" was. Many of you probably remember having the same types of
discussions back in the keyword-stuffing days.
We now know unequivocally that Google has used and continues to use
statistical tools far more advanced than simply looking at where an individual
ranking factor sits on a dial. (We certainly have more than enough Remove'em
users to prove that.) My understanding of Penguin and its content-focused
predecessor Panda is that Google now employs machine-learning techniques across
large data sets to uncover patterns of over-optimization that aren't easily
discerned by the human eye or the crude algorithms of the past. It is with this
understanding that I and my company, Virante, Inc., undertook the Open Penguin
Data project, and ultimately formed our Penguin Vulnerability Score.
The Open Penguin Data Project
Matt Cutts occasionally gives us a heads-up about future updates, and in the
Spring of 2013 we were informed that within a few weeks Penguin 2.0 would roll
out. I remember exactly when the idea hit me. I was reading "How is Big Data
Different from Previous Data" by Bryan Eisenberg, and it occurred to me that the
kind of stuff we were doing at Remove'em to detect bad links just didn't keep
muster with the sophistication of the "big data" analysis Google was using at
the time. So Virante went to work. We started monitoring a huge number of
keywords, so that when Penguin 2.0 hit we could catch winners and losers. In the
end, we used data from three different awesome providers: Authority Labs (for
the initial data set), Stat Search Analytics (for cross-validation) and
SerpMetrics (for determining that we weren't just picking up manual penalties).
We identified around 600 losing URL/keyword pairs and matched them with their
competitors who did not lose rankings.
We then opened the data up to the community at the Open Penguin Data project
website and asked members of the community to contribute their ideas for factors
that might influence the Penguin algorithm. You can go there right now and
download the latest data set, although at present I know there is a bug in the
mozRank and mozTrust columns that needs to be fixed. We have identified over 70
factors that may influence Penguin and are still building upon them, with the
latest variable update being October 14th. Unfortunately, only certain variables
can be added now as fresh data won't be relevant. The data behind the factors
came from a large number of sources beginning with Moz of course, and including
Majestic SEO, Ahrefs, Grep/Words, and Archive.org
We then began to analyze the data in a number of ways. The first was through
standard correlation coefficients to help determine direction of influence
(assuming there was any influence at all). It is important that I deal with the
issue of correlation vs. causation here, because I am sure one of you will bring
it up.
Correlation vs. causation
The purpose of the Open Penguin Data Project was not and is not to determine
which factors cause a Penguin penalty. Rather, we want to determine which
factors predict a Penguin penalty so that we can build a reasonable model of
vulnerability. Once we know a website's vulnerability to Penguin, we can start
applying different techniques to lower that vulnerability that fall closer to
the realm of causal factors.
For example, we will talk about the difference of mozTrust and mozRank as
being a fairly good predictor of Penguin. No one in their right mind believes
that Google consumes Moz's data to determine who and who not to penalize.
However, once we know that a site is likely to be penalized (because we know the
mozTrust and mozRank differential), we can start to apply tactics that will
likely counter Penguin, such as using the disavow tool or removing spammy links.
We aren't talking about causation, we are talking about prediction.
The analysis of the risk factors
We then began analyzing the data using a couple of methods. First, we used
standard mean Spearman correlations to give us an idea of the lay of the land.
This allowed us to also build a crude regression model that actually works quite
well without much tweaking. This model essentially comes from adding up the
correlation coefficients for each of the factors. Obviously, more sophisticated
modeling is better than this, but to build a crude overview, this works quite
nicely and can be done on the fly. The real magic happens, though, when we apply
the same sorts of machine-learning techniques to the data set that Google uses
in building models like Penguin.
Let me be clear, I do not presume to know what statistical techniques Google
used to build their model. However, there are certain types of techniques that
are regularly used to answer these types of multivariate classification problems
and I chose to use them. In particular, I chose to use a gradient boosting
algorithm. You can read up on the methodology or the specific implementation we
used via scikit-learn, but I'll save you the headache and tell you what you need
to know.
Most of us think about statistical analysis as putting some variables in Excel
and making a nice graph with a linear regression that shows an upward or
downward trend. You can see this below. Unfortunately, this grossly
over-simplifies complex problems and often produces a crude result where
everything above the line is considered different from that below the line, when
clearly they are not. As you see in the example graph below, there are plenty of
penalized sites that get missed by falling below the line and completely decent
sites that are above the line that get hit.
Classification systems work differently. We aren't necessarily concerned with
higher or lower numbers, we are concerned with patterns that might predict
something. In this case, we know sites that were hit by Penguin, so now we use a
whole bunch of factors and see how the patterns between them might accurately
predict them. We don't need to draw an arbitrary line, we can individually
analyze the points using machine learning, as you see in the example graph
below.
The hard part is that machine learning tells us a lot about prediction, but
not a lot about how we came to that prediction. That is where some extra work
comes into play. With the Open Penguin Data project, we grouped some of the
factors by common characteristics and measured the effectiveness of their
predictions in isolation from the other factors. For example, we grouped trust
metrics together and anchor text metrics together. We then grouped them in
combinations as well. This then gave us a model we could use to determine not
only increased Penguin vulnerability, but also what factors contributed to that
vulnerability and to what degree.
So, let's talk through some of them here.
Anchor text
By now, everyone and their paid search guy knows that manipulated commercial
anchor text is a risk factor for both algorithmic and manual penalties. So, of
course, we looked at this closely from the start. We actually broke down the
anchor text into three subcategories: exact-match anchor text (meaning the
keyword is exactly the keyword for which you would like to rank), phrase-match
anchor text (meaning the keyword for which you would like to rank occurs
somewhere within the anchor text) and commercial anchor text (the anchor text
has a high CPC value).
Exact-match anchor text
We broke exact-match anchor text down into a couple of metrics:
The most common anchor to the page is exact match
The highest mozRank passed anchor to the page is exact match
There is at least one exact match anchor to the page
The most common anchor to the domain is exact match
The highest mozRank passed anchor to the domain is exact match
There is at least one exact match anchor to the domain
Across the board, every single metric related to anchor text provided some
positive predictive power except for highest mozRank passed anchor to the
domain. Importantly, no single factor had a particularly strong mean Spearman
correlation coefficient. For example, the highest was that the domain merely had
a single link with the exact match anchor text (.11 correlation coefficient).
This is a very weak signal, but our analysis looks to find patterns in these
weak signals, so we are not necessarily hindered because each measurement is not
sufficiently predictive.
For the biggest victims of Penguin, we often see that exact match anchor text
is the second- or third-largest predictor. For example, the below webmaster's
predictive vulnerability score could be lowered by 50% simply by impacting exact
match anchor text links. For this particular webmaster, the anchor text hit most
positive signals we measure regarding anchor text.
Now let me say it one more time: I am not saying that Google is using anchor
text to determine who to penalize, rather that it is a strong predictor.
Prediction is not causation. However, we can say that the groupings of
exact-match anchor text metrics allow us to detect Penguin vulnerability quite
well.
Phrase-match anchor text
We broke down phrase-match anchor text in the exact same fashion. This was one
of the more surprising features we noticed. In many cases, phrase-match anchor
text metrics appeared to be more predictive than exact-match anchor text. Many
SEOs, myself included, have long depended on what we call "brand blend" to
protect against over-optimization penalties. Instead of just building links for
the keyword "SEO", we might build links for "Virante SEO" or "SEO by Virante".
This may have insulated us against manual anchor text over-optimization
penalties, but it does not appear to be the case with Penguin.
In the example I mentioned above, the webmaster hit nearly every exact match
anchor text metric. They also hit every phrase match metric as well. The
combination of these factors increased their prediction of being impact by
Penguin by a full 100%.
Shoving your high-value keywords inside other phrases doesn't guarantee you
any protection. Now, there are a lot of potential takeaways from this. It could
be an artifact of merely doubling the exact match influence (i.e. if you score
high on exact match, you will also score high on phrase match). We do see some
of this occurring, but it doesn't appear to explain all of the additional
predictive power. It could be that they are targeting other related keywords and
thereby increase their exposure to other parts of the Penguin algorithm. All we
know, though, is that the predictive power of the model increases greatly when
we take into account phrase-match anchor text. Nothing more, nothing less.
Commercial anchor text
This is my favorite measure of all, as it shows how Google can use one of its
most powerful ancillary data sets, bid prices for keywords, to detect
manipulation of the link graph. We built 4 metrics around commercial anchor
text.
The page has a high-value anchor in a single link
The majority of the anchors are valuable
The majority of links are very high-value anchors
Has a high CPC site-wide
Both having high-value anchors and very high-value anchors had strong
predictive values of penguin vulnerability. In keeping with the example we have
been using so far, you can see that removing commercial anchor text would have a
profound impact on our prediction as to whether or not the site will be impacted
by Penguin.
If you've been paying close attention, you may have noticed that a lot of
these are related. Having exact-match and phrase-match anchor text likely means
you have highly commercial anchors. All of these metrics are related to one
another and it is their combined weak signals that make it easier to detect
Penguin vulnerability.
Link sources
The next issue we tried to target was the quality of link sources. The most
obvious step was trying to detect commonly spammed link sources: directories,
forums, guestbooks, press releases, articles, and comments. Using a set of
footprints to identify these types of links and spidering all of the backlinks
of the training set, we were able to build a few metrics identifying sites that
either simply had these types of links or had a preponderance of these types of
links.
First, it was interesting that every type of link was positively correlated,
but only very weakly. You can't just look at a bunch of article directory
submissions and assume that is the cause of a Penguin penalty. However, the
combinationâthat is a site that would rely on four or five of these types
of techniques for nearly all of their PageRankâwould appear to have a
greater risk factor.
At this point, I want to stop and draw attention to something: Each of these
groupings of factors appear to have some good predictive value, but none of them
comes even close to explaining the whole vulnerability. Fixing your exact-match
anchor text links, or phrase-match links, or commercial anchor links, or poor
link sources by themselves will not insulate you from detection. It is the
combination of these factors that appears to increase the vulnerability to
Penguin. Most sites that we see hit by Penguin have vulnerability scores that
are 250%+, although in Penguin 2.1 we saw them as low as 150%. To get to these
levels you have to trip a wide variety of factors, but you don't have to be
egregiously violating any one single SEO tactic.
Site-wides
This was one of the most disappointing features we used. I was certain, as
were many, that site-wide links would be the nail in the coffin. Clearly
site-wide links are the culprit behind the Penguin penalty, right? Well, the
data just doesn't bear that out.
Site-wides are just too common. The best sites on the web enjoy tons of
site-wide links, often in the form of Blog-Rolls. In fact, high site-wide rates
correlate negatively with Penguin penalties. Certainly this doesn't mean you
should run out and try to get a bunch of site-wide links, but it does beg the
question: Are site-wides really all that bad?
Here is where we find the real difference: anchor text. Commercial anchor text
site-wides positively correlate with Penguin penalties. While we cannot say they
cause them, there is definitely a predictive leap between just any old site-wide
link and a site-wide link with specific, commercially valuable anchor text.
This also helps illustrate another issue we SEOs often run into: anecdotal
evidence. It is really easy to look at a link profile, see that site-wide, and
immediately assume it is the culprit. It is then seemingly reinforced when we
scratch the surface with too simple an analysis like looking at the
preponderance of that feature among sites that are penalized. It can and does
often lead us down the wrong path.
Trust, trust, trust
Of all the eye-opening, mind-blowing discoveries revealed by the Open Penguin
Data project, this one was the biggest. At minimum, we all need to tip our hats
to the folks at Moz and Majestic for providing us with great link statistics.
Two of the strongest metrics we found in helping predict Penguin vulnerability
were MozRank greater than MozTrust (Moz) and Domain Citation Flow over Domain
Trust Flow (Majestic).
Both Moz and Majestic give us statistics that mimic to a certain degree the
raw flow of PageRank (MozTrust and Citation Flow) and an alternative often
referred to as Trust Rank (MozRank and Trust Flow). They are essentially the
same thing, except Trust metrics start with a trusted set of URLs like .govs and
.edus and gives extra value to sites that get links from these trusted sources.
These metrics by themselves, while useful in other endeavors, don't really give
us much information about Penguin.
However, if we flag URLs and domains where the trust metrics are lower than
the raw link metrics, we score some of the highest correlations of all factors
tested. Even cruder metrics like whether or not the domain has a single .gov
link help predict Penguin vulnerability. While it would be insane to conclude
that Google has a subscription to Moz and Majestic and use them to build their
Penguin algorithm, this appears to be true: In the aggregate, cheap, low quality
links are a Penguin risk factor.
What we should learn
There are some really amazing takeaways that we can build from this kind of
analysisâthe kind of takeaways that should change your understanding of
Penguin and Google's algorithm for many of you who are not yet seasoned
professionals. So let's dive in...
Penguin isn't spam detection, it's you detection
Try this fact on for size. If you hit every anchor text trigger in the Open
Penguin data set, our predictive model actually DROPS in effectiveness. At first
glance this seems counter-intuitive. Certainly Google should catch these extreme
spammers. The reality is, though, that cruder algorithms generally clear out
this type of search spam. If you have done any traditional off-site SEO in the
last three years, it will probably create additional Penguin vulnerability. The
Penguin update is targeted at catching patterns of optimization that aren't so
easily detected. The most egregious offenders are more likely to be caught by
other algorithms than Penguin. So when the next Penguin update comes out and you
hear people complain about how some spam site wasn't affected, you can be
confident that this isn't a flaw in Penguin, rather a deliberate choice on
Google's behalf to create separate algorithms to target different types of
over-optimization.
The rise of the link assassin
It was Ian Curl, a former Virante employee and now head of Link Assassins who
first pointed out to me the clear future of SEO: pruning the link graph. Google
has essentially given us the tools via GWT to both view our links and disavow
them. A new class of link removal and disavow professionals has grown over the
last year: SEOs who can spot a toxic link and guide you through the process of
not just cleaning up a penalty but proactively managing your link profile to
avoid penalties in the first place. These "link assassins" will play a vital
role in the future of SEO in just the same way that one would expect a
professional gardener to prune back excessive growth.
The demise of cheap, scalable white-hat link building
Let me be clear: If it works, Google wants to stop it. We have already heard
the shots across the bow for lily-white link building techniques like guest
posting from Matt Cutts. Right now, the only hold-out I see left is broken link
building which is only scalable under certain circumstances. Google is doing its
best to identify the exact same footprints you use to link-build and adding them
into their own link pattern detection. It isn't an easy task, which is why
Penguin only rolls out every few months, but it appears to be one to which
Google is committed.
The growth of integrated SEO
There is no way around it. If you are interested in long term, effective,
white-hat SEO, you are going to have to build integrated campaigns largely
focused around content marketing that include multiple forms of advertising.
There is a great write up on this by Chris Boggs over at Internet Marketing
Ninjas on Integrating Content Marketing into Traditional Advertising Campaigns.
As Google continues to get better at detecting unnatural patterns, it will be
harder and harder to get away with simply turning one dial at a time.
Next steps
The average webmaster or SEO needs to really step back and make an honest
account of their current SEO footprint. I don't mean to be fear-mongering; only
a fraction of a percent of all websites will ever get hit by Penguin. 75% of
adult males who smoke a pack a day will never get lung cancer, but that doesn't
mean you should keep on smoking because the odds are in your favor. While the
odds are greatly in your favor that Penguin will never strike your site, there
is no reason to not take simple precautions to determine whether your tactics
are putting your site at risk.
Sign up for The Moz Top 10, a semimonthly mailer updating you on the top ten
hottest pieces of SEO news, tips, and rad links uncovered by the Moz team. Think
of it as your exclusive digest of stuff you don't have time to hunt down but
want to read!
You may view the latest post at
http://feedproxy.google.com/~r/seomoz/~3/rjOf43Ytbb4/its-penguin-hunting-season-how-to-be-the-predator-not-the-prey
You received this e-mail because you asked to be notified when new updates are
posted.
Best regards,
Build Great Backlinks
peter.clarke@designed-for-success.com
to Be the Predator and Not the Prey'
Posted by russvirante
Penguin changed everything. For most search engine optimizers like myself,
especially those who operate in the gray areas of optimization, we had long
grown comfortable with using "ratios" and "percentages" as simple litmus tests
to protect ourselves against the wrath of Google. I can't tell you how many
times I both participated in and was questioned about what our current "anchor
text ratio" was. Many of you probably remember having the same types of
discussions back in the keyword-stuffing days.
We now know unequivocally that Google has used and continues to use
statistical tools far more advanced than simply looking at where an individual
ranking factor sits on a dial. (We certainly have more than enough Remove'em
users to prove that.) My understanding of Penguin and its content-focused
predecessor Panda is that Google now employs machine-learning techniques across
large data sets to uncover patterns of over-optimization that aren't easily
discerned by the human eye or the crude algorithms of the past. It is with this
understanding that I and my company, Virante, Inc., undertook the Open Penguin
Data project, and ultimately formed our Penguin Vulnerability Score.
The Open Penguin Data Project
Matt Cutts occasionally gives us a heads-up about future updates, and in the
Spring of 2013 we were informed that within a few weeks Penguin 2.0 would roll
out. I remember exactly when the idea hit me. I was reading "How is Big Data
Different from Previous Data" by Bryan Eisenberg, and it occurred to me that the
kind of stuff we were doing at Remove'em to detect bad links just didn't keep
muster with the sophistication of the "big data" analysis Google was using at
the time. So Virante went to work. We started monitoring a huge number of
keywords, so that when Penguin 2.0 hit we could catch winners and losers. In the
end, we used data from three different awesome providers: Authority Labs (for
the initial data set), Stat Search Analytics (for cross-validation) and
SerpMetrics (for determining that we weren't just picking up manual penalties).
We identified around 600 losing URL/keyword pairs and matched them with their
competitors who did not lose rankings.
We then opened the data up to the community at the Open Penguin Data project
website and asked members of the community to contribute their ideas for factors
that might influence the Penguin algorithm. You can go there right now and
download the latest data set, although at present I know there is a bug in the
mozRank and mozTrust columns that needs to be fixed. We have identified over 70
factors that may influence Penguin and are still building upon them, with the
latest variable update being October 14th. Unfortunately, only certain variables
can be added now as fresh data won't be relevant. The data behind the factors
came from a large number of sources beginning with Moz of course, and including
Majestic SEO, Ahrefs, Grep/Words, and Archive.org
We then began to analyze the data in a number of ways. The first was through
standard correlation coefficients to help determine direction of influence
(assuming there was any influence at all). It is important that I deal with the
issue of correlation vs. causation here, because I am sure one of you will bring
it up.
Correlation vs. causation
The purpose of the Open Penguin Data Project was not and is not to determine
which factors cause a Penguin penalty. Rather, we want to determine which
factors predict a Penguin penalty so that we can build a reasonable model of
vulnerability. Once we know a website's vulnerability to Penguin, we can start
applying different techniques to lower that vulnerability that fall closer to
the realm of causal factors.
For example, we will talk about the difference of mozTrust and mozRank as
being a fairly good predictor of Penguin. No one in their right mind believes
that Google consumes Moz's data to determine who and who not to penalize.
However, once we know that a site is likely to be penalized (because we know the
mozTrust and mozRank differential), we can start to apply tactics that will
likely counter Penguin, such as using the disavow tool or removing spammy links.
We aren't talking about causation, we are talking about prediction.
The analysis of the risk factors
We then began analyzing the data using a couple of methods. First, we used
standard mean Spearman correlations to give us an idea of the lay of the land.
This allowed us to also build a crude regression model that actually works quite
well without much tweaking. This model essentially comes from adding up the
correlation coefficients for each of the factors. Obviously, more sophisticated
modeling is better than this, but to build a crude overview, this works quite
nicely and can be done on the fly. The real magic happens, though, when we apply
the same sorts of machine-learning techniques to the data set that Google uses
in building models like Penguin.
Let me be clear, I do not presume to know what statistical techniques Google
used to build their model. However, there are certain types of techniques that
are regularly used to answer these types of multivariate classification problems
and I chose to use them. In particular, I chose to use a gradient boosting
algorithm. You can read up on the methodology or the specific implementation we
used via scikit-learn, but I'll save you the headache and tell you what you need
to know.
Most of us think about statistical analysis as putting some variables in Excel
and making a nice graph with a linear regression that shows an upward or
downward trend. You can see this below. Unfortunately, this grossly
over-simplifies complex problems and often produces a crude result where
everything above the line is considered different from that below the line, when
clearly they are not. As you see in the example graph below, there are plenty of
penalized sites that get missed by falling below the line and completely decent
sites that are above the line that get hit.
Classification systems work differently. We aren't necessarily concerned with
higher or lower numbers, we are concerned with patterns that might predict
something. In this case, we know sites that were hit by Penguin, so now we use a
whole bunch of factors and see how the patterns between them might accurately
predict them. We don't need to draw an arbitrary line, we can individually
analyze the points using machine learning, as you see in the example graph
below.
The hard part is that machine learning tells us a lot about prediction, but
not a lot about how we came to that prediction. That is where some extra work
comes into play. With the Open Penguin Data project, we grouped some of the
factors by common characteristics and measured the effectiveness of their
predictions in isolation from the other factors. For example, we grouped trust
metrics together and anchor text metrics together. We then grouped them in
combinations as well. This then gave us a model we could use to determine not
only increased Penguin vulnerability, but also what factors contributed to that
vulnerability and to what degree.
So, let's talk through some of them here.
Anchor text
By now, everyone and their paid search guy knows that manipulated commercial
anchor text is a risk factor for both algorithmic and manual penalties. So, of
course, we looked at this closely from the start. We actually broke down the
anchor text into three subcategories: exact-match anchor text (meaning the
keyword is exactly the keyword for which you would like to rank), phrase-match
anchor text (meaning the keyword for which you would like to rank occurs
somewhere within the anchor text) and commercial anchor text (the anchor text
has a high CPC value).
Exact-match anchor text
We broke exact-match anchor text down into a couple of metrics:
The most common anchor to the page is exact match
The highest mozRank passed anchor to the page is exact match
There is at least one exact match anchor to the page
The most common anchor to the domain is exact match
The highest mozRank passed anchor to the domain is exact match
There is at least one exact match anchor to the domain
Across the board, every single metric related to anchor text provided some
positive predictive power except for highest mozRank passed anchor to the
domain. Importantly, no single factor had a particularly strong mean Spearman
correlation coefficient. For example, the highest was that the domain merely had
a single link with the exact match anchor text (.11 correlation coefficient).
This is a very weak signal, but our analysis looks to find patterns in these
weak signals, so we are not necessarily hindered because each measurement is not
sufficiently predictive.
For the biggest victims of Penguin, we often see that exact match anchor text
is the second- or third-largest predictor. For example, the below webmaster's
predictive vulnerability score could be lowered by 50% simply by impacting exact
match anchor text links. For this particular webmaster, the anchor text hit most
positive signals we measure regarding anchor text.
Now let me say it one more time: I am not saying that Google is using anchor
text to determine who to penalize, rather that it is a strong predictor.
Prediction is not causation. However, we can say that the groupings of
exact-match anchor text metrics allow us to detect Penguin vulnerability quite
well.
Phrase-match anchor text
We broke down phrase-match anchor text in the exact same fashion. This was one
of the more surprising features we noticed. In many cases, phrase-match anchor
text metrics appeared to be more predictive than exact-match anchor text. Many
SEOs, myself included, have long depended on what we call "brand blend" to
protect against over-optimization penalties. Instead of just building links for
the keyword "SEO", we might build links for "Virante SEO" or "SEO by Virante".
This may have insulated us against manual anchor text over-optimization
penalties, but it does not appear to be the case with Penguin.
In the example I mentioned above, the webmaster hit nearly every exact match
anchor text metric. They also hit every phrase match metric as well. The
combination of these factors increased their prediction of being impact by
Penguin by a full 100%.
Shoving your high-value keywords inside other phrases doesn't guarantee you
any protection. Now, there are a lot of potential takeaways from this. It could
be an artifact of merely doubling the exact match influence (i.e. if you score
high on exact match, you will also score high on phrase match). We do see some
of this occurring, but it doesn't appear to explain all of the additional
predictive power. It could be that they are targeting other related keywords and
thereby increase their exposure to other parts of the Penguin algorithm. All we
know, though, is that the predictive power of the model increases greatly when
we take into account phrase-match anchor text. Nothing more, nothing less.
Commercial anchor text
This is my favorite measure of all, as it shows how Google can use one of its
most powerful ancillary data sets, bid prices for keywords, to detect
manipulation of the link graph. We built 4 metrics around commercial anchor
text.
The page has a high-value anchor in a single link
The majority of the anchors are valuable
The majority of links are very high-value anchors
Has a high CPC site-wide
Both having high-value anchors and very high-value anchors had strong
predictive values of penguin vulnerability. In keeping with the example we have
been using so far, you can see that removing commercial anchor text would have a
profound impact on our prediction as to whether or not the site will be impacted
by Penguin.
If you've been paying close attention, you may have noticed that a lot of
these are related. Having exact-match and phrase-match anchor text likely means
you have highly commercial anchors. All of these metrics are related to one
another and it is their combined weak signals that make it easier to detect
Penguin vulnerability.
Link sources
The next issue we tried to target was the quality of link sources. The most
obvious step was trying to detect commonly spammed link sources: directories,
forums, guestbooks, press releases, articles, and comments. Using a set of
footprints to identify these types of links and spidering all of the backlinks
of the training set, we were able to build a few metrics identifying sites that
either simply had these types of links or had a preponderance of these types of
links.
First, it was interesting that every type of link was positively correlated,
but only very weakly. You can't just look at a bunch of article directory
submissions and assume that is the cause of a Penguin penalty. However, the
combinationâthat is a site that would rely on four or five of these types
of techniques for nearly all of their PageRankâwould appear to have a
greater risk factor.
At this point, I want to stop and draw attention to something: Each of these
groupings of factors appear to have some good predictive value, but none of them
comes even close to explaining the whole vulnerability. Fixing your exact-match
anchor text links, or phrase-match links, or commercial anchor links, or poor
link sources by themselves will not insulate you from detection. It is the
combination of these factors that appears to increase the vulnerability to
Penguin. Most sites that we see hit by Penguin have vulnerability scores that
are 250%+, although in Penguin 2.1 we saw them as low as 150%. To get to these
levels you have to trip a wide variety of factors, but you don't have to be
egregiously violating any one single SEO tactic.
Site-wides
This was one of the most disappointing features we used. I was certain, as
were many, that site-wide links would be the nail in the coffin. Clearly
site-wide links are the culprit behind the Penguin penalty, right? Well, the
data just doesn't bear that out.
Site-wides are just too common. The best sites on the web enjoy tons of
site-wide links, often in the form of Blog-Rolls. In fact, high site-wide rates
correlate negatively with Penguin penalties. Certainly this doesn't mean you
should run out and try to get a bunch of site-wide links, but it does beg the
question: Are site-wides really all that bad?
Here is where we find the real difference: anchor text. Commercial anchor text
site-wides positively correlate with Penguin penalties. While we cannot say they
cause them, there is definitely a predictive leap between just any old site-wide
link and a site-wide link with specific, commercially valuable anchor text.
This also helps illustrate another issue we SEOs often run into: anecdotal
evidence. It is really easy to look at a link profile, see that site-wide, and
immediately assume it is the culprit. It is then seemingly reinforced when we
scratch the surface with too simple an analysis like looking at the
preponderance of that feature among sites that are penalized. It can and does
often lead us down the wrong path.
Trust, trust, trust
Of all the eye-opening, mind-blowing discoveries revealed by the Open Penguin
Data project, this one was the biggest. At minimum, we all need to tip our hats
to the folks at Moz and Majestic for providing us with great link statistics.
Two of the strongest metrics we found in helping predict Penguin vulnerability
were MozRank greater than MozTrust (Moz) and Domain Citation Flow over Domain
Trust Flow (Majestic).
Both Moz and Majestic give us statistics that mimic to a certain degree the
raw flow of PageRank (MozTrust and Citation Flow) and an alternative often
referred to as Trust Rank (MozRank and Trust Flow). They are essentially the
same thing, except Trust metrics start with a trusted set of URLs like .govs and
.edus and gives extra value to sites that get links from these trusted sources.
These metrics by themselves, while useful in other endeavors, don't really give
us much information about Penguin.
However, if we flag URLs and domains where the trust metrics are lower than
the raw link metrics, we score some of the highest correlations of all factors
tested. Even cruder metrics like whether or not the domain has a single .gov
link help predict Penguin vulnerability. While it would be insane to conclude
that Google has a subscription to Moz and Majestic and use them to build their
Penguin algorithm, this appears to be true: In the aggregate, cheap, low quality
links are a Penguin risk factor.
What we should learn
There are some really amazing takeaways that we can build from this kind of
analysisâthe kind of takeaways that should change your understanding of
Penguin and Google's algorithm for many of you who are not yet seasoned
professionals. So let's dive in...
Penguin isn't spam detection, it's you detection
Try this fact on for size. If you hit every anchor text trigger in the Open
Penguin data set, our predictive model actually DROPS in effectiveness. At first
glance this seems counter-intuitive. Certainly Google should catch these extreme
spammers. The reality is, though, that cruder algorithms generally clear out
this type of search spam. If you have done any traditional off-site SEO in the
last three years, it will probably create additional Penguin vulnerability. The
Penguin update is targeted at catching patterns of optimization that aren't so
easily detected. The most egregious offenders are more likely to be caught by
other algorithms than Penguin. So when the next Penguin update comes out and you
hear people complain about how some spam site wasn't affected, you can be
confident that this isn't a flaw in Penguin, rather a deliberate choice on
Google's behalf to create separate algorithms to target different types of
over-optimization.
The rise of the link assassin
It was Ian Curl, a former Virante employee and now head of Link Assassins who
first pointed out to me the clear future of SEO: pruning the link graph. Google
has essentially given us the tools via GWT to both view our links and disavow
them. A new class of link removal and disavow professionals has grown over the
last year: SEOs who can spot a toxic link and guide you through the process of
not just cleaning up a penalty but proactively managing your link profile to
avoid penalties in the first place. These "link assassins" will play a vital
role in the future of SEO in just the same way that one would expect a
professional gardener to prune back excessive growth.
The demise of cheap, scalable white-hat link building
Let me be clear: If it works, Google wants to stop it. We have already heard
the shots across the bow for lily-white link building techniques like guest
posting from Matt Cutts. Right now, the only hold-out I see left is broken link
building which is only scalable under certain circumstances. Google is doing its
best to identify the exact same footprints you use to link-build and adding them
into their own link pattern detection. It isn't an easy task, which is why
Penguin only rolls out every few months, but it appears to be one to which
Google is committed.
The growth of integrated SEO
There is no way around it. If you are interested in long term, effective,
white-hat SEO, you are going to have to build integrated campaigns largely
focused around content marketing that include multiple forms of advertising.
There is a great write up on this by Chris Boggs over at Internet Marketing
Ninjas on Integrating Content Marketing into Traditional Advertising Campaigns.
As Google continues to get better at detecting unnatural patterns, it will be
harder and harder to get away with simply turning one dial at a time.
Next steps
The average webmaster or SEO needs to really step back and make an honest
account of their current SEO footprint. I don't mean to be fear-mongering; only
a fraction of a percent of all websites will ever get hit by Penguin. 75% of
adult males who smoke a pack a day will never get lung cancer, but that doesn't
mean you should keep on smoking because the odds are in your favor. While the
odds are greatly in your favor that Penguin will never strike your site, there
is no reason to not take simple precautions to determine whether your tactics
are putting your site at risk.
Sign up for The Moz Top 10, a semimonthly mailer updating you on the top ten
hottest pieces of SEO news, tips, and rad links uncovered by the Moz team. Think
of it as your exclusive digest of stuff you don't have time to hunt down but
want to read!
You may view the latest post at
http://feedproxy.google.com/~r/seomoz/~3/rjOf43Ytbb4/its-penguin-hunting-season-how-to-be-the-predator-not-the-prey
You received this e-mail because you asked to be notified when new updates are
posted.
Best regards,
Build Great Backlinks
peter.clarke@designed-for-success.com
Friday, 18 October 2013
[Build Great Backlinks] TITLE
Build Great Backlinks has posted a new item, 'How Google is Changing Long-Tail
Search with Efforts Like Hummingbird - Whiteboard Friday'
Posted by randfish
The Hummingbird update was different from the major algorithm updates like
Penguin and Panda, revising core aspects of how Google understands what it finds
on the pages it crawls. In today's Whiteboard Friday, Rand explains what effect
that has on long-tail searches, and how those continue to evolve.
Whiteboard Friday - How Google is Changing Long-Tail Search with Efforts
Like Hummingbird
For reference, here's a still of this week's whiteboard!
Video Transcription
Howdy, Moz fans and welcome to another edition of Whiteboard Friday. This
week I wanted to talk a little bit about Google Hummingbird slightly, but more
broadly how Google has been making many efforts over the years to change how
they deal with long-tail search.
Now long tail, if you're not familiar already, is those queries that are
usually lengthier in terms of number of words in the phrase and refer to more
specific kinds of queries than the sort of head of the demand curve, which would
be shorter queries, many more people performing them, and, generally speaking,
the ones that in our profession, especially in the SEO world, the ones that we
tend to care about. So those are the shorter phrases, the head of the demand
curve, or the chunky middle of the demand curve versus the long tail.
Long tail, as Google has often mentioned, is a very big proportion of the
Web search traffic. It's anywhere from 20% to maybe 40% or even 50% of all the
queries on the Web are in that long tail, sort of fewer than maybe 10 to 50
searches per month, in that bucket. Somewhere around 18% or 20% of all searches
Google says are extremely long tail, meaning they've never seen them before,
extremely unique kinds of searches.
I think Google struggles with this a little bit. They struggle from an
advertising perspective because they'd like to be able to serve up great ads
targeting those long-tail phrases, but inside of AdWords, Google's Keyword Tool,
for self-service advertising, it's tough to choose those. Google doesn't often
show volume around them. Google themselves might have a tough time figuring out,
"hey, is this query relevant to these types of results," especially if it's in a
long tail.
So we've seen them get more and more sophisticated with content, context,
and textual analysis over the years such that today, with the release of, in
August according to Google, Hummingbird, which was an infrastructure update more
so than an algorithmic update. You can think of Penguin or Panda as being
algorithmic style updates, and Google Caffeine, which upgraded their speed, or
Hummingbird, which they say upgrades their text processing and their content and
context understanding mechanisms is affecting things today.
I'll try and illustrate this with an example. Let's say Google gets two
search queries, "best restaurants SEA," Seattle's airport, that's the airport
code, the three-letter code, and "where to eat at Sea-Tac Airport in Terminal
C." Let's say then that we've got a page here that's been produced by someone
who has listed the best restaurants at Sea-Tac, and they've ordered them by
terminals.
So if you're in Terminal A, Terminal B, Terminal C, it's actually easy to
walk between most of them except for N and S. I hope you never have to go N.
It's just a pain. S is even more of a pain. But in Terminal C, which I assume
would be Beecher's Cheese, because that place is incredible. It just opened.
It's super good. In Terminal C, they've got a Beecher's Cheese, so they've got a
listing for this.
A smart Google, an intelligent engineer at Google would go, "Man, you know,
I'd really like to be able to serve up this page for this result. But it doesn't
target the words 'where to eat' or 'Terminal C' specifically, especially not in
the title or the headline, the page title. How am I going to figure that out?"
Well, with upgrades like what we've seen with Hummingbird, Google may be able to
do more of this. So they essentially say, "I want to understand that this page
can satisfy both of these kinds of results."
This has some implications for the SEO world. On top of this, we're also
getting kind of biased away from long-tail search, because keyword (not
provided) means it's harder for an individual marketer to say: "Oh, are people
searching for this? Are people searching for that? Is this bringing me traffic?
Maybe I can optimize my page more towards it, optimize my content for it."
So this kind of combination and this direction that we're feeling from
Google has a few impacts. Those include more traffic opportunities,
opportunities for great content that isn't necessarily doing a fantastic job at
specific keyword targeting.
So this is kind of interesting from an SEO perspective, because we're not
saying, and I'm definitely not saying, stop doing keyword targeting, stop
putting good keywords in your titles and making your pages contextually relevant
to search queries. But I am saying if you do a good job of targeting this, best
restaurants at SEA or best restaurants Sea-Tac, you might find yourself getting
a lot more traffic for things like this. So there's almost an increased benefit
to producing that great content around this and serving, satisfying a number of
needs that a search query's intent might have.
Unfortunately, for some of us in the SEO world, it could get rougher for
sites that are targeting a lot of mid and long-tail queries through keyword
targeting that aren't necessarily doing a fantastic job from a content
perspective or from other algorithmic inputs. So if it's the case that I just
have to be ranking for a lot of long-tail phrases like this, but I don't have a
lot of the brand signals, link signals, social signals, user usage signals, I
just have strong keyword signals, well, Google might be trying to say, "Hey,
strong keyword signals doesn't mean as much to us anymore because now we can
take pages that we previously couldn't connect to that query and connect them
up."
In general, what we're talking about is Google rewarding better content over
more content, and that's kind of the way that things are trending in the SEO
world today.
So I'm sure there's going to be some great discussion. I really appreciate
the input of people who have done extensive analysis on top of Hummingbird.
Those folks include folks like Dr. Pete, of course, from Moz, Bill Slawski from
SEO by the Sea, Ammon Johns, who wrote a great post about this. I think there'll
be more great discussion in the comments. I look forward to joining you there.
Take care.
Video transcription by Speechpad.com
Sign up for The Moz Top 10, a semimonthly mailer updating you on the top ten
hottest pieces of SEO news, tips, and rad links uncovered by the Moz team. Think
of it as your exclusive digest of stuff you don't have time to hunt down but
want to read!
You may view the latest post at
http://feedproxy.google.com/~r/seomoz/~3/PU5oUTlWS4U/google-is-changing-long-tail-search-with-efforts-like-hummingbird-whiteboard-friday
You received this e-mail because you asked to be notified when new updates are
posted.
Best regards,
Build Great Backlinks
peter.clarke@designed-for-success.com
Search with Efforts Like Hummingbird - Whiteboard Friday'
Posted by randfish
The Hummingbird update was different from the major algorithm updates like
Penguin and Panda, revising core aspects of how Google understands what it finds
on the pages it crawls. In today's Whiteboard Friday, Rand explains what effect
that has on long-tail searches, and how those continue to evolve.
Whiteboard Friday - How Google is Changing Long-Tail Search with Efforts
Like Hummingbird
For reference, here's a still of this week's whiteboard!
Video Transcription
Howdy, Moz fans and welcome to another edition of Whiteboard Friday. This
week I wanted to talk a little bit about Google Hummingbird slightly, but more
broadly how Google has been making many efforts over the years to change how
they deal with long-tail search.
Now long tail, if you're not familiar already, is those queries that are
usually lengthier in terms of number of words in the phrase and refer to more
specific kinds of queries than the sort of head of the demand curve, which would
be shorter queries, many more people performing them, and, generally speaking,
the ones that in our profession, especially in the SEO world, the ones that we
tend to care about. So those are the shorter phrases, the head of the demand
curve, or the chunky middle of the demand curve versus the long tail.
Long tail, as Google has often mentioned, is a very big proportion of the
Web search traffic. It's anywhere from 20% to maybe 40% or even 50% of all the
queries on the Web are in that long tail, sort of fewer than maybe 10 to 50
searches per month, in that bucket. Somewhere around 18% or 20% of all searches
Google says are extremely long tail, meaning they've never seen them before,
extremely unique kinds of searches.
I think Google struggles with this a little bit. They struggle from an
advertising perspective because they'd like to be able to serve up great ads
targeting those long-tail phrases, but inside of AdWords, Google's Keyword Tool,
for self-service advertising, it's tough to choose those. Google doesn't often
show volume around them. Google themselves might have a tough time figuring out,
"hey, is this query relevant to these types of results," especially if it's in a
long tail.
So we've seen them get more and more sophisticated with content, context,
and textual analysis over the years such that today, with the release of, in
August according to Google, Hummingbird, which was an infrastructure update more
so than an algorithmic update. You can think of Penguin or Panda as being
algorithmic style updates, and Google Caffeine, which upgraded their speed, or
Hummingbird, which they say upgrades their text processing and their content and
context understanding mechanisms is affecting things today.
I'll try and illustrate this with an example. Let's say Google gets two
search queries, "best restaurants SEA," Seattle's airport, that's the airport
code, the three-letter code, and "where to eat at Sea-Tac Airport in Terminal
C." Let's say then that we've got a page here that's been produced by someone
who has listed the best restaurants at Sea-Tac, and they've ordered them by
terminals.
So if you're in Terminal A, Terminal B, Terminal C, it's actually easy to
walk between most of them except for N and S. I hope you never have to go N.
It's just a pain. S is even more of a pain. But in Terminal C, which I assume
would be Beecher's Cheese, because that place is incredible. It just opened.
It's super good. In Terminal C, they've got a Beecher's Cheese, so they've got a
listing for this.
A smart Google, an intelligent engineer at Google would go, "Man, you know,
I'd really like to be able to serve up this page for this result. But it doesn't
target the words 'where to eat' or 'Terminal C' specifically, especially not in
the title or the headline, the page title. How am I going to figure that out?"
Well, with upgrades like what we've seen with Hummingbird, Google may be able to
do more of this. So they essentially say, "I want to understand that this page
can satisfy both of these kinds of results."
This has some implications for the SEO world. On top of this, we're also
getting kind of biased away from long-tail search, because keyword (not
provided) means it's harder for an individual marketer to say: "Oh, are people
searching for this? Are people searching for that? Is this bringing me traffic?
Maybe I can optimize my page more towards it, optimize my content for it."
So this kind of combination and this direction that we're feeling from
Google has a few impacts. Those include more traffic opportunities,
opportunities for great content that isn't necessarily doing a fantastic job at
specific keyword targeting.
So this is kind of interesting from an SEO perspective, because we're not
saying, and I'm definitely not saying, stop doing keyword targeting, stop
putting good keywords in your titles and making your pages contextually relevant
to search queries. But I am saying if you do a good job of targeting this, best
restaurants at SEA or best restaurants Sea-Tac, you might find yourself getting
a lot more traffic for things like this. So there's almost an increased benefit
to producing that great content around this and serving, satisfying a number of
needs that a search query's intent might have.
Unfortunately, for some of us in the SEO world, it could get rougher for
sites that are targeting a lot of mid and long-tail queries through keyword
targeting that aren't necessarily doing a fantastic job from a content
perspective or from other algorithmic inputs. So if it's the case that I just
have to be ranking for a lot of long-tail phrases like this, but I don't have a
lot of the brand signals, link signals, social signals, user usage signals, I
just have strong keyword signals, well, Google might be trying to say, "Hey,
strong keyword signals doesn't mean as much to us anymore because now we can
take pages that we previously couldn't connect to that query and connect them
up."
In general, what we're talking about is Google rewarding better content over
more content, and that's kind of the way that things are trending in the SEO
world today.
So I'm sure there's going to be some great discussion. I really appreciate
the input of people who have done extensive analysis on top of Hummingbird.
Those folks include folks like Dr. Pete, of course, from Moz, Bill Slawski from
SEO by the Sea, Ammon Johns, who wrote a great post about this. I think there'll
be more great discussion in the comments. I look forward to joining you there.
Take care.
Video transcription by Speechpad.com
Sign up for The Moz Top 10, a semimonthly mailer updating you on the top ten
hottest pieces of SEO news, tips, and rad links uncovered by the Moz team. Think
of it as your exclusive digest of stuff you don't have time to hunt down but
want to read!
You may view the latest post at
http://feedproxy.google.com/~r/seomoz/~3/PU5oUTlWS4U/google-is-changing-long-tail-search-with-efforts-like-hummingbird-whiteboard-friday
You received this e-mail because you asked to be notified when new updates are
posted.
Best regards,
Build Great Backlinks
peter.clarke@designed-for-success.com
Thursday, 17 October 2013
[Build Great Backlinks] TITLE
Build Great Backlinks has posted a new item, 'Why Visitor Analytics Aren't
Enough for Modern Marketers'
Posted by randfish
For the first two decades of the web, the vast majority of those performing
web marketing tasks used visitor analytics tools (from log files and hit
counters all the way up to today's full-featured visitor analytics tools) to do
t...
You may view the latest post at
http://feedproxy.google.com/~r/seomoz/~3/6jLhQ-MUb54/visitor-analytics-arent-enough-for-modern-marketers
You received this e-mail because you asked to be notified when new updates are
posted.
Best regards,
Build Great Backlinks
peter.clarke@designed-for-success.com
Enough for Modern Marketers'
Posted by randfish
For the first two decades of the web, the vast majority of those performing
web marketing tasks used visitor analytics tools (from log files and hit
counters all the way up to today's full-featured visitor analytics tools) to do
t...
You may view the latest post at
http://feedproxy.google.com/~r/seomoz/~3/6jLhQ-MUb54/visitor-analytics-arent-enough-for-modern-marketers
You received this e-mail because you asked to be notified when new updates are
posted.
Best regards,
Build Great Backlinks
peter.clarke@designed-for-success.com
[Build Great Backlinks] TITLE
Build Great Backlinks has posted a new item, 'Introducing Our Niche Site Case
Study (With a Twist)'
One of the things that undoubtedly helped me to succeed online was seeing that
other people were having success. It wasnt so much specific examples not many
people shared them when I was starting out to be honest but just knowing that
someone had figured out this making money on the internet thing [...]
You may view the latest post at
http://www.viperchill.com/niche-site-twist/
You received this e-mail because you asked to be notified when new updates are
posted.
Best regards,
Build Great Backlinks
peter.clarke@designed-for-success.com
Study (With a Twist)'
One of the things that undoubtedly helped me to succeed online was seeing that
other people were having success. It wasnt so much specific examples not many
people shared them when I was starting out to be honest but just knowing that
someone had figured out this making money on the internet thing [...]
You may view the latest post at
http://www.viperchill.com/niche-site-twist/
You received this e-mail because you asked to be notified when new updates are
posted.
Best regards,
Build Great Backlinks
peter.clarke@designed-for-success.com
Wednesday, 16 October 2013
[Build Great Backlinks] TITLE
Build Great Backlinks has posted a new item, 'A Guide to Spanish Content
Marketing'
Posted by ZephSnappThis post was originally in YouMoz, and was promoted to the
main blog because it provides great value and interest to our community. The
author's views are entirely his or her own and may not reflect the views of Moz,
Inc.
Si prefieres leer este post en Espaol, se encuentra en el blog de Altura
Interactive.
Just like the rest of the SEO/inbound/internet marketing world, we have spent
the last year learning how to shift from link building to link earning, and
despite the fact that this stuff is really, really hard, weâve found some
success by building out processes. One challenge (advantage?) that we have is
that we work exclusively on Spanish-language projects. This means that while
many of the strategies are the same, some of the tactics vary. This post is
primarily meant for marketers interested in targeting the Spanish-speaking
world, but should also be helpful to full-stack marketers no matter the
language.
Are you ready for Spanish content marketing?
There are a ton of great reasons to get started on Spanish-language content
marketing. The Hispanic community in the US grew 67% from 2000 to 2011 according
to Pew Hispanic, and cleared 50 million people for the first time (although
reaching them does not necessarily mean you need to start marketing in Spanish).
Also, while growth has slowed in Latin American countries over the past couple
of years, their economies are stable enough that they arenât as affected
by downturns in the US economy as they once were. Just because Hispanic
marketing is hot, though, is not a good reason for your business to invest time,
money, and sweat equity in marketing to Spanish speakers. You need to validate
the concept and ensure it's the right move for you.
First, translate your main keywords. In some cases this can be fairly
straightforward, but there are some products that shouldn't be translated, since
the term exists on its own. A great example is âe-commerce:" While there
are ways to translate this term, most of the time we leave it in English. But
please, a word of advice: Donât use a machine translation. Get a human
being to translate your terms for you, then have someone else check their
translations. It is of paramount importance that your terms preserve the same
query intent, otherwise, any work on keyword research will be wasted.
Next, make sure that your website is in order, and that you have decided on an
international strategy. If you need more help on that front, check out
Aleydaâs Whiteboard Friday about International SEO Doâs and
Dontâs and her International SEO Checklist. They are both excellent
resources if you are thinking about taking your business abroad.
The research phase
We believe in doing persona-based marketing at all times. There is no reason
to belabor the point of how to build personas, since this topic has been written
about extensively. Suffice it to say, we follow the process explained by Mike
King almost to a T. The main difference in our technique is that in addition to
this process, we have to think about the country/region towards which we will be
targeting the content. This informs the type of data we should use for a given
piece of content. For example, if you are going after US-based Hispanics, you
may not even need to create the content in Spanish!
Armed with these personas, we find actual people who are active on social
media and see what type of content they are sharing. Followerwonk is a great way
to do this. These are not necessarily prospects, but Itâs absolutely
necessary to drill down as much as possible, otherwise your outreach will not be
nearly as effective.
Arm yourself with information
If you are going to create interesting content for Latin American audiences,
you are going to need data. Lots of it. Luckily for you, weâve gathered a
ton of data resources from all over Latin America. Some of them are country
specific, but others look at the region as a whole. The information is in
Spanish, but as we say in Mexico, "gajes del oficio" (comes with the territory).
At least weâve translated the description of the databases so youâll
be able to find what you are looking for. It is also a living document. As we
find more data sets, they will be added (and if you have any suggestions, please
put them in the comments, either here or on that post).
Since you already have your personas built, you can easily decide the data
that makes the most sense for your project, and then move on to another
important step:
Building the content
If you are a data driven marketer (the best kind in my opinion), when you are
diving into the data, your aim has to be to understand the story that the data
is telling you, and how you can use it to promote your client. Once you have the
story in place, we start thinking about how to best present the data. In some
many cases, a great blog post will do the trick. In those cases, we have one
person start writing titles. We write a minimum of five, because we want to
stimulate creative thoughtâit is rare that the first idea is the best.
Our lead editor reviews the proposals with the author, and together they
decide which best fits the subject, as well as the websites/people the post will
be targeting. Then the post is written, reviewed by the editor, and then another
content creator to ensure that the piece is focused, creative, and grammatically
sound.
In many cases, users will respond more favorably to a visualization than to
text. This is especially true if you are explaining a process or giving
instructions. Weâve found that video can be an awesome way get through to
these people. If you donât have the budget or the ability to shoot a video
yourself (although you shouldâas Phil Nottingham explained at MozCon, good
video can be created pretty cheaply), PowToon allows you to create an animated
explanation video, even if you donât have incredible design chops.
If you must create an infographic, at least try to be original in how you
present it. Weâve used Piktochart and Visual.ly just like everyone else,
but there are a ton of other ways to present data. Weâve created a list of
data visualization resources that includes some very unusual ways of presenting
data. In many cases, the main investment is in learning how to use the platform.
Shameless Plug: In my Mozinar next Tuesday Iâll be sharing the easiest
way to build resources with outreach prospects built in. Itâs seriously
awesome. You should sign up now. Por favor!
Prospecting for outreach
Generally speaking, we are looking for:
People
Usually the best way to find experts in a given vertical is to look at
Twitter, and the best way to qualify them is via Followerwonk. Enough blog posts
have been written about this already, so there is no need for us to get into
that here.
Websites
If you are really strapped for cash, all you need is a list of keywords for
your vertical and Googleâs advanced operators. We use these on occasion,
but most of the time, it is faster and more efficient to lean on tools built by
others.
Link Prospector supports multilingual queries, and if you want to get a great
list of prospects quickly, this is a great way to find them. (Full disclosure:
We helped build the multilingual tool, and while we didnât profit from it,
we do get to use it for free. Still, if you told me I could only use Moz and one
other tool, this would be it).
Buzzstream is an awesome tool which also supports multilingual queries, and
doubles as a way to remember what prospects are in what stage of a relationship.
We have found that the contact information that the tool pulls is not
particularly accurate for websites in Spanish, so if you are using this tool
donât depend on themâgo get the information for yourself. Another
platform that weâve been using that has proven helpful is GroupHigh. Their
platform is pricey, but the prospects that you can get from here are excellent,
especially if you are doing a bilingual English/Spanish outreach campaign. The
metrics they provide are based on Mozâs stats as well as social shares,
but they donât always coincide with what we find when we check sites by
hand.
To be sure, we prequalify every single website we are going to do outreach to.
And we craft every single pitch individually to ensure that they are more likely
to looked upon favorably by our prospective partners.
Once we have our prospects, we separate them into tiers. The top tier is of
the most important people and websites in a sphere. We know that getting in
touch with and convincing these targets to share our content will be
extraordinarily hard, simply because they are pitched to so often. The advantage
we have is that most of the pitches they receive totally suck. Knowing how to
approach each influencer can make or break your outreach efforts, which leads to
our next point:
Outreach to influencers
The goal of any outreach campaign is to get the person/website youâve
targeted to share your content piece, right? In most cases, no matter the
quality of your pitch, it will be ignored. This is because some websites are
abandoned, the webmaster might be too busy with other work (like a day job), or
they simply might not care enough to respond. These are the facts.
And then there is the question of culture and language. Weâve used
templates developed by some of the best link builders in the US and seen zero or
even negative response. So, it is crucial to localize not just the content, but
also the approach. By following our process, you can increase your engagement
rate when doing outreach, especially when it is for a piece of content you have
created. Here are a few tips that weâve found to be effective when doing
outreach to Spanish-speaking webmasters, bloggers, and journalists:
1) Write it in Spanish
I know that this might seem obvious, but my friends who are
bloggersâincluding for the oldest blog in Mexicoâreceive dozens of
pitches from professional PR companies IN ENGLISH. Unforgivable.
2) Make it relevant
Even if the piece of content that you are promoting is only loosely related to
the target site, make sure that you make an argument for why it would be
interesting to the readership of that site. Yes, this means you canât just
blast emails. Too bad.
3) Keep it short
In Spanish, we have a tendency to be a bit verbose. In fact, we use more words
to explain something than people usually do in English. That being said, it is
still better to be concise.
4) Have a hook
Whenever you are doing outreach, the goal is to provide value to your client
or company. Keep in mind, however, that webmasters donât care about how
great it will be for you if they share your latest infographic about dog food.
They care about their readers and community, so make sure that your pitch
addresses the benefits for them, not for you.
5) Address the webmaster how (s)he addresses users
In Spanish, you can address readers either formally or informally. By making
your outreach consistent with how they address their readers, you can be sure
that your pitch fits their style.
6) Be legit, be honest
Despite what Iâve heard about other markets, weâve found that
being TAGFEE is the best way to get results from an outreach campaign. That
doesnât mean that you canât sugarcoat your outreach ("Links, Please"
is probably not the best subject line), but we send emails from our own domain,
and own up to working on behalf of a client. We even link back to our profile
pages in our outreach emails.
7) Prioritize outreach method
The best method for outreach depends on who you are reaching out to. This is
our priority list when reaching out to bloggers, for example:
Contact form
Facebook
Email
Twitter
In our experience, the first two methods are easily the most effective. This is
another place where being open and honest works to our advantage. Since we are
using our own Facebook profiles to conduct outreach, prospects can look at our
pictures, read our updates, and see that we are human beings, just like them.
They are far less likely to say no to someone who likes the same band as them,
right?
Of course, if you are reaching out to a journalist (or even a web-based
magazine) it is probably going to be best to reach out via phone. Having a
prioritized list of methods makes things easier for the outreach specialist to
work.
There is obviously a lot more that goes into outstanding Spanish content
marketing, but this guide is here to give you the basics. If you want to dig
deeper into our Spanish digital marketing processes, please sign up for my
Mozinar. Muchas Gracias!
If you would prefer to read this post in Spanish, check it out on the Altura
Interactive blog.
Si quieres leer sobre estrategia de contenido en espaol, este post
tambi©n se encuentra en el blog de Altura Interactive.
Sign up for The Moz Top 10, a semimonthly mailer updating you on the top ten
hottest pieces of SEO news, tips, and rad links uncovered by the Moz team. Think
of it as your exclusive digest of stuff you don't have time to hunt down but
want to read!
You may view the latest post at
http://feedproxy.google.com/~r/seomoz/~3/BslXG7lcQEY/a-guide-to-spanish-content-marketing
You received this e-mail because you asked to be notified when new updates are
posted.
Best regards,
Build Great Backlinks
peter.clarke@designed-for-success.com
Marketing'
Posted by ZephSnappThis post was originally in YouMoz, and was promoted to the
main blog because it provides great value and interest to our community. The
author's views are entirely his or her own and may not reflect the views of Moz,
Inc.
Si prefieres leer este post en Espaol, se encuentra en el blog de Altura
Interactive.
Just like the rest of the SEO/inbound/internet marketing world, we have spent
the last year learning how to shift from link building to link earning, and
despite the fact that this stuff is really, really hard, weâve found some
success by building out processes. One challenge (advantage?) that we have is
that we work exclusively on Spanish-language projects. This means that while
many of the strategies are the same, some of the tactics vary. This post is
primarily meant for marketers interested in targeting the Spanish-speaking
world, but should also be helpful to full-stack marketers no matter the
language.
Are you ready for Spanish content marketing?
There are a ton of great reasons to get started on Spanish-language content
marketing. The Hispanic community in the US grew 67% from 2000 to 2011 according
to Pew Hispanic, and cleared 50 million people for the first time (although
reaching them does not necessarily mean you need to start marketing in Spanish).
Also, while growth has slowed in Latin American countries over the past couple
of years, their economies are stable enough that they arenât as affected
by downturns in the US economy as they once were. Just because Hispanic
marketing is hot, though, is not a good reason for your business to invest time,
money, and sweat equity in marketing to Spanish speakers. You need to validate
the concept and ensure it's the right move for you.
First, translate your main keywords. In some cases this can be fairly
straightforward, but there are some products that shouldn't be translated, since
the term exists on its own. A great example is âe-commerce:" While there
are ways to translate this term, most of the time we leave it in English. But
please, a word of advice: Donât use a machine translation. Get a human
being to translate your terms for you, then have someone else check their
translations. It is of paramount importance that your terms preserve the same
query intent, otherwise, any work on keyword research will be wasted.
Next, make sure that your website is in order, and that you have decided on an
international strategy. If you need more help on that front, check out
Aleydaâs Whiteboard Friday about International SEO Doâs and
Dontâs and her International SEO Checklist. They are both excellent
resources if you are thinking about taking your business abroad.
The research phase
We believe in doing persona-based marketing at all times. There is no reason
to belabor the point of how to build personas, since this topic has been written
about extensively. Suffice it to say, we follow the process explained by Mike
King almost to a T. The main difference in our technique is that in addition to
this process, we have to think about the country/region towards which we will be
targeting the content. This informs the type of data we should use for a given
piece of content. For example, if you are going after US-based Hispanics, you
may not even need to create the content in Spanish!
Armed with these personas, we find actual people who are active on social
media and see what type of content they are sharing. Followerwonk is a great way
to do this. These are not necessarily prospects, but Itâs absolutely
necessary to drill down as much as possible, otherwise your outreach will not be
nearly as effective.
Arm yourself with information
If you are going to create interesting content for Latin American audiences,
you are going to need data. Lots of it. Luckily for you, weâve gathered a
ton of data resources from all over Latin America. Some of them are country
specific, but others look at the region as a whole. The information is in
Spanish, but as we say in Mexico, "gajes del oficio" (comes with the territory).
At least weâve translated the description of the databases so youâll
be able to find what you are looking for. It is also a living document. As we
find more data sets, they will be added (and if you have any suggestions, please
put them in the comments, either here or on that post).
Since you already have your personas built, you can easily decide the data
that makes the most sense for your project, and then move on to another
important step:
Building the content
If you are a data driven marketer (the best kind in my opinion), when you are
diving into the data, your aim has to be to understand the story that the data
is telling you, and how you can use it to promote your client. Once you have the
story in place, we start thinking about how to best present the data. In some
many cases, a great blog post will do the trick. In those cases, we have one
person start writing titles. We write a minimum of five, because we want to
stimulate creative thoughtâit is rare that the first idea is the best.
Our lead editor reviews the proposals with the author, and together they
decide which best fits the subject, as well as the websites/people the post will
be targeting. Then the post is written, reviewed by the editor, and then another
content creator to ensure that the piece is focused, creative, and grammatically
sound.
In many cases, users will respond more favorably to a visualization than to
text. This is especially true if you are explaining a process or giving
instructions. Weâve found that video can be an awesome way get through to
these people. If you donât have the budget or the ability to shoot a video
yourself (although you shouldâas Phil Nottingham explained at MozCon, good
video can be created pretty cheaply), PowToon allows you to create an animated
explanation video, even if you donât have incredible design chops.
If you must create an infographic, at least try to be original in how you
present it. Weâve used Piktochart and Visual.ly just like everyone else,
but there are a ton of other ways to present data. Weâve created a list of
data visualization resources that includes some very unusual ways of presenting
data. In many cases, the main investment is in learning how to use the platform.
Shameless Plug: In my Mozinar next Tuesday Iâll be sharing the easiest
way to build resources with outreach prospects built in. Itâs seriously
awesome. You should sign up now. Por favor!
Prospecting for outreach
Generally speaking, we are looking for:
People
Usually the best way to find experts in a given vertical is to look at
Twitter, and the best way to qualify them is via Followerwonk. Enough blog posts
have been written about this already, so there is no need for us to get into
that here.
Websites
If you are really strapped for cash, all you need is a list of keywords for
your vertical and Googleâs advanced operators. We use these on occasion,
but most of the time, it is faster and more efficient to lean on tools built by
others.
Link Prospector supports multilingual queries, and if you want to get a great
list of prospects quickly, this is a great way to find them. (Full disclosure:
We helped build the multilingual tool, and while we didnât profit from it,
we do get to use it for free. Still, if you told me I could only use Moz and one
other tool, this would be it).
Buzzstream is an awesome tool which also supports multilingual queries, and
doubles as a way to remember what prospects are in what stage of a relationship.
We have found that the contact information that the tool pulls is not
particularly accurate for websites in Spanish, so if you are using this tool
donât depend on themâgo get the information for yourself. Another
platform that weâve been using that has proven helpful is GroupHigh. Their
platform is pricey, but the prospects that you can get from here are excellent,
especially if you are doing a bilingual English/Spanish outreach campaign. The
metrics they provide are based on Mozâs stats as well as social shares,
but they donât always coincide with what we find when we check sites by
hand.
To be sure, we prequalify every single website we are going to do outreach to.
And we craft every single pitch individually to ensure that they are more likely
to looked upon favorably by our prospective partners.
Once we have our prospects, we separate them into tiers. The top tier is of
the most important people and websites in a sphere. We know that getting in
touch with and convincing these targets to share our content will be
extraordinarily hard, simply because they are pitched to so often. The advantage
we have is that most of the pitches they receive totally suck. Knowing how to
approach each influencer can make or break your outreach efforts, which leads to
our next point:
Outreach to influencers
The goal of any outreach campaign is to get the person/website youâve
targeted to share your content piece, right? In most cases, no matter the
quality of your pitch, it will be ignored. This is because some websites are
abandoned, the webmaster might be too busy with other work (like a day job), or
they simply might not care enough to respond. These are the facts.
And then there is the question of culture and language. Weâve used
templates developed by some of the best link builders in the US and seen zero or
even negative response. So, it is crucial to localize not just the content, but
also the approach. By following our process, you can increase your engagement
rate when doing outreach, especially when it is for a piece of content you have
created. Here are a few tips that weâve found to be effective when doing
outreach to Spanish-speaking webmasters, bloggers, and journalists:
1) Write it in Spanish
I know that this might seem obvious, but my friends who are
bloggersâincluding for the oldest blog in Mexicoâreceive dozens of
pitches from professional PR companies IN ENGLISH. Unforgivable.
2) Make it relevant
Even if the piece of content that you are promoting is only loosely related to
the target site, make sure that you make an argument for why it would be
interesting to the readership of that site. Yes, this means you canât just
blast emails. Too bad.
3) Keep it short
In Spanish, we have a tendency to be a bit verbose. In fact, we use more words
to explain something than people usually do in English. That being said, it is
still better to be concise.
4) Have a hook
Whenever you are doing outreach, the goal is to provide value to your client
or company. Keep in mind, however, that webmasters donât care about how
great it will be for you if they share your latest infographic about dog food.
They care about their readers and community, so make sure that your pitch
addresses the benefits for them, not for you.
5) Address the webmaster how (s)he addresses users
In Spanish, you can address readers either formally or informally. By making
your outreach consistent with how they address their readers, you can be sure
that your pitch fits their style.
6) Be legit, be honest
Despite what Iâve heard about other markets, weâve found that
being TAGFEE is the best way to get results from an outreach campaign. That
doesnât mean that you canât sugarcoat your outreach ("Links, Please"
is probably not the best subject line), but we send emails from our own domain,
and own up to working on behalf of a client. We even link back to our profile
pages in our outreach emails.
7) Prioritize outreach method
The best method for outreach depends on who you are reaching out to. This is
our priority list when reaching out to bloggers, for example:
Contact form
In our experience, the first two methods are easily the most effective. This is
another place where being open and honest works to our advantage. Since we are
using our own Facebook profiles to conduct outreach, prospects can look at our
pictures, read our updates, and see that we are human beings, just like them.
They are far less likely to say no to someone who likes the same band as them,
right?
Of course, if you are reaching out to a journalist (or even a web-based
magazine) it is probably going to be best to reach out via phone. Having a
prioritized list of methods makes things easier for the outreach specialist to
work.
There is obviously a lot more that goes into outstanding Spanish content
marketing, but this guide is here to give you the basics. If you want to dig
deeper into our Spanish digital marketing processes, please sign up for my
Mozinar. Muchas Gracias!
If you would prefer to read this post in Spanish, check it out on the Altura
Interactive blog.
Si quieres leer sobre estrategia de contenido en espaol, este post
tambi©n se encuentra en el blog de Altura Interactive.
Sign up for The Moz Top 10, a semimonthly mailer updating you on the top ten
hottest pieces of SEO news, tips, and rad links uncovered by the Moz team. Think
of it as your exclusive digest of stuff you don't have time to hunt down but
want to read!
You may view the latest post at
http://feedproxy.google.com/~r/seomoz/~3/BslXG7lcQEY/a-guide-to-spanish-content-marketing
You received this e-mail because you asked to be notified when new updates are
posted.
Best regards,
Build Great Backlinks
peter.clarke@designed-for-success.com
Subscribe to:
Posts (Atom)
