Google Ranking #6 Penalty / Filter

A New Google Filter is Born

In early December some astute webmasters noticed that some of their longterm (in some cases many years) #1 or #2 ranking pages in Google now rank at #6. Just like with the Google -30 and the Google -950 penalties, some people will maintain this is fiction, but too many smart people experienced the same thing at the same time for it to be such.

Background Information

Tedster started a WMW thread on the topic on December 26th. From Tedster's post, some of the sites that were hit:

  1. Well established site with a long history.
  2. Long time good rankings for a big search term - usually #1
  3. Other searches that returned the same url at #1 may also be sent to #6, but not all of them
  4. Some reports of a #2 result going to #6.

My Site That Got Hit

My site which saw a ranking dive on December 18th had the homepage hit, and interior pages hit for some (but not all) related phrases. Here are some noteworthy conditions with my site that was hit:

  • The site was entirely ranked on SEO. There is no ad budget outside of PPC ads or link buying, and no brand recognition outside of the search results. Outside of one linkbait there is nothing remarkable about the site.
  • The homepage did not get any new quality links in over a year.
  • Much of the link building was done years ago when I was far spammier and far more aggressive with anchor text than I would be today, though I did use some semantic variation to pick up rankings for many different keyword permutations.
  • The internal pages still rank #1 for some semi-related longer queries, while they are also filtered and ranking #6 for some more obviously connected shorter search queries.
  • The site continues to buy PPC ads and gets decent conversion rates for the keywords that were hit, and gets great conversion rates for more focused related terms, some of which the site was hit for and some of which the site still ranks great for. This conversion data is being sent to Google via the AdWords conversion tracker.
  • This affected alternate permutations of acronyms (letters strung together or pulled apart).
  • For my site this affects rankings on alternate versions of words (ie: single vs plural). For at least one person on WMW they did not see it affect both single and plural versions of their keywords.
  • This affects words if mixed into a different order.
  • This affects many longer search query containing the core words or closely related words.
  • This did not affect obvious domain name or brand related queries, even if the brand contained one of the words overlapping with the penalized set. If a filtered word outside of the domain name / brand name is appended to the query then the rankings are killed, and the site is stuck at #6.

Usage Data or Improved Phrase Relationship Detection of Anchor Text?

Why I do Not Think it is Usage Data

Based on feedback in the WMW thread it is hard to isolate this to any one variable with certainty. Two possibilities that have been thrown out are rolling more usage data into the search results or a better understanding of word and phrase relationships. It is easy to think of usage data as a possibility given my site's lack of marketing and lack of integration into the organic web, but that would not explain why some pages and queries were hit while some similar pages and queries still rank, with Google getting strong conversion data via AdWords on some of these pages. Also, for that homepage I wrote an aggressive page title and meta description that draws in many clicks, and the landing page is exceptionally relevant for the query.

Why I Think it is Phrase Relationships

I think this issue is likely tied to a stagnant link profile with a too tightly aligned anchor text profile, with the anchor text being overly-optimized when compared against competing sites.

The fact that some related queries were hit, but not all, makes me think that rather than being about usage data this is about word and phrase relationship improvements. I think if Google got better at understanding word relationships, many of the pages that once fit the criteria to rank may now have anchor text that is too focused and too well aligned with the target keywords, especially if they compare your anchor text to the anchor text of other sites competing for the same phrases. Once possible manipulation is identified via artificial anchor text your rankings across the site can be suppressed for a basket of semantically related terms, as noted in some of Google's phrase based indexing patents.

Matt Cutts Does Not Know What Happened

This filter was also called the minus 5 penalty, but many of the sites that were hit still rank at #6 even if they were ranking #2 or #3 before they were hit. When Barry posted about this Matt Cutts said "Hmm. I'm not aware of anything that would exhibit that sort of behavior," but some past SEO issues, like the famed Google sandbox have been accidentally introduced as a side effect of Google upgrades:

What's a sandbox, Matt?

"Some people have asked, "does this apply to newer sites?" Essentially, the way to think about it is, around 2003 Google switched to a new method of updating its index. Before that we had monthly Google dances. So as a result, new data is always being folded into the index. It's not like there was one pivotal moment when anyone can say, "Hah! This is the change!" In fact, even at different data centers we have different binaries, different algorithms, different types of data always being tested.

"I think a lot of what's perceived as the sandbox is artifacts where, in our indexing, some data may take longer to be computed than other data."

Great Comments About the Filter

3 great posts from the WMW thread:

Your Feedback Needed

With my sample set of one site my current hypothesis might be out to lunch. If you have any sites that you feel were hit and want to share them for helping everyone figure out what is going on please do so in the comments below. If you have any ideas or feedback on what happened please leave a comment with that too.

We Will Not Make Editorial Judgements, But We Desire to Rank Our Content #1

With the announcement of Knol, Google displayed their desire to become a publisher. Why? To make free information more accessible. It doesn't hurt that publishers dominate other industries, like music - where in some cases giving artists nothing, while some artist get less than nothing, even if they made millions in sales.

Danny Sullivan had some reservations on Knol, as does Rich Skrenta, and just about every other successful results oriented independent web author.

While claiming Google will not make any editorial judgements of quality, and Google will treat Knols like any other web pages, Google's Udi Manber had this to say:

A knol on a particular topic is meant to be the first thing someone who searches for this topic for the first time will want to read. The goal is for knols to cover all topics, from scientific concepts, to medical information, from geographical and historical, to entertainment, from product information, to how-to-fix-it instructions. Google will not serve as an editor in any way, and will not bless any content.

They desire it to be a starting point for searchers and yet they will not promote it?

Think back to the YouTube purchase. After Google bought the site, did they start blessing / featuring any YouTube content? Yes they did. Google's Uinversal Search integrated YouTube so tightly in their search results that now people add YouTube to the search query for many music searches . Don't believe me that they shifted user behavior? Try using Google Suggest for music searches and see where YouTube shows up.

Manber wrote not to worry about spam, as Google has that issue covered:

Our job in Search Quality will be to rank the knols appropriately when they appear in Google search results. We are quite experienced with ranking web pages, and we feel confident that we will be up to the challenge. We are very excited by the potential to substantially increase the dissemination of knowledge.

Sure they will filter out some of the garbage people submit, but the good stuff will rank better than it should. I am not a betting man, but if I were I would bet that Knols get ranked right at the top, next to Youtube. As John Andrews describes it:

As TrustRank (the Google version, not the Yahoo! version) takes hold as the #1 or #2 ranking factor for SEO, this Knol thing steps in and bingo… who could be more trusted than Google itself?

Wikipedia has amazing momentum in Google, and is poised to rank for everything. How will Google compete?

How can Google come late to the game, offer no pay, desire to throw their ads on it right out of the gate, and expect to win marketshare UNLESS they rank this content better than it deserves to rank on merit? Put another way, what person who gets paid to create content is going to prefer putting it on Google Knol for free UNLESS Google gives Knol preferential treatment? If you are producing content out of passion with no profit motive, why would you put it on Google instead of your own server? If you desire peer review with your name attached to it why not publish it on YourName.com?

Offline media has always been biased and aggressively consolidated, it looks like the web is going to suffer the same fate, but worse, unless you are a Google stakeholder. Or, if Google gets too aggressive with this cross integration maybe they will hurt their relevancy enough that people search elsewhere.

Google, Subdomains, and Branding

In the past any large company could use subdomains as an effective reputation management strategy. As eBay and others have aggressively used subdomains to dominate branded AND unbranded search results, and Google has improved their sitelinks technology, any relevancy gain by treating subdomains as a separate site has gone away. Google is going to start treating subdomains like subfolders, and limit the number of results from any site to two.

There is still an upside to using subdomains because they allow you to feature standout content, but that upside relates to how marketable the content on that subdomain is, whereas in the past using lots of subdomains allowed eBay to get 20 of the top 30 listings for some queries, even if the subdomain was recycled garbage.

This move adds value to regionalizing sites and creating niche brands (like MobileCrunch), since currently I believe ebay.ca and ebay.com will be seen as two separate sites. If sites are too aggressive with regionalization or creating niche brands and start double dipping that way then Google might eventually look to devalue that move as well, although that will be more of a challenge because it would create a lot of collateral damage.

Official announcement by Matt Cutts at Pubcon, reported first by Barry.

By Far, The Worst Gmail Ad I've Ever Seen

Post by Giovanna Wall

I really like Google's Gmail program. It's truly my favorite email service. They also do a great job scanning through my emails for relevant keywords and phrases that they match with their advertisers. However, recently they have fouled and have gone out of bounds. In my recent emails to friends and family, I've used the terms "wife", "husband" and "happily married" a lot referring to my recent marriage.

This was Google's response:

google promtes infidelity

Perhaps I'm a little sensitive or maybe it's because I was raised as a conservative Catholic. But regardless of anyone's background, why would Google, with their "Do No Evil" policy promote cheating and infidelity? It's also ironic that the Google founders recently got married (I think one will wed next month).

It is an issue of money vs. morality when exposing disturbing ads to married people for ad revenue.

Google P2P Network? (or, Its Easy to Score Relevancy When You OWN the Network)

Google, already has a near infinite number of data points to compute relevancy for the active parts of the web, and is looking to gather even more user data information. The WSJ has background on the story:

Google is preparing a service that would let users store on its computers essentially all of the files they might keep on their personal-computer hard drives -- such as word-processing documents, digital music, video clips and images, say people familiar with the matter. The service could let users access their files via the Internet from different computers and mobile devices when they sign on with a password, and share them online with friends.

They also mentioned the C word:

Google will likely have to address copyright issues. Allowing consumers to share different types of files such as music with other users could trigger the sort of copyright complaints the company already faces over videos on its YouTube video sharing site. One person familiar with the matter says Google is discussing with copyright holders how to approach the issue and has some preliminary solutions.

This is going to move Google up the value system by

  • giving them a unique data source
  • giving them unique relevancy signals
  • keeping users locked into their services and using their services longer
  • shift power from copyright holders to Google
  • eventually allow Google to sell content (if they want to - the Google Video trial did not work too well)

But there will also be a big upside, especially to marketers and content creators who are willing to give away high value content to gain mindshare and marketshare. By creating content that people would want to store and share on Google, you get cheap or free exposure for your business interests.

As DaveN said, Google eventually has to move away from links because links are too polluted. What better relevancy signals can they come up with than attention data and how often people cite and share data ON THEIR NETWORK? Feedburner, Google Reader, iGoogle, Gmail, and Youtube are already part of the Google network. Soon your hard drive will be too.

Understanding Google's Mindset on Classifying Spam

If...

  • people would not notice it when Google removes your site from the search results
  • Google can clone your business model without paying writers to produce content or carrying physical inventory

... then your site is spam. Maybe not by today's standards, but eventually.

As the web evolves, a once whitelisted site can become a site that is easy to penalize. Evolve with the web, or grow irrelevant by the day.

This could sound like a scaremongering post, or it could be taken as a sign of the importance of connecting with people on an emotional level, and offering an experience worth sharing.

The Great Google Data Grab of 2007

If your Google AdWords quality score is too low, Google will allow you to compete in the auction with reasonable ad pricing ONLY if you give them your conversion data:

The arguement from my representative was that your pages are terrible so if we can’t see how well you convert our users then we will need you to pay $10.00 per click to make up for your low QS (my average keyword price was $2.75 at the time). That of course would have put me out of business.

Once I caved and allowed them to snoop on my conversions they allowed me to keep buying at or near my original keyword price.

If your copyright content is being uploaded to YouTube, Google will protect you if you upload your copyright content to Google:

I see a monetization in the works.

a) All of the big companies will make the effort to supply Youtube with good qualities of their videos. Movies, Shows etc.
b) YouTube gathers all that stuff, and builds the largest database of top quality videos.
c) Youtube offers the media companies to enter into a partnership. “Hey guys, you already have the stuff uploaded…why not sell the premium content to our millions of users?

If people are scraping and stealing your content Google will eventually allow you to rank for your own work if you sign up with Google Webmaster Central and register your copyright work.

After all your sites are registered with Google, how easy will it be for them to force compliance on smaller webmasters? Given the indiscriminate attitude exhibited when Google recently hand edited PageRank scores, it seems there is good reason to not register with the borg.

[Video] Using Google Date Based Filters

Tips on How to Use Google Indexing Date Filters

  • The Google advanced search page allows you to search for pages that were recently indexed, letting you filter through days, weeks, months, and years. Here are pages from SeoBook.com indexed in the last week.
  • In the URL they place as_qdr=w as_qdr=d (day) as_qdr=w (week) as_qdr=m (month) as_qdr=y (year). You can also search for multiples of these units, like search for pages indexed in the last 2 weeks by placing as_qdr=w2 in the URL string.
  • If you change your content management system or add new sections to your site you can see how quickly they are getting indexed, and look for any duplicate content issues as the new pages are getting indexed by looking for pages indexed under multiple URLs.
  • If you have never checked your site for duplicate content issues, but recently published content, that might also show any content duplication issues or Google indexed pages that you do not want in Google's index.
  • In addition to using date based filters to find how well your site is getting indexed, you can search to see who is mentioning an idea with a footprint, or use date based filters for doing link research.

Aesthetic Google PageRank Update in Google Toolbars Worldwide!!!

Search Engine Land recently listed a bunch of sites that had their PageRank scores manually edited for selling links. Of course, if you are the publisher of one of these sites, you don't care about an algorithm relevancy score so meaningless that it is edited by hand. You care about traffic.

Rankings Never Changed

SERoundtable was on the list of sites which saw their toolbar PageRank scores reduced, but I just looked at some of the terms they were ranking for, and they are still right at or near the top of the results for everything they were ranking for, even the competitive terms.

Traffic Matters, PageRank Does Not

Watch the Compete.com traffic stats for the sites listed in that Search Engine Land post. Google even provides large sites like the NYT over 20% of their traffic, so if these sites were really penalized you will see a plunge in traffic. If you do not see a plunge in traffic across these sites then toolbar PageRank scores are proved irrelevant as a measure of quality and trust.

Since When Are Publishing Networks Bad?

Many network blogs had their PageRank scores dropped too. Again, you can simply check the traffic stats to see if there is any real impact, or if Google is just polluting their Toolbar PageRank scores.

Since Google is demoting PageRank's viability as a site's global authority score perhaps this is a time for Yahoo to bring back WebRank, or Ask to launch something like CommunityRank. The Google-webmaster relationship is fraying. This presents an opportunity for whoever wants to take it.

Publishing is About Networks

If Google is penalizing blogs for being part of a network of sites, how long until they penalize IAC for owning 20 travel sites? Or Monster.com for owning 100 thin lead generation education sites? Or BankRate for owning white label sites with similar names? Networks have always been a part of publishing based business models.

It seems someone or something inside of Google is melting down. Choppy times ahead for webmasters worldwide.

Google Corrects Domain Name Spelling Errors (Sometimes, Anyway)

SEL highlighted that Google is correcting domain spelling errors. Which works to block some typos, but in some instances is pushing traffic away from smaller domains toward more authoritative websites.

Good Job Google

Here is an example of the spell correction working right...

Lets say you want to go to my blog located at www.seobook.com/node, but misspelled node as nodd. When you search for www.seobook.com/nodd they offer the correct URL as a suggestion.

Bad Job Google

Now lets say that I misspell a filename. What if I typed www.seibook.com/bok (If you add a second o to the word book in the filename this URL exists). What does Google do? Even when I am not signed in, Google STILL recommends people go to SeoBook.com, to the URL www.seobook.com/blog instead of recommending they go to seibook.com/book/

In that last case correcting the URL and keeping the people on the same site only took changing 1 letter, but Google decided instead to change a letter in the domain name, and change 3 in the filename!

Why Not Fix This?

What about errors in the domain extension? If you type in ASP.nt (leaving out the e in net) Google does not correct that spelling error. If you type in ebay.cm (ebay.com leaving out the o) Google does not correct that error. Why launch a feature such as this without correcting the most common errors?

Pages