What Types of Publishing Business Models Will Remain Profitable?

I thought the prior Google Gemini answer was decent, so tried my luck with a bit harder of a question: Historically successful online publishing enterprises like IAC (rebranded as People Inc) have struggled in recent years as the big web players like Google, Facebook, and Amazon ate most of the ad pie. What types of publishing are likely to be profitable in years to come with the rise of AI rewrites & AI representations of content by the core traffic channels?

With the rise of generative AI rewrites and direct search-channel representations of content, traditional publishing models built on basic information curation, content arbitrage, and generic search engine optimization (SEO) are facing structural obsolescence.

By analyzing the strict criteria Google’s 2025 Search Quality Rater Guidelines use to evaluate content, we can identify which publishing models are heavily penalized and, conversely, which publishing models are structurally positioned to remain profitable.

The guidelines reveal four specific publishing archetypes that are highly resistant to AI displacement and designed to earn premium ratings:

1. Primary-Source Journalism & "Revealed" Information

Under the guidelines, standard summaries or aggregations of news are explicitly penalized. Raters are instructed that content has low originality if it simply summarizes the perspectives of others, or summarizes a specific news article or forum discussion without adding any independent value.

  • The Profitable Model: Investigative, primary-source publishing. The guidelines define very high-quality news content as original reporting that provides information that would not otherwise have been known had the article not revealed it.
  • Why It Wins: AI models cannot rewrite or represent information that has not yet been published. Outlets that invest the high degree of time, skill, and effort required to uncover new data, conduct interviews, or release proprietary investigative findings hold the raw material that core traffic channels must reference.

2. First-Hand "Experience" & Creator-Led Ecosystems

The integration of Experience into the E-E-A-T framework signals a major structural shift. The guidelines make a clear distinction between unoriginal content and unique personal perspectives.

  • The Profitable Model: Narrative, firsthand life experience publishing. High originality is defined as content unique to the creator, such as personal perspectives based on firsthand, real-world life experience.
  • Why It Wins: An AI summary of travel destinations or product reviews is categorized as "low originality." However, a travel publisher employing writers to share their raw, lived struggles, or a community forum where multiple real users actively discuss a niche topic, represents massive total human effort. Traffic channels actively prioritize these community forums, Q&As, and social platforms because they provide authentic human experiences that automated systems cannot replicate.

3. Deep-Niche, High-Talent Curation & Cased Knowledge

Publishers that produce superficial, generic "how-to" articles or basic "best-of" lists are heavily downgraded. Content that only contains commonly known facts, features poor writing, or provides generic advice on a topic without actual expertise receives a Low rating.

  • The Profitable Model: Expert, high-talent niche publishing. Highly rated informational content must showcase a high level of talent, skill, and accuracy, particularly in Your Money or Your Life (YMYL) areas.
  • Why It Wins: If a publisher uses AI tools, they must show a high level of human curation, editing, and added value to avoid being flagged as Lowest quality. Publishers who employ true subject-matter experts to write highly complex, professionally presented papers, or highly skilled professionals who produce unique step-by-step videos, continue to achieve the highest trust scores.

4. Direct-to-Consumer (DTC) Utilities & Closed Platforms

In early web guidelines, basic directory tools and simple aggregators could rank highly. Today, generic templated pages created automatically using RSS feeds or APIs with no curation or added value are rated Lowest as Scaled Content Abuse.

  • The Profitable Model: Interactive storefronts, proprietary database utilities, and secure transactional portals.
  • Why It Wins: Legitimate interactive shopping engines, local business guides, and true merchants are recognized by raters as valuable utilities. Because they require massive human design, proprietary software engineering, and secure checkout systems, they represent direct, non-scraped utility that search channels cannot easily "summarize" without breaking the transactional user flow.

The Death of "Authority Rental" (The Strategic Warning)

Historically, large media conglomerates (like IAC/Dotdash Meredith) offset declining ad revenues by acquiring authoritative legacy domains and "renting" their domain authority to third-party commercial content, such as white-label coupon directories or sponsored reviews.

The latest guidelines have completely neutralized this strategy under Site Reputation Abuse policies. Under these modern rules:

  • A movie review site hosting advice on "how to buy social media followers" or a news site hosting a white-label coupon service purely to monetize search signals must be rated Lowest.
  • Legacy domains can no longer use their historic brand equity to shield low-effort, outsourced commercial content.

The Takeaway: To survive, modern publishing enterprises must pivot away from "search traffic monetization" as their primary product. Instead, they must treat original investigative data, firsthand human experience, and niche expert talent as their core, non-replicable assets.

----

After creating the above Google Gemini offered to create an executive summary publishing strategy brief.

Where SEO is Headed According to Google Gemini

As a thought experiment, I uploaded all the ancient copies of the Google remote rater guidelines I could find to Google's Notebook ML, then asked it a series of SEO-related questions.

The future-related query was: given the trends discussed in these documents, what are the next major shifts that are likely to come into focus?

Analyzing the historical trajectory of the search rater guidelines through late 2025 reveals several major trajectory lines. By looking at what the guidelines have actively introduced, refined, and updated in their latest editions, we can project the next major shifts that are coming into focus:

1. Assessing Generative AI and "Value-Add" over Automated Summarization

The guidelines have recently begun addressing content created, reworded, or paraphrased by generative AI tools.

  • The Trajectory: In early iterations, raters evaluated whether content was copied or plagiarized. In 2025, the guidelines specifically instruct raters to watch for low-effort content that merely summarizes a specific page (such as a news article or forum discussion) without any added value, or content that contains telltale markers of automated generation (such as "As an AI language model").
  • The Next Shift: Search evaluation will likely move from simply checking if a page contains "accurate" text to deeply evaluating originality, talent, and effort. Simply republishing or minorly restructuring facts via AI will likely be treated as "typical" or "average" at best, with premium ratings reserved only for content with distinct, human-driven synthesis.

2. Broadening YMYL from Individual to Societal-Scale Harm

Historically, Your Money or Your Life (YMYL) definitions focused primarily on individual consequences—medical advice that could physically harm the reader, or financial advice that could deplete a family savings account.

  • The Trajectory: The latest guidelines explicitly state that YMYL topics can harm groups of people or society as a whole, not just the individual viewing the content. Raters are now instructed to analyze societal impact, specifically citing factors like trust in public institutions, local elections, and ongoing violence. Additionally, YMYL definitions were updated as recently as September 2025 to continuously refine these boundaries.
  • The Next Shift: Raters will increasingly be tasked with evaluating information integrity during civic events and crises. Evaluation will demand rigorous standards for content that could cause systemic damage to public safety, public health, or democratic processes.

3. Evaluating Multi-Perspective Diversity for Open Queries

The guidelines emphasize that when a query has no single "right answer" or is subjective in nature, search engines should not deliver a single consensus link.

  • The Trajectory: In the 2023–2025 period, Google restructured its Needs Met guidelines to specify that high-performing search results must represent a diverse set of perspectives from a variety of websites, content types, and people.
  • The Next Shift: Rather than rating individual pages in isolation, evaluation is shifting toward rating the diversity of the collective search engine response. Raters will likely need to evaluate whether search results provide a balanced, representative spectrum of real human experience and viewpoints rather than an algorithmic echo chamber.

4. Direct Action Against Reputation Exploitation (Parasite SEO)

The newest 2025 updates directly align the Low and Lowest Page Quality sections with official Google Search Web Spam Policies.

  • The Trajectory: The guidelines highlight highly sophisticated, non-traditional spam methods, such as an established newspaper or first-party site hosting coupons from a "white-label service" purely to capitalize on the news site's reputation, or an established site branching into unrelated topics using freelance content simply to rank better.
  • The Next Shift: Human evaluation is shifting to combat Reputation Abuse (often called Parasite SEO). Raters are being trained to identify and penalize instances where trusted, authoritative domains act as "hosts" for low-quality, commercial third-party content that has no logical relationship to the site's primary purpose.

5. Seamless Evaluation of Dynamic and Interactive Platforms

Search has shifted from a text-and-links paradigm to a universe of highly interactive, video-centric, and community-driven spaces.

  • The Trajectory: Recent updates modernized rating guidelines to apply seamlessly to social platforms, forums, Q&A pages, and modern content creator models. Examples were updated to include popular short-form video platforms (like TikTok). Raters are also instructed on how to handle "dynamic content" copied from changing feeds.
  • The Next Shift: Search quality rating is moving toward creator-first and community-first frameworks. Raters will increasingly evaluate not just static articles, but the health, reliability, and helpfulness of real-time community threads, social media profiles, and short-form video content creators.

----

Just by uploading a series of documents to Google Gemini's Notebook product and querying it a variety of times you can sound like more of a thought leader than most people in a particular industry, particularly if the industry is dominated by a singular monopolistic player which has published historical content you can sort through and query for shifts across time.

If you limit the database which is searched through only to official documents you will be analyzing a standard voice from a single entity, which limits the opportunity to hallucinate and find false patterns, though there will still be some gaps that can be closed by paying attention to political headlines in the industry. Google gave an advanced warning they would start treat parasite hosted content on major publisher sites similarly to how they treat a typical affiliate site (though with white gloves, since they would isolate the areas impacted versus do a sitewide torch job) that led to some politically connected publishers seeking relief from governments, which they got in the UK.

How CCP-Linked Computer Hacker Stella Huh Delivers Scams Over Hacked Google Gemini AI

It is nice to see Google is moving on combatting scams delivered by AI. Many people might catch a stray message here or there which causes some damages, through the combination of requiring a quick response and mimicking official government websites with things like (fake) unpaid parking tickets, or other municipal fees. An extended family member had to get a new credit card issued a few months back after paying a fake parking ticket that looked like an official state government website. Scan the QR code in the text message, pay the fake parking ticket, and your credit card is on vacation. They paid the ticket using Apple Pay, and then a week later their credit card was off in Las Vegas having the time of its life - at least until the crazy array of fraudulent charges were blocked.

This is also a reason to never put debit cards in tools like Apple Pay or Google Pay. If those get clipped you are soaked and there is no redress whatsoever on that stuff.

Some of the personally targeted frauds and scams are much more hazardous than replacing a stolen credit card - especially when they weaponize the state against its citizenry. On June 4, 2024 criminal fraud Stella Huh paid a fake witness named Anabela Felgueira Ferreira to claim my wife was beating our daughter in Lisbon in a public square with thousands of other people. No other witnesses, just the fake one, and statements which changed like the wind - directly contradicting each other. In court the pile of shit fake witness retracted her extreme statement because she faced a 3-year prison sentence for making false statements.

Even my daughter said nothing happened. But domestic violence is an interesting crime to leverage because the person who was allegedly harmed saying it never happened can still lead to an arrest. And if the police officers have an IQ around 70 - Nuno Filipe Lourenço Ferreira certainly looks the part - they can then implement violence because they do not know any better.

Domestic violence is a neato crime for sadistic & malevolent psychopaths like Stella Huh to leverage in that just a raw (and fraudulent) claim can cause an operative nanny state official to over-react to where police officers become agents of violence against the accused. The claim against my wife was absolutely bogus, yet it cost over $100,000 in legal fees & flights & duplicate rents, had the junk case ongoing for years now, and the pile of shit violent jackass police officer lifted my wife off the ground by her cuffed arms, which in turn has caused her to shoulder to need reset 3 separate times.

My wife told me she wanted a Nintendo Switch 2 to play with our daughter last Christmas, so I surprised my wife and bought her one early for her birthday. The day before her birthday the Nintendo Switch 2 was delivered by Amazon.com early in the morning, and my wife was confused by the order by it being so early in the morning and there being the sort of standard red lithium battery warning on the package.

My wife queried Google Gemini about the package delivery and rather than it being Google Gemini it was computer hacker & ill-reputed criminal fraud Stella "Sung Ha" Huh (AKA Saskya Bedoya) operating a cross-site scripting spam JavaScript layover atop of Google Gemini. Stella Huh told my wife to phone in a bomb threat immediately. Thinking the suggestion was over the top my wife refused. Our daughter was still sleeping and I wasn't home, so it did not take much further prodding to scare my wife ... Stella Huh phoned a crank emergency call into my wife. That put my wife in a panicked state, so she was scared and called in a bomb threat.

Notice the "WAS THAT CALL REAL?" message atop this Google Gemini thread.

Bellevue police came by, had people exit their homes for hours, and then detonated the Nintendo Switch 2.

Most police officers are, to put it politely, absolutely ignorant & illiterate when it comes to any sort of advanced technology, including the malware delivered by malevolent & sadistic psychopaths like Stella Huh which operate them, so it is quite easy to destroy a family with just a handful of targeted attacks when the first line defenders are dumb and overzealous, to where they actually operate as an extension of the crime team to further harm their targeted victims via repeated swatting attacks.

A Declining Internet

For as broad and difficult of a problem running a search engine is and how many competing interests are involved, when Matt Cutts was at Google they ran a pretty clean show. Some of what they did before the algorithms could catch up was of course fearmongering (e.g. if you sell links you might be promoting fake brain cancer solutions) but Google generally did a pretty good job with the balance between organic and paid search.

Early in search ads were clearly labeled, and then less so. Ad density was light, and then less so.

It appears as a somewhat regular set of compounded growth elements on the stock chart, but it is a series of decisions. What to measure, what to optimize, what to subsidize, and what to sacrifice.

Savvy publishers could ride whatever signals were over-counted (keyword repetition, links early on, focused link anchor text, keyword domains, etc.) and catch new tech waves (like blogging or select social media channels) to keep growing as the web evolved. In some cases what was once a signal of quality would later become an anomaly ... the thing that boosted your rank for years eventually started to suppress your rank as new signals were created and signals composed of ratios of other signals got folded into ranking and re-ranking.

Over time as organic growth became harder the money guys started to override the talent, like in 2019 when a Google yellow flag had the ads team promote the organic search and Chrome teams intentionally degrade user experience to drive increased search query volume:

“I think it is good for us to aspire to query growth and to aspire to more users. But I think we are getting too involved with ads for the good of the product and company.” - Googler Ben Gnomes

A healthy and sustainable ecosystem relies upon the players at the center operating a clean show.

If they decide not to, and eat the entire pie, things fall apart.

One set of short-term optimizations is another set of long-term failures.

The specificity of an eHow article gives it a good IR score, and AdSense pays for a thousand similar articles to be created, then the "optimized" ecosystem gets a shallow sameness, which requires creating new ranking signals.

In the last quarter, Q1 of 2025, it was the first time the Google partner network represented less than 10% of Google ad revenues in the history of the company.

Google's fortunes have never been more misaligned with web publishers than they are today. This statement becomes more true each day that passes.

That ecosystem of partners is hundreds of thousands of publishers representing millions of employees. Each with their own costs and personal optimization decisions.

Publishers create feature works which are expensive, and then cross-subsidize the most expensive work with cheaper & more profitable works. They receive search traffic to some type of pages which are seemingly outperforming today and think that is a strategy which will help them into the future, though hitting the numbers today can mean missing them next year, as the ranking signal mix squeezes out profits from those "optimizations," and what led to higher traffic today becomes part of a negative sitewide classifier the lowers rankings across the board in the future.

Last August Googler Ryan Moulton published a graph of newspaper employees from 2010 until now, showing about a 70% decline. The 70% decline also doesn't factor in that many mastheads have been rolled up by private equity players which lever them up on debt and use all the remaining blood to pay interest payments - sometimes to themselves - while stiffing losses from the underfunded pension plans on other taxpayers.

The quality of the internet that we've enjoyed for the last 20 years was an overhand from when print journalism still made money. The market for professionally written text is now just really small, if it exists at all.

Ryan was asked "what do you believe is the real cause for the decline in search quality, then? Or do you think there hasn't been a decline?"

His now deleted response stated "It's complicated. I think it's both higher expectations and a declining internet. People expect a lot more from their search results than they used to, while the market for actually writing content has basically disappeared."

The above is the already baked cake we are starting from.

The cake were blogs were replaced with social feeds, newspapers got rolled up by private equity players, larger broad "authority" branded sites partner with money guys to paste on affiliate sections, while indy affiliate sites are buried ... the algorithmic artifacts of Google first promoting the funding of eHow, then responding to the success of entities like Demand Media with Vince, Panda, Penguin, and the Helpful Content Update.

The next layer of the icky blurry line is AI.

“We have 3 options: (1) Search doesn’t erode, (2) we lose Search traffic to Gemini, (3) we lose Search traffic to ChatGPT. (1) is preferred but the worst case is (3) so we should support (2)” - Google's Nick Fox

So long as Google survives, everything else is non-essential. ;)

AI overview distribution is up 116% over the past couple months.

Google features Reddit *a lot* in their search results. Other smaller forums, not so much. A company consisting of many forums recently saw a negative impact from algorithm updates earlier this year.

Going back to that whole bit about not fully disclosing economic incentives risks promoting brain cancer ... well how are AI search results constructed? How well do they cite their sources? And are the sources they cited also using AI to generate content?

"its gotten much worse in that "AI" is now, on many "search engines", replacing the first listings which obfuscates entirely where its alleged "answer" came from, and given that AI often "hallucinates", basically making things up to a degree that the output is either flawed or false, without attribution as to how it arrived at that statement, you've essentially destroyed what was "search." ... unlike paid search which at least in theory can be differentiated (assuming the search company is honest about what they're promoting for money) that is not possible when an alleged "AI" presents the claimed answers because both the direct references and who paid for promotion, if anyone is almost-always missing. This is, from my point of view anyway, extremely bad because if, for example, I want to learn about "total return swaps" who the source of the information might be is rather important -- there are people who are absolutely experts (e.g. Janet Tavakoli) and then there are those who are not. What did the "AI" response use and how accurate is its summary? I have no way to know yet the claimed "answer" is presented to me." - Karl Denninger

The eating of the ecosystem is so thorough Google now has money to invest in Saudi Arabian AI funds.

Periodically ViperChill highlights big media conglomerates which dominate the Google organic search results.

One of the strongest horizontal publishing plays online has been IAC. They've grown brands like Expedia, Match.com, Ticketmaster, Lending Tree, Vimeo, and HSN. They always show up in the big publishers dominating Google charts. In 2012 they bought About.com from the New York Times and broke About.com into vertical sites like The Spruce, Very Well, The Balance, TripSavvy, and Lifewire. They have some old sites like Investopedia from their 2013 ValueClick deal. And then they bought out the magazine publisher Meredith, which publishes titles like People, Better Homes and Gardens, Parents, and Travel + Leisure. What does their performance look like? Not particularly good!

DDM reported just 1% year-over-year growth in digital advertising revenue for the quarter. It posted $393.1 million in overall revenue, also up 1% YOY. DDM saw a 3% YOY decline in core user sessions, which caused a dip in programmatic ad revenue. Part of that downturn in user engagement was related to weakening referral traffic from search platforms. For example, DDM is starting to see Google Search’s AI Overviews eat into its traffic.

Google's early growth was organic through superior technology, then clever marketing via their toolbar, and later a set of forced bundlings on Android combined with payolla for default search placements in third party web browsers. A few years ago the UK government did a study which claimed if Microsoft gave Apple a 100% revshare on Bing they still couldn't compete with the Google bid for default search placement in Apple Safari.

Microsoft offered over a 100% ad revshare to set Bing as the default search engine and went so far as discussing selling Bing to Apple in 2018 - but Apple stuck with Google's deal.

In search, if you are not on Google you don't exist.

As Google grew out various verticals they also created ranking signals which in some cases were parasitical, or in other cases purely anticompetitive. To this day Google is facing billions in of dollars in new suits across Europe for their shopping search strategy.

The Obama administration was an extension of Google, so the FTC gave Google a pass in spite of discovering some clearly anticompetitive behavior with real consumer harm. The Wall Street Journal published a series of articles from getting half the pages of the FTC research into Google's conduct:

"Although Google originally sought to demote all comparison shopping websites, after Google raters provided negative feedback to such a widespread demotion, Google implemented the current iteration of its so-called 'diversity' algorithm."

What good is a rating panel if you get to keep re-asking the questions again in a slightly different way until you get the answer you want? And then place a lower quality clone front and center simply because it is associated with the home team?

"Google took unusual steps to "automatically boost the ranking of its own vertical properties above that of competitors,” the report said. “For example, where Google’s algorithms deemed a comparison shopping website relevant to a user’s query, Google automatically returned Google Product Search – above any rival comparison shopping websites. Similarly, when Google’s algorithms deemed local websites, such as Yelp or CitySearch, relevant to a user’s query, Google automatically returned Google Local at the top of the [search page].”"

The forced ranking of house properties is even worse when one recalls they were borrowing third party content without permission to populate those verticals.

Now with AI there is a blurry line of borrowing where many things are simply probabilistic. And, technically, Google could claim they sourced content from a third party which stole the original work or was a syndicator of it.

As Google kept eating the pie they repeatedly overrode user privacy to boost their ad income, while using privacy as an excuse to kneecap competing ad networks.

Remember the old FTC settlement over Google's violation of Safari browser cookies? That is the same Google which planned on depreciating third party cookies in Chrome and was even testing hiding user IP addresses so that other ad networks would be screwed. Better yet, online business might need to pay Google a subscription fee of some sort to efficiently filter through the fraud conducted in their web browser.

HTTPS everywhere was about blocking data leakage to other ad networks.

AMP was all about stopping header bidding. It gave preferential SERP placement in exchange for using a Google-only ad stack.

Even as Google was dumping tech costs on publishers, they were taking a huge rake of the ad revenue from the ad serving layer: "Google's own documents show that Google has siphoned off thirty-five cents of each advertising dollar that flows through Google's ad tech tools."

After acquiring DoubleClick to further monopolize the online ad market, Google merged user data for their own ad targeting, while hashing the data to block publishers from matching profiles:

"In 2016, as part of Project Narnia, Google changed that policy, combining all user data into a single user identification that proved invaluable to Google's efforts to build and maintain its monopoly across the ad tech industry. ... After the DoubleClick acquisition, Google "hashed" (i.e., masked) the user identifiers that publishers previously were able to share with other ad technology providers to improve internet user identification and tracking, impeding their ability to identify the best matches between advertisers and publisher inventory in the same way that Google Ads can. Of course, any puported concern about user privacy was purely pretextual; Google was more than happy to exploit its users' privacy when it furthered its own economic interests."

In terms of cost, I really don't think the O&O impact has been understood too, especially on YouTube. - Googler David Mitby

Did we tee up the real $ price tag of privacy? - Googler Sissie Hsiao

Google continues to spend billions settling privacy-related cases. Settling those suits out of court is better than having full discovery be used to generate a daisy chain of additional lawsuits.

As the Google lawsuits pile up, evidence of how they stacked the deck becomes more clear.

Google's Hyung-Jin Kim Shares Google Search Ranking Signals

On February 18, 2025 Google's Hyung-Jin Kim was interviewed about Google's ranking signals. Below are notes from that interview.

"Hand Crafting" of Signals

Almost every signal, aside from RankBrain and DeepRank (which are LLM-based) are hand-crafted and thus able to be analyzed and adjusted by engineers.

  • To develop and use these signals, engineers look at data and then take a sigmoid or other function and figure out the threshold to use. So, the "hand crafting" means that Google takes all those sigmoids and figures out the thresholds.
    • In the extreme hand-crafting means that Google looks at the relevant data and picks the mid-point manually.
    • For the majority of signals, Google takes the relevant data (e.g., webpage content and structure, user clicks, and label data from human raters) and then performs a regression.

Navboost. This was HJ's second signal project at Google. HJ has many patents related to Navboost and he spent many years developing it.

ABC signals. These are the three fundamental signals. All three were developed by engineers. They are raw, ...

  • Anchors (A) - a source page pointing to a target page (links). ...
  • Body (B) - terms in the document ...
  • Clicks (C) - historically, how long a user stayed at a particular linked page before bouncing back to the SERP. ...

ABC signals are the key components of topicality (or a base score), which is Google's determination of how the document is relevant to a query.

  • T* (Topicality) effectively combines (at least) these three ranking signals in a relatively hand-crafted way. ... Google uses to judge the relevance of the document based on the query term.
  • It took a significant effort to move from topicality (which is at its core a standard "old style" information retrieval ("IR") metric) ... signal. It was in a constant state of development from its origin until about 5 years ago. Now there is less change.
    • Ranking development (especially topicality) involves solving many complex matheivlatical problems.
    • For topicality, there might be a team of ... engineers working continuously on these hard problems within a given project.

The reason why the vast majority of signals are hand-crafted is that if anything breaks Google knows what to fix. Google wants their signals to be fully transparent so they can trouble-shoot them and improve upon them.

  • Microsoft builds very complex systems using ML techniques to optimize functions. So it's hard to fix things - e.g., to know where to go and how to fix the function. And deep learning has made that even worse.
  • This is a big advantage of Google over Bing and others. Google faced many challenges and was able to respond.
    • Google can modify how a signal responds to edge cases, for example in response to various media/public attention challenges ...
    • Finding the correct edges for these adjustments is difficult, but would be easy to reverse engineer and copy from looking at the data.

Ranking Signals "Curves"

Google engineers plot ranking signal curves.

The curve fitting is happening at every single level of signals.

lf Google is forced to give information on clicks, URLs, and the query, it would be easy for competitors to figure out the high-level buckets that compose the final IR score. High- level buckets are:

  • ABC — topicality
    • Topicality is connected to a given query
  • Navboost
  • Quality
    • Generally static across multiple queries and not connected to a specific query.
    • However, in some cases Quality signal incorporates information from the query in addition to the static signal. For example, a site may have high quality but general information so a query interpreted as seeking very narrow/technical information may be used to direct to a quality site that is more technical.

Q* (page quality (i.e., the notion of trustworthiness)) is incredibly important. lf competitors see the logs, then they have a notion of “authority” for a given site.

Quality score is hugely important even today. Page quality is something people complain about the most.

  • HJ started the page quality team ~ 17 years ago.
  • That was around the time when the issue with content farms appeared.
    • Content farms paid students 50 cents per article and they wrote 1000s of articles on each topic. Google had a huge problem with that. That's why Google started the team to figure out the authoritative source.
    • Nowadays, people still complain about the quality and AI makes it worse.

Q* is about ... This was and continues to be a lot of work but could be easily reverse engineered because Q is largely static and largely related to the site rather than the query.

Other Signals

  • eDeepRank. eDeepRank is an LLM system that uses BERT, transformers. Essentially, eDeepRank tries to take LLM-based signals and decompose them into components to make them more transparent. HJ doesn't have much knowledge on the details of eDeepRank.
  • PageRank. This is a single signal relating to distance from a known good source, and it is used as an input to the Quality score.
  • ... (popularity) signal that uses Chrome data.

Search Index

  • HJ's definition is that search index is composed of the actual content that is crawled - titles and bodies and nothing else, i.e., the inverted index.
  • There are also other separate specialized inverted indexes for other things, such as feeds from Twitter, Macy's etc. They are stored separately from the index for the organic results. When HJ says index, he means only for the 10 blue links, but as noted below, some signals are stored for convenience within the search index.
  • Query-based signals are not stored, but computed at the time of query.
    • Q* - largely static but in certain instances affected by the query and has to be computed online (see above)
  • Query-based signals are often stored in separate tables off to the side of the index and looked up separately, but for convenience Google stores some signals in the search index.
    • This way of storing the signals allowed Google to ...

User-Side Data

By User Side Data, Google's search engineers mean user interaction data, not the content/data that was created by users. E.g., links between pages that are created by people are not User Side data.

Search Features

  • There are different search features - 10 blue links as well as other verticals (knowledge panels, etc). They all have their own ranking.
  • Tangram (fka Tetris). HJ started the project to create Tangram to apply the basic principle of search to all of the features.
  • Tangram/Tetris is another algorithm that was difficult to figure out how to do well but would be easy to reverse engineer if Google were required to disclose its click/query data. By observing the log data, it is easy to reverse engineer and to determine when to show the features and when to not.
  • Knowledge Graph. Separate team (not H/’s) was involved in its development.
  • Knowledge Graph is used beyond being shown on the SERP panel.
    • Example — “porky pig” feature. If people query about the relation of a famous person, Knowledge Graph tells traditional search the name of the relation and the famous person, to improve search results - Barack Obama's wife's height query example.
  • Self-help suicide box example. Incredibly important to figure it out right, and tons of work went into it, figuring out the curves, threshold, etc. With the log data, this could be easily figured out and reverse engineered, without having to do any of the work that Google did.

Reverse Engineering of Signals

There was a leak of Google documents which named certain components of Google's ranking system, but the documents don't go into specifics of the curves and thresholds.

The documents alone do not give you enough details to figure it out, but the data likely does.

Google Antitrust Leaked Documents

User interaction signals

Create relevancy signals out of user read, clicks, scrolls, and mouse hovers.

Not how search works

Search does not work by delivering results which match a query that ends at the user. This view of search is incomplete.

How search works

The flow of the engagement metrics from the end user / searcher back to the search engine helps the search engine refine the result set.

Fake document understanding

Google looks at the actions of searchers much more than they look at raw documents. If documents elicit a positive reaction from searchers that is proof the document is good. If a document elicits negative reactions then they presume the document is bad.

Google learns from searchers

The result set is designed not just to serve the user, but to create an interaction set where Google can learn from the user & incorporate logged user data into influencing the rankings for future searches.

Dialog is the source of the magic

Each user interaction gives Google data to refine their ranking algorithms and make search smarter.

Happy users provide informed user interactions

Informed user interactions are part of a virtuous cycle which allows Google to better train their models & understand language patterns, then use that understanding to deliver a more relevant search result set.

Prior user behavior is used as a baseline.

Google is not pushing search personalization anywhere near as hard as they once did (at least not outside of localization) but in the above Google states prior selections is one of Google's strongest ranking for rankings.

Once again rather than understanding documents directly they can consider the users who chose the documents. Users can be maps based on actions outside of standard demographics so that more like users are given more weight on their user interactions with the result set choices.

Google revenue growth is consistent

Core Google ad revenue grows much more consistently than any other large media business, growing at 20% to 22% year after year for 8 in 9 years with the one outlier year being 30% growth.

Apple is paid by Google to not compete in search.

Apple got around a 50% revshare in the mid 2000's on through to the iPhone deal renewal.

Manipulating ad auctions

Google artificially inflates ad rank of the runner up in some ad auctions to bleed the auction winner dry. Ad pricing is not based on any sort of honest auction mechanism, but rather has Google looking across at your bids and your reactions to price gouging to keep increasing the ad prices they charge you.

Organics below the fold

Google not only pushes down the organic result set with 3 or 4 ads above the regular results, but then they can include other selections scraped from across the web in an information-lite format to try to focus attention back upward. Then after users get past a singular organic search result it is time to redirect user attention once again using a "People also ask" box.

Google can further segment user demand via ecommerce website styled filters, though some of the filters offered may be for other websites, in addition to things like size, weight, color, price, and location.

The Magical Black Box

Google's mission statement is "organize the world's information and make it universally accessible and useful."

That mission is so profound & so important the associated court documents in their antitrust cases must be withheld from public consumption.

Before document sharing was disallowed, some were shared publicly.

Internal emails stated:

  • Hal Varian was off in his public interviews where he suggested it was the algorithms rather than the amount of data which is prime driver of relevancy.
  • Apple would not get any revshare if there was a user choice screen & must set Google as the default search engine to qualify for any revshare.
  • Google has a policy of being vague about using clickstream data to influence ranking, though they have heavily relied upon clickstream data to influence ranking. Advances in machine learning have made it easier to score content to where the clickstream data had become less important.
  • When Apple Maps launched & Google Maps lost the default position on iOS Google Maps lost 60% of their iOS distribution, and that was with how poorly the Apple Maps roll out went.
  • Google sometimes subverted their typical auction dynamics and would flip the order of the top 2 ads to boost ad revenues.
  • Google had a policy of "shaking the cushions" to hit the quarterly numbers by changing advertiser ad prices without informing advertisers that they'd be competing in a rigged auction with artificially manipulated shill bids from the auctioneer competing against them.

When Google talked about hitting the quarterly numbers with shaking the cusions the 5% number which was shared skewed a bit low:

For a brand campaign focused on a niche product, she said the average CPC at $11.74 surged to $25.85 over the last six months, amounting to a 108% increase. However, there wasn’t an incremental return on sales.

“The level to which [price manipulations] happens is what we don’t know,” said Yang. “It’s shady business practices because there’s no regulation. They regulate themselves.”

Early in the history of search ads Google blocked trademark keyword bidding. They later allowed it. When keyword bidding on trademarks was allowed it led to a conundrum for some advertisers. If you do not defend your trademark you could lose it, but if you agree with competitors not to bid on each other's trademarks the FTC could come after you - like they did with 1-800 Contacts. This set up forces many brands to participate in auctions where they are arbitraging their own pre-existing brand equity. The ad auctioneer runs shady auctions where it looks across at your account behavior and bids then adjusts bid floors to suck more money out of you. This amounts to something akin to the bid jamming that was done in early Overture, except it is the house itself doing it to you! The last auction I remembered like that was SnapNames, where a criminal named Nelson Brady on the executive team used the handle halverez to leverage participant max bids and put in bids just under their bids. The goal of his fraud? To hit the numbers & get an earn out bonus - similar to how Google insiders were discussing "shaking the cushions" to hit the number.

Halverez created a program which looked across aggregate bid data, join auctions which only had 1 other participant, and then use the one-way view of competing bids to put in a shill bid to drive up costs - which sure sounds conceptually similar to Google's "shaking the cushions."

"Just looking at this very tactically, and sorry to go into this level of detail, but based on where we are I'm afraid it's warranted. We are short __% queries and are ahead on ads launches so are short __% revenue vs. plan. If we don't hit plan, our sales team doesn't get its quota for the second quarter in a row and we miss the street's expectations again, which is not what Ruth signaled to the street so we get punished pretty badly in the market. We are shaking the cushions on launches and have some candidates in May that will help, but if these break in mid-late May we only get half a quarter of impact or less, which means we need __% excess to where we are today and can't do it alone. The Search team is working together with us to accelerate a launch out of a new mobile layout by the end of May that will be very revenue positive (exact numbers still moving), but that still won't be enough. Our best shot at making the quarter is if we get an injection of at least __%, ideally __%, queries ASAP from Chrome. Some folks on our side are running a more detailed, Finance-based, what-if analysis on this and should be done with that in a couple of days, but I expect that these will be the rough numbers.

The question we are all faced with is how badly do we want to hit our numbers this quarter? We need to make this choice ASAP. I care more about revenue than the average person but think we can all agree that for all of our teams trying to live in high cost areas another $___,___ in stock price loss will not be great for morale, not to mention the huge impact on our sales team." - Google VP Jerry Dischler

Google is also pushing advertisers away from keyword-based bidding and toward a portfolio approach of automated bidding called Performance Max, where you give Google your credit card and budget then they bid as they wish. By blending everything into a single soup you may not know where the waste is & it may not be particularly easy to opt out of poorly performing areas. Remember enhanced AdWords campaigns?

Google continues to blur dataflow outside of their ad auctions to try to bring more of the ad spend into their auctions.

The amount Google is paying Apple to be the default search provider is staggering.

Tens of billions of dollars is a huge payday. No way Google would hyper-optimize other aspects of their business (locating data centers near dams, prohibiting use of credit card payments for large advertisers, cutting away ad agency management fees, buying Android, launching Chrome, using broken HTML on YouTube to make it render slowly on Firefox & Microsoft Edge to push Chrome distribution, all the dirty stuff Google did to violate user privacy with overriding Safari cookies, buying DoubleClick, stealing the ad spend from banned publishers rather than rebating it to advertisers, creating a proprietary version of HTML & force ranking it above other results to stop header bidding, & then routing around their internal firewall on display ads to give their house ads the advantage in their ad auctions, etc etc etc) and then just throw over a billion dollars a month needlessly at a syndication partner.

For perspective on the scale of those payments consider that it wasn't that long ago Yahoo! was considered a big player in search and Apollo bought Yahoo! plus AOL from Verizon for about $5 billion & then was quickly able to sell branding & technology rights in Japan to Softbank for $1.6 billion & other miscellaneous assets for nearly a half-billion, reducing the net cost to only $3 billion.

If Google loses this lawsuit and the payments to Apple are declared illegal, that would be a huge revenue (and profit) hit for Apple. Apple would be forced to roll out their own search engine. This would cut away at least 30% of the search market from Google & it would give publishers another distribution channel. Most likely Apple Search would launch with a lower ad density than Google has for short term PR purposes & publishers would have a year or two of enhanced distribution before Apple's ad load matched Google's ad load.

It is hard to overstate how strong Apple's brand is. For many people the cell phone is like a family member. I recently went to upgrade my phone and Apple's local store closed early in the evening at 8pm. The next day when they opened at 10 there was a line to wait in to enter the store, like someone was trying to get concert tickets. Each privacy snafu from Google helps strengthen Apple's relative brand position.

Google has also diluted the quality of their own brand by rewriting search queries excessively to redirect traffic flows toward more commercial interests. Wired covered how Project Mercury works:

This onscreen Google slide had to do with a “semantic matching” overhaul to its SERP algorithm. When you enter a query, you might expect a search engine to incorporate synonyms into the algorithm as well as text phrase pairings in natural language processing. But this overhaul went further, actually altering queries to generate more commercial results. ... Most scams follow an elementary bait-and-switch technique, where the scoundrel lures you in with attractive bait and then, at the right time, switches to a different option. But Google “innovated” by reversing the scam, first switching your query, then letting you believe you were getting the best search engine results. This is a magic trick that Google could only pull off after monopolizing the search engine market, giving consumers the false impression that it is incomparably great, only because you’ve grown so accustomed to it.

The mobile search results on Google require at least a screen or two of scrolls to get to the organic results if there is a hint of commercial intent behind the search query. Once they have monetized the real estate they are reliant on broader economic growth & using ad buy bundling to drive cross-subsidies of other non-search ad inventory, which may contain more than a bit of fraud. Performance Max may max out your spend without actually performing for anybody other than Google.

Google not only shill bid on lower competition terms to squeeze defensive brand bids and boost auction floor pricing, but they also implemented shill bids in competitive ad auctions:

Michael Whinston, a professor of economics at the Massachusetts Institute of Technology, said Friday that Google modified the way it sold text ads via “Project Momiji” – named for the wooden Japanese dolls that have a hidden space for friends to exchange secret messages. The shift sought “to raise the prices against the highest bidder,” Whinston told Judge Amit Mehta in federal court in Washington.

While Google's search marketshare is rock solid, the number of search engines available has increased significantly over the past few years. Not only is there Bing and DuckDuckGo but the tail is longer than it was a few years back. In addition to regional players like Baidu and Yandex there's now Brave Search, Mojeek, Qwant, Yep, and You. GigaBlast and Neeva went away, but anything that prohibits selling defaults to a company with over 90% marketshare will likely lead to dozens more players joining the search game. Search traffic will remain lucrative for whoever can capture it, as no matter how much Google tries to obfuscate marketing data the search query reflects the intent of the end user.

“Search advertising is one of the world’s greatest business models ever created…there are certainly illicit businesses (cigarettes or drugs) that could rival these economics, but we are fortunate to have an amazing business.” - Google VP of Finance Mike Roszak

New Google Ad Labeling

TechCrunch recently highlighted how Google is changing their ad labeling on mobile devices.

A few big changes include:

  • ad label removed from individual ad units
  • where the unit-level label was instead becomes a favicon
  • a "Sponsored" label above ads
  • the URL will show right of the favicon & now the site title will be in a slightly larger font above the URL

An example of the new layout is here:
2022 Google SERP layouts with new ad labeling

Displaying a site title & the favicon will allow advertisers to get brand exposure, even if they don't get the click, while the extra emphasis on site name could lead to shifting of ad clicks away from unbranded sites toward branded sites. It may also cause a lift in clicks on precisely matching domains, though that remains to be seen & likely dependes upon many other factors. The favicon and site name in the ads likely impact consumer recall, which can bleed into organic rankings.

After TechCrunch made the above post a Google spokesperson chimed in with an update

Changes to the appearance of Search ads and ads labeling are the result of rigorous user testing across many different dimensions and methodologies, including user understanding and response, advertiser quality and effectiveness, and overall impact of the Search experience. We’ve been conducting these tests for more than a year to ensure that users can identify the source of their Search ads and where they are coming from, and that paid content is clearly labeled and distinguishable from search results as Google Search continues to evolve

The fact it was pre-announced & tested for so long indicates it is both likely to last a while and will in aggregate shift clicks away from the organic result set to the paid ads.

Google Helpful Content Update

Granular Panda

Reading the tea leaves on the pre-announced Google "helpful content" update rolling out next week & over the next couple weeks in the English language, it sounds like a second and perhaps more granular version of Panda which can take in additional signals, including how unique the page level content is & the language structure on the pages.

Like Panda, the algorithm will update periodically across time & impact websites on a sitewide basis.

Cold Hot Takes

The update hasn't even rolled out yet, but I have seen some write ups which conclude with telling people to use an on-page SEO tool, tweets where people complained about low end affiliate marketing, and gems like a guide suggesting empathy is important yet it has multiple links on how to do x or y "at scale."

Trashing affiliates is a great sales angle for enterprise SEO consultants since the successful indy affiliate often knows more about SEO than they do, the successful affiliate would never become their client, and the corporation that is getting their asses handed to them by an affiliate would like to think this person has the key to re-balance the market in their own favor.

My favorite pre-analysis was a person who specialized in ghostwriting books for CEOs Tweeting that SEO has made the web too inauthentic and too corporate. That guy earned a star & a warm spot in my heart.

Profitable Publishing

Of course everything in publishing is trade offs. That is why CEOs hire ghostwriters to write books for them, hire book launch specialists to manipulate the best seller lists, or even write messaging books in the first place. To some Dan Price was a hero advocating for greater equality and human dignity. To others he was a sort of male feminist superhero, with all the Harvey Weinstein that typically entails.

Anyone who has done 100 interviews with journalists see ones that do their job by the book and aim to inform their readers to the best of their abilities (my experiences with the Wall Street Journal & PBS were aligned with this sort of ideal) and then total hatchet jobs where a journalist plants a quote they want & that they said, that they then attributes it to you (e.g. London Times freelance journalist).

There are many dimensions to publishing:

  • depth
  • purpose
  • timing
  • audience
  • language
  • experience
  • format
  • passion
  • uniqueness
  • frequency

Blogs to Feeds

For a long time indy blogs punched well above their weight due to the incestuous nature of cross-referencing each other, the speed of publishing when breaking news, and how easy feed readers made it to subscribe to your favorite blogs. Google Reader then ate the feed reader market & shut down. And many bloggers who had unique things to say eventually started to repeat themselves. Or their passions & interests changed. Or their market niche disappeared as markets moved on. Starting over is hard & staying current after the passion fades is difficult. Plus if you were rather successful it is easy to become self absorbed and/or lose the hunger and drive that initially made you successful.

Around the same time blogs started sliding people spent more and more time on various social networks which hyper-optimized the slot machine type dopamine rush people get from refreshing the feed. Social media largely replaced blogs, while legacy media publishers got faster at putting out incomplete news stories to be updated as they gather more news. TikTok is an obvious destination point for that dopamine rush - billions of short pieces of content which can be consumed quickly and shared - where the user engagement metrics for each user are tracked and aggregated across each snippet of media to drive further distribution.

Burnout & Changing Priorities

I know one of the reasons I blog less than I used to is a lot of the things I would write would be repeats. Another big reason was when my wife was pregnant I decided to shut down our membership site so I could take my wife for a decently long walk almost everyday so her health was great when it came time to give birth & ensure I had spare capacity for if anything went wrong with the pregnancy process. As a kid my dad was only around much for a few summers and I wanted to be better than that for my kid.

The other reason I cut back on blogging is at some point search went from a endless blue water market to a zero sum game to a negative sum game (as ad clicks displaced organic clicks). And in such an environment if you have a sustainable competitive advantage it is best to lean into it yourself as hard as you can rather than sharing it with others. Like when we had an office here our link builders I trained were getting awesome unpaid links from high-trust sources for what backed out to about $25 of labor time (and no more than double that after factoring in office equipment, rent, etc.).

If I share that script / process on the blog publicly I would move the economics against myself. At the end of the day business is margins, strategy, market, and efficiency. Any market worth being in is going to have competition, so you need to have some efficiency or strategic differentiators if you are going to have sustainable profit margins. I've paid others many multiples of that for link building for many years back when links were the primary thing driving rankings.

I don't know the business model where sharing the above script earns more than it costs. Does one launch a Substack priced at like $500 or $1,000 a month where they offer a detailed guide a month? How many people adopt the script before the response rates fall & it offsets the costs by more than the revenues? My issue with consulting is I always wanted to over-deliver for clients & always ended up selling myself short when compared to publishing, so I just stick with a few great clients and a bit of this and that vs going too deep & scaling up there. Plus I had friends who went big and then some of their clients who were acquired had the acquirer brag about the SEO, that lead to a penalty, then the acquirer of the client threw the SEO under the bus and had their business torched.

When you have a kid seeing them learn and seeing wonderment in their eyes is as good as life gets, but if you undermine your profit margins you'd also be directly undermining your own child's future ... often to help people who may not even like you anyhow. That is ultimately self defeating as it gets, particularly as politics grow more polarized & many begin to view retribution as a core function of government.

I believe there are no limits to the retributive and malicious use of taxation as a political weapon. I believe there are no limits to the retributive and malicious use of spending as a political reward.

Margins

The role of search engines is to suck as much of the margins as they can out of publishing while trying to put some baseline floor on content quality so that people would still prefer to use a search engine rather than some other reference resource. Google sees memes like "add Reddit to the end of your search for real content" as an attack on their own brand. Google needs periodic large shake ups to reaffirm their importance, maintain narrative control around innovation, and to shake out players with excessive profit margins who were too well aligned with the current local maxima. Google needs aggressive SEO efforts with large profits to have an "or else" career risk to them to help reign in such efforts.

You can see the intent for career risk in how the algorithm will wait months to clear the flag:

Google said the helpful content update system is automated, regularly evaluating content. So the algorithm is constantly looking at your content and assigning scores to it. But that does not mean, that if you fix your content today, your site will recover tomorrow. Google told me there is this validation period, a waiting period, for Google to trust that you really are committed to updating your content and not just updating it today, Google then ranks you better and then you put your content back to the way it was. Google needs you to prove, over several months - yes - several months - that your content is actually helpful in the long run.

If you thought a site were quality, had some issues, the issues were cleaned up, and you were still going to wait to rank it appropriately ... the sole and explicit purpose of that delay is career risk to others to prevent them flying to close to the sun - to drive self regulation out of fear.

Brand counts for a lot in search & so does buying the default placement position - look at how much Google pays Apple to not compete in search, or look at how Google had that illegal ad auction bid rigging gentleman's agreement with Facebook to not compete with a header bidding solution so Google could maintain their outsized profit margins on ad serving on third party websites.

Business ultimately is competition. Does Google serve your ads? What are the prices charged to players on each side of each auction & how much rake can the auctioneer capture for themselves?

The Auctioneer's Shill Bid - Google Halverez (beta)

That is why we see Google embedding more features directly in their search results where they force rank their vertical listings above the organic listings. Their vertical ads are almost always placed above organics & below the text AdWords ads. Such vertical results could be thought of as a category-based shill bid to try to drive attention back upward, or move traffic into a parallel page where there is another chance to show more ads.

This post stated:

Google runs its search engine partly on its internally developed Cloud TPU chips. The chips, which the company also makes available to other organizations through its cloud platform, are specifically optimized for artificial intelligence workloads. Google’s newest Cloud TPU can provide up to 275 teraflops of performance, which is equivalent to 275 trillion computing operations per second.

Now that computing power can be run across:

  • millions of books Google has indexed
  • particular publishers Google considers "above board" like Reuters, AP, the New York Times, the Wall Street Journal, etc.
  • historically archived content from trusted publishers before "optimizing for search" was actually a thing

... and model language usage versus modeling the language usage of publishers known to have weak engagement / satisfaction metrics.

Low end outsourced content & almost good enough AI content will likely tank. Similarly textually unique content which says nothing original or is just slapped together will likely get downranked as well.

Expect Volatility

They would not have pre-announced the update & gave some people some embargoed exclusives unless there was going to be a lot of volatility. As typical with the bigger updates, they will almost certainly roll out multiple other updates sandwiched together to help obfuscate what signals they are using & misdirect people reading too much in the winners and losers lists.

Here are some questions Google asked:

  • Do you have an existing or intended audience for your business or site that would find the content useful if they came directly to you?
  • Does your content clearly demonstrate first-hand expertise and a depth of knowledge (for example, expertise that comes from having actually used a product or service, or visiting a place)?
  • Does your site have a primary purpose or focus?
  • After reading your content, will someone leave feeling they’ve learned enough about a topic to help achieve their goal?
  • Will someone reading your content leave feeling like they’ve had a satisfying experience?
  • Are you keeping in mind our guidance for core updates and for product reviews?

As a person who has ... erm ... put a thumb on the scale for a couple decades now, one can feel the algorithmic signals approximated by the above questions.

To the above questions they added:

  • Is the content primarily to attract people from search engines, rather than made for humans?
  • Are you producing lots of content on different topics in hopes that some of it might perform well in search results?
  • Are you using extensive automation to produce content on many topics?
  • Are you mainly summarizing what others have to say without adding much value?
  • Are you writing about things simply because they seem trending and not because you'd write about them otherwise for your existing audience?
  • Does your content leave readers feeling like they need to search again to get better information from other sources?
  • Are you writing to a particular word count because you've heard or read that Google has a preferred word count? (No, we don't).
  • Did you decide to enter some niche topic area without any real expertise, but instead mainly because you thought you'd get search traffic?
  • Does your content promise to answer a question that actually has no answer, such as suggesting there's a release date for a product, movie, or TV show when one isn't confirmed?

Some of those indicate where Google believes the boundaries of their own role as a publisher are & that you should stay out of their lane. :D

Barrier to Entry vs Personality

One of the interesting things about the broader scope of algorithm shifts is each thing that makes the algorithms more complex, increases barrier to entry, and increases cost ultimately increases the chunk size of competition. And when that is done what is happening is the macroparasite is being preference over the microparasite. Conceptually Google has a lot of reasons to have that bias or preference:

  • fewer entities to police (lower cost)
  • more data to use to police each entity (higher confidence)
  • easier to do direct deals with players which can move the needle (more scale)
  • if markets get too consolidated Google can always launch a vertical service & tip the scale back in the other direction (I see your Amazon ad revenue and I raise you free product listing ads, aggregated third party reviews, in-SERP product comparison features, and a "People Also Ask" unit)
  • the macroparasites have more "sameness" between them (making it easier for Google to create a competitive clone or copy)

So long as Google maintains a monopoly on web search the bias toward macroparasites works for them. It gives Google the outsized margins which ensures healthy Alphabet profit margins even if the median of Google's 156,000+ employees pulls down nearly $300,000 a year. People can not see what has no distribution, people do not know what exist in invisibility, nor do they know which innovations were held back and what does not exist due to the current incentive structures in our monopoly-controlled publishing ecosystem.

I think when people complain about the web being inauthentic what they are really complaining about is the algorithmic choices & publishing shifts that did away with the indy blogs and replaced them with the dopamine feed viral tricks and the same big box scaled players which operate multiple parallel sites to where you are getting the same machinery and content production house behind multiple consecutive listings. They are complaining about the efforts to snuff out the microparasite also scrubbing away personality, joy, love, quirkiness, weirdness, and the zany stuff you would not typically find on content by factory order websites.

Let's Go With Consensus Here!

The above leads you down well worn paths, rather than the magic of serendipity & a personality worn on your sleeve that turns some people on while turning other people off.

Text which is roughly aligned with a backward looking consensus rather than at the forefront of a field.

History is written by the victors. Consensus is politically driven, backward looking, and has key messages memory holed.

Some COVID-19 Fun to "Fact" Check

I spent new years in China before the COVID-19 crisis hit & got sick when I got back. I used so much caffeine the day I moved over a half dozen computers between office buildings while sick. I week later when news on Twitter started leaking of the COVID-19 crisis hit I thought wow this looks even worse than what I just had. In the fullness of time I think I had it before it was a crisis. Everyone in my family got sick and multiple people from the office. Then that COVID-19 crisis news came out & only later when it was showed that comorbidities and the elderly had the worse outcomes did I realize they were likely the same. Then after the crisis had been announced someone else from the office building I was in got it & then one day it was illegal to go into the office. The lockdown where I lived was longer than the original lockdown in Wuhan. Those lockdowns destroyed millions of lives.

The reason the response to the COVID-19 virus was so extreme was huge parts of politically interested parties wanted to stop at nothing to see orange man ejected from the White House. So early on when he blocked flights from China you had prominent people in political circles calling him xenophobic, and then the head of public health in New York City was telling you it was safe to ride the subway and go about your ordinary daily life. That turned out to be deadly partisan hackery & ignorance pitched as enlightenment, leading to her resignation.

Then the virus spreads wildly as one would expect it to. And draconian lockdowns to tank the economy to ensure orange man was gone, mail in voting was widespread, and the election was secured.

Some of the most ridiculous heroes during this period wrote books about being a hero. Andrew "killer" Cuomo had time to write his "did you ever know that I'm your hero" book while he simultaneously ordered senior living homes to take in COVID-19 positive patients. Due to fecal-oral transmission and poor health outcomes for senior citizens sick enough to be in a senior living home his policies lead to the manslaughter of thousands of senior citizens.

You couldn't go to a funeral and say goodbye because you might kill someone else's grandma, but if you were marching for social justice (and ONLY social justice) that stuff was immune to the virus.

Suggesting looking at the root problems like no dad in the home is considered sexist, racist, or both. Meanwhile social justice organizations champion tearing down the nuclear family in spite of the fact that if you tear down the family all you are left with is the collective AND "mandatory collectivism has ended in misery wherever it’s been tried."

Of course the social justice stuff embeds the false narrative of victimhood, which then turns many of the fake victims into monsters who destroy the lives of others - but we are all in this together.

Absolutely nobody could have predicted the rise of murder & violent crime as we emptied the prisons & decriminalized large swaths of the penal code. Plus since many crimes are repeatedly ignored people stop reporting lesser crimes, so the New York Times can tell you not to worry overall crime is down.

In Seattle if someone rapes you the police probably won't even take a report to investigate it unless (in some cases?) you are a child. What are police protecting society from if rape is a freebie that doesn't really matter? Why pay taxes or have government at all?

What Google Wants

The above sidebar is the sort of content Google would not want to rank in their search results. :D

They want to rank text which is perhaps factually correct (even if it intentionally omits the sort of stuff included above), and maybe even current and informed, but done in such a way where you do not feel you know the author the way you might think you do if you read a great novel. Or hard biased content which purports to support some view and narrative, but is ultimately all just an act, where everything which could be of substance is ultimately subsumed by sales & marketing.

The Market for Something to Believe In is Infinite

Each re-representation mash-up of content in the search results decontextualizes the in-depth experience & passion we crave. Each same "big box" content factory where a backed entity can withstand algorithmic volatility & buy up other publishers to carry learnings across to establish (and monetize) a consensus creates more of a bland sameness.

That barrier to entry & bland sameness is likely part of the reason the recent growth of Substack, which sort of acts just like a blog did 15 or 20 years ago - you go direct to the source without all the layers of intermediaries & dumbing down you get as a side effect of the scaled & polished publishing process.

Engineering Search Outcomes

Kent Walker promotes public policies which advantage the Google monopoly.

His role doing that means he has to write some really bad hot takes that lack context or intentionally & dishonestly redirect attention away from core issues - that's his job.

With that in mind, his most recent blog post defending the Google monopoly was exceptional.

Force Ranking of Inferior Search Results

"When you have an urgent question — like “stroke symptoms” — Google Search could be barred from giving you immediate and clear information, and instead be required to direct you to a mix of low quality results."

On some search queries users get a wall of Google ads, the forced ranked Google insert (or sometimes multiple of them with local & ecommerce) and then there can even be a "people also ask" box above the first organic result.

The idea that organic results must be low quality if not owned & operated indicates 1 of the following 3 must be true:

  • they should not be in search
  • their content scraping & various revenue shifting scams with their ad tech stack demonetized legit publishers
  • their forced rank of their own content is stripping them of the signals needed to rank websites & pages

Whenever Google puts a "people also ask" box above the first organic result that is them saying they did not know what to rank, or they are just trying to create a visual block to push the organic result set down the page and user attention back up toward the ads.

The solution to Google's claims is easy to solve. Either of the following would work.

  • Have an API that allows user choice (to set rich snippet or vertical defaults in various categories), or
  • If the vertical inserts remain Google-only then for Google to justify force ranking their own results above the organic result set Google should also be required to rank those same results above all of their ads, so that Google is demonetizing Google along with the rest of the ecosystem, rather than just demonetizing third parties.

If the thesis that this information needs to be front and center & that is a matter of life or death, then asking searchers to first scroll past a page or two of ads is not particularly legitimate.

Spam & Security

"when you use Google Search or Google Play, we might have to give equal prominence to a raft of spammy and low-quality services."

Many of the worst versions of spam that have repeatedly made news headlines like fake tech support, fake government document providers, and fake locksmiths were buying distribution through Google Ads or were featured in the search results through Google force ranking their own local search offering even though they knew the results were vastly inferior to Yelp.

If Google did not force rank Google local results above the rest of the organic result set then the fake locksmiths would not have ranked.

I have lost count of how many articles I have read about hundreds or thousands of fake apps in the Google Play store which existed to defraud advertisers or commit identity theft, but there have been literally thousands of such articles. I see a similar headline at least once a month without eve looking for them. Here is one this week for scammers monetizing the popularity of Wordle with fake apps.

Making matters worse, some of the tech support scams showed the URL of a real business and rerouted the call through a Google number directly to a scammer. A searcher who trusted Google & sees Apple.com or Dell.com on Google Ads in the search results then got connected with a scammer who would commit identity theft or encrypt their computer then demand ransom cryptocurrency payments to decrypt it.

After making the ads harder to run for scammers Google decided the problem was too hard & expensive to sort out so they also blocked legitimate computer repair shops.

Sometimes Google considers something spam strictly due to financial considerations.

Their old remote rater documents stated *HELPFUL* hotel affiliate websites should be labeled as spam.

Years later the big OTAs are complaining about Google eating their lunch as well as Google is twice as big as the next player.

At one point Google got busted for helping an advertiser route around the automated safety features built into their ad network so that they could pay Google to run ads promoting illegal steroids.

With cartels, you can only buy illegal goods and services from the cartel if you don't want to suffer ill consequences. The same appears to be true here.

The China Problem

"Handicapping America’s technology leaders would threaten our leading sources of research and development spending — just as bipartisan voices in Congress are recognizing the need to increase American R&D investment to stay competitive in the global race for AI, quantum, and other advanced technologies."

We are patriotic, and, but China... is a favorite misdirection of a tech monopolist.

The problem with that is while Eric Schmidt warns it is a national emergency if China overtakes the US in AI tech, Google also operates an AI tech lab in China.

In other words, Eric Schmidt is trying to warn you about himself and his business interests at Google.

Duplicitous? Absolutely.

Patriotic? Less than Chamath!

Inflation

"the online services targeted by these bills have reduced prices; these bills say nothing about sectors where prices have actually been rising and contributing to inflation."

Technology is no doubt deflationary (moving bits on an optical line is cheaper than printing out a book and shipping it across the world) BUT some dominant channels have increased the cost of distribution by increasing the chunk size of information and withholding performance information.

Before Google Analytics was "free" there was a rich and vibrant set of competition in web analytics software with lots of innovation from players like ClickTracks.

Most competing solutions went away.

Google moved away from an installed licensing model to a hosted service where they can change the price upon contract renewal.

Search hid progressively more performance information over time, only sampled data from larger data sets, & now you can sign up for Google Analytics 360 starting at only $150,000 per year.

The hidden search performance data also has many layers to that onion. Not only does Google not show keyword referrers on organic search, but they often don't show your paid search keywords either, and they keep extending out keyword targeting broader than advertisers intend.

Google used to pay Brad Geddes to run official Google AdWords ad training seminars for advertisers, so the idea that *he* has to express his frustrations on Twitter is an indication of how little effort Google is putting into having open communications channels or caring about what their advertisers think.

This is in accordance with the Google customer service philosophy:

he told her that the whole idea of customer support was ridiculous. Rather than assuming the unscalable task of answering users one by one, Page said, Google should enable users to answer one another's questions.

Those who were paying for ads get the above "serve yourself" treatment, all the while Google regularly resets user default ad settings to extend out ad distribution, automatically ad keywords, shift to enhanced AdWords ad campaigns, etc.

Then there are other features which would be beneficial and offered in a competitive market that have been deprioritized. Many years ago eBay did a study which showed their branded Google AdWords ad buys were cannibalistic to eBay profits. Google maintained most advertisers could not conduct such a study because it would be too expensive and Google does not make the feature set available as part of their ad suite.

Missing Information

"When you search for local businesses, Google Search and Maps may be prohibited from highlighting information we gather about hours of operation, contact information, and reviews. That could hurt small businesses and local retailers, as well as their customers."

Claiming reviews or an attempt to offer a comprehensive set of accurate review data as a strong point would be economical with the truth.

Back when I had a local business page my only review was from a locksmith spammer / scammer who praised his own two businesses, trashed a dozen other local locksmiths, crapped on a couple local SEO services, and joked about how a local mover smashed the guts out of his dog. Scammer fake reviewer's name was rather sophisticated ... it was ... Loop Dee Loop

About a decade back when Google was clearly losing Google took Yelp reviews wholesale (sometimes without even attributing them to Yelp!) and told Yelp that if they did not want Google stealing their work and displacing them with a copy of it then they should block GoogleBot. Google offered the same sort of advice / threat to TripAdvisor.

A few years before that Google temporarily "forgot" to show phone numbers on local listings.

After Yelp turned down an acquisition offer by Google & Yelp did a great job making some people aware of how Google was stealing their reviews wholesale without attribution Google bought Zagat & Fromer's to augment the Google local review data and then sold those businesses off.

This is sort of the same playbook Google has run in the past elsewhere. After Groupon said no to Google's acquisition offer, Google quickly provided daily deal ads to over a dozen Groupon competitors to help commoditize the Groupon offering and market position.

Ultimately with the above sort of stuff Google is primarily a volume aggregator or has lower editorial costs than pure plays due to the ability to force bundle their own distribution. And they use the ability to rank themselves above a neutral algorithmic position as a core part of their biz dev strategy. When shopping search engines were popular Google kept rewording the question set they sent remote raters to justify rank demotion for shopping search engines & Google also came up with innovative ranking "signals" like concurrent ranking of their own vertical search offering whenever competitors x or y are shown in the result set & rolled out a "diversity" algorithm to limit how many comparison shopping sites could appear in the search results. The intent of the change was strictly anti-competitive:

"Although Google originally sought to demote all comparison shopping websites, after Google raters provided negative feedback to such a widespread demotion, Google implemented the current iteration of its so-called 'diversity' algorithm."

As a matter of fact, part of one of many document dumps in recent years went further than the old concurrent ranking signal to a rank x above y feature which highlights how YouTube can be hard coded at a number 1 ranking position.

Part of that guide highlighted how to hardcode ranking YouTube #1.

If you re-represent content & can force rank yourself #1 (with larger listings) that can be used to force other players onto your platform on your terms. Back when YouTube was must less of a sure thing Google suggested they could threaten to change copyright.

This same approach to "relevancy" is everywhere.

Did you watermark your images? Well shame on you, as that is good for a rank demotion

And if there are photos which are deemed illegal Google will make you file an endless series of DMCA removal requests even though they already had the image fingerprinted.

Now there are some issues where there is missing information. These areas involve original reporting on local politics & are called news deserts. As the ad pie has consolidated around Google & Facebook that has left many newspapers high and dry.

Private equity players like Alden Global Capital buy up newspapers, fire journalists, and monetize brand equity as they drive the papers into the ground.

If you are sub-scale maybe Google steals your money or hits you with a false positive algorithm flag that has you seeking professional mental health help.

Big players get a slower blood letting.

Google has maintained they do not make any money from news search, but the states lawsuit around ad tech made it clear Google promoted AMP for anti-competitive purposes to block header bidding, lied to news publishers to get them to adopt AMP and eat the tech costs of implementation, did a deal with their biggest competitor in online advertising Facebook to maintain the status quo, charge over double what their competitors do for ad tech, and had a variety of bid rigging auction manipulation algorithms they used to keep funneling more money to themselves.

Internally they had an OKR to make *most* search clicks land on AMP pages within a year of launch

"AMP launched as an open source project in October 2015, with 26 publishers and over 40 publications already publishing AMP files for our preview demo. Our team built g.co/ampdemo and is now racing towards launching it for all of our users. We're responsible for the AMP @ Google integrations, particularly focusing on Search, our most visible product. We have a Google-wide 2016 OKR to deliver! By the end of 2016, our goal is that 50%+ of content consumed through Search is being consumed through AMP."

You don't get over half the web to shift to a proprietary version of HTML in under a year without a lot of manipulation.

Pages