---
title: "AI Content Risks"
description: "Why reckless AI content production harms SEO: Google's spam policy, hallucination, E-E-A-T erosion, and why chasing citations with AI-generated volume tends to backfire."
tldr: "AI-generated content is not inherently penalised by Google, but the patterns that come with reckless production, thin coverage, factual errors, absent authorship and no original expertise, are what Google's spam systems demote, and its AI features run on the same systems. Mass-produced AI content aimed at ranking or citation volume is a named spam policy violation. A compounding risk sits underneath: AI content becomes the source other AI tools retrieve from, so an unverified claim is repeated until repetition passes for corroboration, and models invent citation URLs that resolve to nothing. The defences: check what a model is sourcing, and open every URL before citing it."
publishDate: 2026-05-05
updatedDate: 2026-09-18
author: "Liam Hayward"
---
AI tools have made content production faster and cheaper. Used well, they can help with research, drafting, and editing. Used recklessly (as a pipeline for publishing at scale without genuine editorial input), they introduce risks that tend to compound over time: algorithmic demotion, loss of trust signals, and reduced visibility in the AI-generated answers they were often deployed to target.

## What does Google's spam policy say?

Google's position is frequently misquoted. The policy is not that AI-generated content is prohibited. It is that content produced primarily to manipulate search rankings, regardless of how it was created, violates the spam policy. Google has said so in its own words since February 2023: "Using automation, including AI, to generate content with the primary purpose of manipulating ranking in search results is a violation of our spam policies", alongside the principle that its focus is "on the quality of content, rather than how content is produced".[^6]

The relevant category is "scaled content abuse", which Google defines as generating many pages "for the primary purpose of manipulating search rankings and not helping users", typically large amounts of unoriginal content "no matter how it's created"; its first named example is using generative AI tools to produce many pages without adding value.[^7] That is: generating large volumes of content, with or without AI, where the primary purpose is ranking rather than serving readers. A site publishing 300 AI-generated articles on loosely related topics, with thin coverage and no meaningful editorial review, falls under this policy whether or not the content is technically accurate.

The distinction that matters is intent and quality, not the tool. AI-assisted content, where a subject matter expert drafts with AI help, reviews the output, and publishes under their name, sits in a different category entirely from an automated pipeline publishing content at scale.

## Hallucination and factual accuracy

AI language models generate text that is confident, well-formatted, and sometimes wrong. This is the hallucination problem: the model produces a plausible-sounding claim that is factually incorrect, and the error is often invisible to anyone without domain knowledge.

For general topics, a hallucinated sentence might be embarrassing. For specialist content - SEO, health, finance, legal - it can be a material trust liability. A page stating an incorrect algorithm date, a misattributed statistic, or a tool feature that no longer exists damages credibility in ways that are hard to recover from, particularly on a site making authority claims.

The risk is not that AI writes poorly. It is that AI writes poorly in a way that passes a surface-level read. Catching hallucinations requires someone who knows enough to know what to check, which is the domain expertise that a purely automated content pipeline typically removes from the process.

## Outdated training data

AI models have knowledge cutoffs. They write about the world as it was when their training data was collected, not as it is now. For stable topics this is a minor issue. For anything that changes regularly - Google features, algorithm behaviour, tool interfaces, crawler support - the gap between training data and current reality can be significant.

SEO content is particularly exposed. AI Overviews roll out and change behaviour month by month. Google Search Console adds and removes reports. Third-party tools update their interfaces. An AI-drafted article about any of these topics, published without review by someone tracking current developments, may be factually incorrect from day one and progressively more so as time passes.

The fix is a review layer with genuine current knowledge. Proofreading is insufficient; the reviewer must understand the domain well enough to catch errors the writer didn't see.

## What happens when AI content becomes AI's source?

Hallucination is usually described as one model inventing one fact. The compounding problem is what happens next, when that output becomes a source other models retrieve.

An AI-drafted article publishes an unverified claim. The article gets indexed. The next model retrieves it alongside others repeating the same claim, all tracing back to the same unchecked origin, and reports it as settled fact rather than a contested assertion, because from the inside it looks like consensus. Ten sources saying the same thing are worth no more than one if all ten copied that one, and a search summary across them cannot tell the difference.

Both sides of that loop have been measured, with the same instrument and the same limit. Ahrefs sampled 900,000 newly published pages in April 2025 and found 74.2% contained some AI-generated content, though only 2.5% were written entirely by AI and most were a human-machine blend, so "three-quarters of the web is AI-written" misreads the finding.[^1] Graphite estimated that 42.7% of the references ChatGPT cited in June 2026 were themselves AI-generated, up from 38.9% five months earlier.[^2] Both figures come from AI-content detectors, which infer from style and misclassify individual pages, so the direction is credible and the precise shares are not. Provider-side marking is the stronger signal and is arriving under the EU AI Act, but it changes nothing you can check today; [AI watermarking](/ai-search/ai-watermarking/) covers it.

Graphite also simulated where the loop ends. Once references a model had itself authored entered the retrieval pool, answers collapsed onto the same homogenised result in 79.6% of runs, driven by self-authorship rather than AI generation as such.[^2] The authors state they "do not definitively prove that AI search collapse is already happening": a demonstrated mechanism, not a measured state of the live web. [The collapse study](/seo-news-updates/ai-search-collapse-graphite-study/) has the detail; the loop's other effect, propagating errors, is the subject here.

### The 360Brew case

In 2025, a research paper described "360Brew", a large language model for ranking, written by authors from LinkedIn. The paper was later withdrawn, with arXiv removing every version because the submitter did not have the right to agree to the licence.[^3] LinkedIn's VP of Engineering then addressed the claim directly. Asked whether 360Brew was used in ranking, he answered: "Short answer: No!", explaining that the model had been tested with a small group of members, judged not the right fit, and the test shut down.[^4]

None of that stopped it. Marketing articles continued to announce 360Brew as LinkedIn's new algorithm, AI tools retrieving from those articles repeated it as fact, and practitioners built ranking advice on top of it.[^5] The primary source had been retracted and the platform had publicly denied the claim, yet the derivative layer registered neither, because nothing in that layer checks.

### Two failure modes that follow

**Retractions do not propagate.** When a paper is withdrawn or a company corrects the record, the correction lands on the original source. The articles, summaries and model outputs already built on it are not revisited, so the claim continues to circulate in the derivative layer after its basis has gone.

**Models fabricate citations.** Asked to source a claim, a language model will produce a plausible URL that has never existed, assembled from the pattern of real URLs on that domain. A fabricated citation is indistinguishable from a genuine one until it is opened, which is why opening it is the only check that catches it.

### How do you defend against it?

Three habits, all mechanical rather than attitudinal.

**Check what the model is sourcing, not only what it is saying.** A confident answer with no source, or with a source nobody has opened, is not evidence.

**Open every URL before citing it.** This catches fabricated references, which no other check will.

**Treat heavy repetition of a specific claim as a reason for more scrutiny, not less.** Widespread agreement about a named internal detail, an algorithm, a model, a system, that appears only in marketing content and never in a primary source, is a signal to trace the claim back to its origin and confirm the origin still stands.

## E-E-A-T erosion

Experience, expertise, authoritativeness, and trustworthiness are assessed at the page level and the site level. Content produced at scale without named authors, without first-hand experience, and without genuine expertise behind it is structurally weak on every E-E-A-T dimension.

A site that publishes 500 articles with no bylines has made a choice to trade long-term authority for short-term output. Google's quality systems are designed to identify this pattern, and its self-assessment questions for helpful content ask whether it is "self-evident to your visitors who authored your content".[^8] Google's AI features are "rooted in our core Search ranking and quality systems",[^9] so the same questions apply there. Beyond Google, no platform documents an authorship signal and no study has measured one; the argument for named authors on AI surfaces is that a page a reader can attribute is a page a reader can check.

## The core irony

The sites most commonly targeted by AI content strategies - AI Overviews, Perplexity, ChatGPT Search - retrieve and cite based on: factual accuracy, clear attribution, verifiable expertise, and well-sourced claims. These are the signals that reckless AI content production systematically fails to provide.

A site publishing AI-generated content at volume, without meaningful human review, is producing what Google's AI guide names directly: creating pages for every variation of a query "primarily to manipulate rankings or generative AI responses in Google Search violates Google's scaled content abuse spam policy", and "a high quantity of pages doesn't make a website higher quality or more relevant to users".[^9] The strategy sold as a way to win AI search visibility is a named policy violation on the largest AI surface.

This is not a coincidence. Google says its generative features are "rooted in our core Search ranking and quality systems",[^9] and the overlap is measured for the others too: Semrush found Google's AI Mode shared around 54% of its cited domains with the organic top 10 in July 2025, and Evertune's May 2026 review of the most-cited URLs across six models concluded that "pages that perform well in human search results also tend to perform well in bot-driven searches".[^10][^11] On Google, the content that earns citations is largely the content that would have ranked well anyway.

## What does responsible AI-assisted content look like?

The question is not whether to use AI tools. It is what the human layer looks like.

Content produced with AI assistance and published responsibly typically involves: a subject matter expert directing or reviewing the output, fact-checking against current sources rather than relying on the model's training data, a named author with genuine credentials, and editorial judgment about what to publish and what to cut.

The practical test: would a knowledgeable person in the relevant field put their name to this and stand behind it? If not, the content is not ready to publish, regardless of how it was produced.

AI tools work well for first drafts, research summaries, structural suggestions, and editing passes. They work poorly as a replacement for the expertise and judgment that determines whether a piece of content is worth publishing.

The same split applies to the SEO work around the content. Using ChatGPT for SEO is reliable for organising, summarising and drafting; it is unreliable for anything that requires live data, because the model will produce confident keyword volumes and rankings it has no way to know.

## Frequently asked questions

### Is AI-generated content against Google's guidelines?

Not inherently. Google's own FAQ answers this directly: "Appropriate use of AI or automation is not against our guidelines. This means that it is not used to generate content primarily to manipulate search rankings, which is against our spam policies."[^6] The policy targets content produced primarily to manipulate rankings, regardless of the tool used. High-quality AI-assisted content, reviewed and published by a named expert, is not in conflict with any Google guideline. Mass-produced, low-quality AI content published at scale is.

### How can I use AI tools without risking penalties?

Treat AI as a drafting and research aid, not a publishing pipeline. Have a subject matter expert review every piece before publication. Fact-check against current sources. Publish under a named author with verifiable credentials. Maintain topical depth rather than breadth: cover a defined area well rather than everything superficially.

### Does AI-generated content affect E-E-A-T?

Only if it lacks the signals E-E-A-T depends on: named authorship, demonstrated expertise, first-hand experience, and factual accuracy. Those signals can be present in AI-assisted content if a genuine expert is involved in producing and reviewing it. They are typically absent from fully automated content pipelines.

[^1]: [What percentage of new content is AI-generated? — Ahrefs](https://ahrefs.com/blog/what-percentage-of-new-content-is-ai-generated/)
[^2]: [AI search collapse: AI responses collapse when AI retrieves its own generations — Graphite](https://graphite.io/five-percent/ai-search-collapse), Gregory Druck and Ethan Smith, June 2026; academic version [RAG Collapse, arXiv 2608.22118](https://arxiv.org/abs/2608.22118), 22 August 2026. Collapse figures from Graphite's own simulations (three models, 1,019 questions, 1,528 runs); the AI-generated-reference shares are live ChatGPT references from January and June 2026, classified with the GPTZero detector.
[^3]: [360Brew: a decoder-only foundation model for personalized ranking and recommendation (withdrawn) — arXiv](https://arxiv.org/abs/2501.16450)
[^4]: [Is 360Brew used as part of your ranking system? Short answer: No — Tim Jurka, VP Engineering, LinkedIn](https://www.linkedin.com/posts/timjurka_every-day-i-talk-with-members-and-get-to-activity-7449929754679545857-O7Er)
[^5]: [LLMs and combatting misinformation about LinkedIn's algorithm — Phil Szomszor](https://www.linkedin.com/posts/philszomszor_llms-and-combatting-misinformation-about-activity-7472992526031884288-4k6I)
[^6]: [Google Search's guidance about AI-generated content — Google Search Central](https://developers.google.com/search/blog/2023/02/google-search-and-ai-content)
[^8]: [Creating helpful, reliable, people-first content — Google Search Central](https://developers.google.com/search/docs/fundamentals/creating-helpful-content)
[^9]: [Optimizing your website for generative AI features on Google Search — Google Search Central](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide)
[^10]: [How Google's AI Mode Compares to Traditional Search and Other LLMs — Semrush](https://www.semrush.com/blog/ai-mode-comparison-study/). 5,000 keywords, 150,000+ unique citations, July 2025.
[^11]: [AI search loves listicles: What 25,000 URLs reveal about citations — Evertune, via Search Engine Land](https://searchengineland.com/ai-search-loves-listicles-what-25000-urls-reveal-about-citations-477682). Sponsored content; the 6,000 most-cited URLs per model, March and April 2026, from Evertune's own brand-tracking panel.
[^7]: [Scaled content abuse — Google Search Central spam policies](https://developers.google.com/search/docs/essentials/spam-policies#scaled-content)