Black-Hat SEO
Last updated
Black-hat SEO refers to techniques that attempt to game search engine rankings by violating the guidelines search engines publish. The terms originate from old Western films, where villains wore black hats and heroes wore white hats. The convention passed into computing as shorthand for malicious versus ethical behaviour, and from there into SEO.
These techniques are not always ineffective in the short term, which is why they persist. The risk is not that they never work. They produce fragile rankings that can disappear in a single algorithm update, and expose sites to manual actions that can take months to recover from.
Keyword stuffing
Keyword stuffing involves cramming target keywords into a page at an unnaturally high density. Early search engines ranked pages partly on keyword frequency, so repeating a phrase dozens of times across a page, in body text, headings, footers, and meta tags, could push that page into top positions.
Google’s language models now understand topic coverage, not keyword frequency. A page stuffed with a phrase reads as lower quality than one that covers the topic naturally. Keyword stuffing is both ineffective and identifiable.
Scaled content abuse
Scaled content abuse is generating large numbers of pages primarily to manipulate rankings rather than to help users. Google made it a named spam policy in March 2024, broadening the older thin content framing to capture mass production by any method.1
The method is not what defines the violation. Using generative AI to publish hundreds of pages is not penalised because the content is AI-written; it is penalised when pages are produced at scale to capture search traffic without adding value. Google has been explicit that using AI is not itself against its guidelines, the line is whether the output is genuinely useful or simply manufactured volume.2 Scraping and lightly rewording feeds, stitching passages together from other pages, and spinning one article into dozens of near-duplicates all fall under the same policy, whether a human or a model produced them.
Scraping and duplicate content
Scraping is taking content from other sites, often by automated means, and republishing it to manipulate rankings. Google names it as a spam policy in its own right.3 It covers republishing articles with no original value or credit, copying content and modifying it only slightly with synonyms or automated tools, reproducing feeds wholesale, and compiling other people’s images or videos without adding anything of substance.
This is the deliberate, manipulative side of duplicate content, distinct from the technical duplication Google handles routinely. Copying a competitor’s pages, or generating near-identical templates that change only a city name to capture local searches, is a spam tactic, not a housekeeping issue. Unintentional duplication, syndication, CMS-generated URL variants, printer-friendly versions, is resolved with canonical tags and redirects and carries no penalty; scraping and large-scale copying invite a manual action.
Cloaking
Cloaking means serving different content to search engine crawlers than to human visitors. A page might show Googlebot keyword-rich text while showing users a thin or entirely unrelated page.
This is one of the most explicitly prohibited techniques in Google’s spam policies. When detected, it typically results in a manual action rather than just algorithmic suppression. Google’s rendering pipeline has become sophisticated enough to detect simple cloaking reliably, which has made it both riskier and less effective.
Sneaky redirects
A sneaky redirect sends users to a different destination than the one search engines were shown, or somewhere they did not expect.4 A page might present search-friendly content to Googlebot while redirecting human visitors to an unrelated page, or serve desktop users a normal page while pushing mobile users to a different spam domain entirely.
It is closely related to cloaking: both turn on showing one thing to search engines and another to users. The difference is mechanism, cloaking serves divergent content in place, while a sneaky redirect moves the user elsewhere. Redirects are a normal and necessary part of running a site; the violation is the deceptive intent, using a redirect to mislead either users or search engines.
Hidden text and links
Placing white text on a white background, setting font size to zero, or positioning content off-screen creates text that users cannot see but that crawlers could historically read. This was used to add keyword-rich content or additional links without affecting the visible page.
Modern rendering detects hidden text, and it is listed explicitly in Google’s spam policies. Hidden links are treated as manipulative, regardless of how they are implemented.
Doorway pages
Doorway pages are low-quality pages created to rank for specific search queries and then redirect users to a different destination. A site might create hundreds of pages each targeting a different city or keyword variant, with the sole purpose of capturing traffic and passing it somewhere else.
These pages provide no value at the query level: they exist only as a funnel. Google’s documentation explicitly calls them out as a violation, and large-scale doorway page networks are routinely detected and removed from the index.
Private blog networks (PBNs)
A private blog network is a collection of sites created or acquired specifically to point links at a target site. The operator controls all the sites and manufactures link equity rather than earning it.
PBNs exploit the principle that backlinks signal authority. A site with hundreds of links from “independent” domains looks well-cited. The fraud is that the independence is manufactured. Google’s link quality assessment has improved significantly: patterns of shared hosting, registrar data, linking footprints, and content similarity help detect coordinated networks. Detected PBN links are devalued or generate manual actions against the receiving site.
Manipulative link schemes
Beyond PBNs, Google’s spam policies cover a broader category of manipulative link activity:
- Buying links that pass PageRank (as opposed to clearly marked paid or sponsored links)
- Excessive reciprocal linking where sites systematically exchange links in a way that mimics editorial citation without actually providing it
- Large-scale guest posting purely for link placement, particularly on low-quality or irrelevant sites
- Automated link building through directory submissions, forum signatures, comment spam, and similar mass-submission tactics
The common thread is links obtained through a scheme rather than earned through genuine editorial endorsement.
Free products for reviews are a link scheme
Sending a blogger or reviewer a free product in exchange for coverage that includes a followed link is a link scheme under Google's guidelines, the same category as buying links: Google's own example of link spam is "sending someone a product in exchange for them writing about it and including a link". It is common in influencer and PR outreach and rarely recognised as a violation. The fix is not to stop gifting products, but to mark any resulting link rel="sponsored" or rel="nofollow". In the UK there is a second, separate duty: the ASA and CAP Code treat a gifted product as advertising, so the post must be clearly labelled as an ad (not just "#gifted"), and since the DMCC Act 2024 the CMA can enforce that directly. The endorsement can be genuine, but both the link and the disclosure have to be honest that something of value changed hands.
Entity stacking
Entity stacking, also called Google entity stacking, Google entity stacks, or authority stacking, is a link-building tactic that interlinks Google’s own properties, Google Docs, Sheets, Slides, Sites, and custom Google Maps, then points them at a target site. It evolved from the older “Web 2.0” link schemes, and a cottage industry of services sells it as a shortcut, particularly in local SEO.
The pitch is dressed in the language of entity SEO and the Knowledge Graph: the claim is that interlinking Google-owned properties teaches Google to recognise a brand as an authoritative entity. The underlying argument is cruder, that Google is unlikely to discount links sitting on its own high-authority domains. It does. Google assesses links at the level of the individual page and its content, not the registrable domain, so a thin, auto-generated document on a shared host carries no editorial endorsement and is exactly the kind of manufactured signal Google’s link spam systems are built to ignore.5
The deeper flaw is conceptual. Genuine entity strength comes from independent third parties referencing a brand, corroboration Google can see the brand does not control. A folder of self-created Google documents is not independent corroboration; it is the same self-referential signal as a private blog network, hosted on Google. Kalicube, among the most established authorities on entity SEO, explicitly does not use or recommend the technique.6 Vendors market it as white-hat because it uses Google’s platforms, but the platform is not what makes a link legitimate. Where the stacked properties exist solely to host links, the tactic is a link scheme by Google’s definition. At best the links are ignored; at worst a pattern of them contributes to a manual action.
What is cloud stacking?
Cloud stacking is the broader, platform-agnostic version of the same idea: rather than restricting itself to Google properties, it hosts thin pages or documents on any high-authority cloud platform, Amazon Web Services, Microsoft Azure, Google Cloud, GitHub, Netlify, then links them back to the target site. The selling point is that a URL on an s3.amazonaws.com or cloud.google.com host inherits the trust of a high-authority domain.
It fails for the same reason entity stacking does. The host’s authority does not transfer to a page that no one links to, references, or visits, and mass-producing such pages to manufacture links is a link scheme regardless of how reputable the underlying infrastructure is. The distinction between the two terms is one of platform scope and marketing vocabulary, not of mechanism or risk.
Expired domain abuse
Expired domain abuse is buying a lapsed domain name and repurposing it primarily to exploit the ranking signals it built up under its previous owner. Like scaled content abuse, it became a named Google spam policy in March 2024.7
The value being mined is the domain’s history: its backlinks, age, and residual trust. An expired charity domain repurposed for affiliate content, or a former school domain filled with casino pages, inherits authority the new content never earned. That mismatch between a domain’s reputation and the content now hosted on it is the signal Google uses to identify the abuse. It is distinct from legitimately acquiring a domain to continue or genuinely rebuild what was there.
How Google responded: Panda and Penguin
Understanding black-hat SEO historically means understanding why Google launched its two most significant spam-targeting updates.
Panda (2011) was Google’s response to thin, low-quality, and duplicated content that had gamed early ranking signals. Sites built around scraped content, keyword-stuffed articles, and content farms were ranking for valuable queries. Panda applied a quality assessment to sites and demoted those with a high proportion of low-quality pages. It updated periodically until 2016, when it was integrated into Google’s core ranking systems permanently. It now runs continuously.
Penguin (2012) was Google’s response to manipulative link building. Sites with unnatural link profiles, characterised by keyword-rich anchor text patterns, low-quality linking domains, and sudden link velocity spikes, were suppressed. Like Panda, Penguin was eventually integrated into core systems in 2016 and now operates in real time, assessing link profiles continuously rather than in periodic batches.
Both updates were direct algorithmic responses to widespread black-hat practices. Their integration into core systems means there is no longer a “safe window” between updates where manipulative techniques can be used without risk.
The current risk profile
The practical risk of black-hat SEO today goes beyond ineffectiveness:
Algorithmic suppression: Techniques that trigger Google’s spam classifiers result in rankings being demoted, often without any explicit notification. A site may simply stop ranking for queries it previously appeared in.
Manual actions: Google’s spam team issues manual actions for egregious violations. These appear in Google Search Console under “Manual actions” and suppress affected pages or the entire site. Recovery requires removing the violating content or links, then submitting a reconsideration request, a process that typically takes weeks to months.
Deindexation: For the most severe violations, particularly cloaking and mass-scale spam, Google removes the site from its index entirely.
Lingering manipulative links: Sites that built manipulative link profiles may keep some of those links long after abandoning the tactic. Since Penguin became part of core in 2016, Google generally neutralises manipulative links by ignoring them rather than penalising the site they point at, so in the ordinary case they need no active cleanup. The disavow file is reserved for the narrow situation it still exists for: an unnatural-links manual action where the links cannot be removed at source.
The asymmetry matters: the short-term rankings gained through black-hat techniques are fragile and temporary; the recovery process if penalised is slow and uncertain. White-hat SEO produces the reverse: rankings built on genuine quality signals tend to hold and strengthen as Google’s ability to measure those signals improves.
If you have inherited a site in the aftermath of this, the first task is diagnosis rather than remediation: a manual action, an algorithmic demotion and an unrelated technical fault all present as the same fall in traffic. The SEO recovery guide separates them before prescribing anything.