what is duplicate content? the seo ghost that mostly haunts e-commerce and copy-paste sites
seo · Mar 29, 2026 · 4 min read
duplicate content is the same (or nearly the same) text reachable at multiple urls — inside your site or across sites. the myth says google punishes it. the reality is quieter: google does not penalize duplication; it de-duplicates — it picks one version to rank and quietly ignores the rest. the cost is not a penalty; it is dilution and lost control of which version gets picked.
where duplication actually comes from
- the same page at several urls. with-and-without trailing slash, http and https both live, print versions, parameters ( ?color=red creating a new url for the same product) — plumbing problems, not content problems.
- e-commerce filters and categories. the same product in three categories creates three urls; the manufacturer’s description pasted on forty products creates forty near-identical pages. this is where duplication actually costs money.
- copy-paste content. the manufacturer text, the legal boilerplate, the service description reused across city pages. google’s index has seen the text before; the page competes against whoever published it first — usually not you.
- scrapers. someone republishing your content. annoying, mostly harmless to you: the original usually wins. file a takedown when it matters; do not lose sleep.
when it hurts and when it does not
hurts: when your important page competes with its own twins for ranking (dilution), when whole pages exist only as pasted text (they rank for nothing), when the “wrong” version is the one google shows. does not hurt: boilerplate snippets (a legal line, a shipping note) repeated across pages — that is normal site furniture.
the fixes, mostly plumbing
- one canonical url per page: redirect the http→https, the slash variants, the parameter twins — the redirect tool is written in what is a 301 redirect.
- unique value on every page you want ranked: your words, your photos, your prices, your town — the specificity argument is the backbone of does copying text hurt my seo.
- for genuinely identical pages that must exist: the canonical link tag, which tells google which twin is the real one. one line, done.
FAQ
is my italian page a duplicate of my english one?
no — different languages are different content, and hreflang tells google they are alternatives, not duplicates. the cluster logic is in forty-two repositories.
i copied my own brochure text to the site — is that duplication?
only if the brochure text is also online somewhere else. if the words exist nowhere else, the page is original — originality is about the index, not about who wrote it.
what about ai-written pages?
ai paraphrases what exists, so mass-produced ai pages tend to converge toward the median — which is already indexed. thin and derivative loses regardless of authorship; the ai question is handled honestly in my take on ai.
the closing thought
duplicate content is rarely the emergency seo vendors describe — it is mostly plumbing plus the habit of pasting instead of writing. one url per page, your own words on the pages that matter, canonicals where twins are unavoidable: the ghost dissolves into configuration.
if you want the site audited for what google actually sees:
- web development in Parma — clean urls by construction
- what is google search console — the tool that shows the duplication, free