Duplicate Content Checker: Find Exact & Near-Duplicate Pages
How a duplicate content checker finds exact and near-duplicate pages with SimHash, what it misses, and data from 1,332 pages on 57 sites. Free Python script.
A duplicate content checker compares the main text of every page on your site and reports two things. Exact duplicates are pages whose body text is identical, found by hashing it. Near-duplicates are pages that are almost the same, found with a fingerprint such as SimHash. A good checker also flags pages that share the same <title>, meta description or H1, and it ignores pages that are already consolidated with rel=canonical or excluded with noindex.
That last part decides whether the report is useful or noise. When we crawled 1,332 pages across 57 websites, canonical tags had already merged 65 URL variants on 25 of the sites. A checker that ignored those would have reported dozens of "duplicates" that search engines had already resolved.
This guide explains how exact and near-duplicate detection actually work, how sensitive SimHash really is (we benchmarked it), what duplicate content looks like on real sites, and how to run the same check yourself with a copy-paste Python script.
What Counts as Duplicate Content (and the Penalty Myth)
Duplicate content is substantial blocks of content that match or closely resemble other content on the same site or across sites. In most cases it isn't a penalty. Google's documentation on consolidating duplicate URLs describes what actually happens: Google groups the duplicates, picks one canonical URL to show in results, and filters out the rest.
The cost is in what you lose control of:
- Google may pick a different canonical than you intended, which shows up in Search Console as "Duplicate, Google chose different canonical than user".
- Signals split across URLs: internal links and backlinks point at several versions instead of one.
- Crawl requests are wasted on URLs that will never rank.
- The wrong page ranks. On a store, that might be a filtered category URL instead of the clean one.
Penalties are reserved for deliberately scraped or mass-generated copies meant to manipulate rankings. Your /index.html duplicate is a hygiene problem, not spam.
How Duplicate Content Checkers Work
Exact duplicates: hash the body
Normalise the page's main content (lowercase, collapse whitespace), hash it, and group pages with the same hash. That's fast and has no false positives, but a single changed character breaks the match. So a date in the sidebar or a rotating "related posts" block is enough to hide an exact duplicate.
Near-duplicates: SimHash
SimHash (Charikar's algorithm) turns a document into a 64-bit fingerprint where similar documents differ in only a few bits:
- Split the text into overlapping shingles, for example 4-word sequences.
- Hash each shingle to 64 bits.
- For each bit position, add +1 if the shingle's hash has a 1 there and β1 if it has a 0.
- The fingerprint has a 1 wherever the total is positive.
Two pages are near-duplicates when the Hamming distance between their fingerprints (the number of differing bits) is at or below a threshold. BugViso uses β€ 6 of 64 bits. At scale, fingerprints are split into four 16-bit bands and only pages sharing a band are compared, so a 10,000-page crawl doesn't need 50 million pairwise checks.
Element duplicates: titles, metas, H1s
The cheapest and often most useful check: group pages by their exact <title>, meta description and <h1>. Duplicate titles are the most visible symptom of templated or thin pages, and they're what Google shows in results.
π‘ Fingerprint the main content, not the whole page. Navigation, footers and cookie text are identical on every URL. If you hash the full
<body>, every pair of short pages looks like a near-duplicate.
How Sensitive Is SimHash? A Benchmark
Thresholds are easy to quote and hard to reason about, so we tested one. We took 12 real long-form articles (900 words each) and created controlled variants, then measured the Hamming distance from the original using 64-bit SimHash over 4-word shingles, the same settings as the script below.
| Variant of the original article | Median distance | Range | Flagged at β€6 bits |
|---|---|---|---|
| Identical copy | 0 | 0β0 | 12 / 12 |
| Same article + a 40-word boilerplate footer | 5 | 1β9 | 9 / 12 |
| 2% of the text rewritten (one block) | 3.5 | 1β13 | 10 / 12 |
| 5% of the text rewritten | 6 | 4β11 | 7 / 12 |
| One word swapped every 40 words ("city page" template) | 8.5 | 5β13 | 3 / 12 |
| 10% of the text rewritten | 9.5 | 3β12 | 3 / 12 |
| 25% of the text rewritten | 14.5 | 7β23 | 0 / 12 |
| 50% of the text rewritten | 21 | 11β28 | 0 / 12 |
| A different article from the same site | 32.5 | 26β37 | 0 / 12 |
Three practical conclusions:
- A 6-bit threshold catches pages that are about 95% or more identical. It reliably flags copies, boilerplate-only variants and light edits. It does not flag pages with a quarter of their text changed, and it shouldn't.
- Templated "city pages" sit right on the edge. Swapping one word (the city name) every 40 words pushed the median distance to 8.5, so most of them escape a strict check. If you run location or programmatic pages, test them with a looser threshold (around 10 bits) as a separate audit.
- Unrelated pages land around 32 bits, which is what you'd expect for random 64-bit fingerprints. Anything well below that has real overlap.
What Duplicate Content Looks Like on Real Sites
We sampled 57 websites from the Tranco top-sites list (ranks 1,001β50,000). For each, we rendered the homepage and up to 25 pages linked from it in Chromium, and ran BugViso's duplicate-content analysis on the main-content text: 1,332 pages in total. Sites where most pages redirected off-site, to a consent wall for example, were excluded.
| Finding (57 sites, ~25 top-level pages each) | Sites | Share |
|---|---|---|
β₯1 duplicate <title> | 13 | 22.8% |
| β₯1 duplicate meta description | 15 | 26.3% |
β₯1 duplicate <h1> | 10 | 17.5% |
| Any title / meta / H1 duplicate | 24 | 42.1% |
| Near-duplicate body content (β€6 bits) | 5 | 8.8% |
| Exact duplicate body content | 2 | 3.5% |
URL variants consolidated by rel=canonical | 25 sites | 65 variants |
At the page level, 5.0% of pages shared their title with another page, 8.4% shared a meta description and 7.1% shared an H1. 18.7% had no meta description and 18.5% had no H1.
These are only the pages one click from the homepage, the best-maintained part of any site. Duplication usually grows deeper in, in tag archives, filtered listings and paginated series.
The duplicates we found fell into five recognisable patterns:
| Pattern | Example (paths only) | Type |
|---|---|---|
| Default document duplicate | / and /index.html | Exact |
| Same path on two hostnames | /about-us.html on www. and the bare domain | Exact |
| Tracking or session parameters | /preferences/edit?ref_=β¦ vs /preferences/edit?β¦ReturnUrl=β¦ | Exact |
| Tag pages with almost no unique content | /tag/β¦?lang=β¦ Γ 5 | Near (1β2 bits) |
| Two listing pages showing the same items | /category/game-guides/ vs /trending-guides/ | Near (5 bits) |
Look at the distances. Most near-duplicate pairs were 1β2 bits apart: templates with a few words changed, not "similar topics". That's the problem a near-duplicate check is built for.
Run Your Own Duplicate Content Check
This script reads your sitemap (or a list of URLs), extracts the main content, and reports exact duplicates, near-duplicates and duplicate titles/metas/H1s. It skips noindex pages. Install two packages and point it at your sitemap.
#!/usr/bin/env python3
"""Duplicate content checker: exact + near-duplicate pages, duplicate titles/metas/H1s.
Reads URLs from a sitemap (or a text file with one URL per line), extracts the main
content text, fingerprints it with a 64-bit SimHash over 4-word shingles, and reports:
* exact duplicates (identical normalised body text)
* near-duplicates (SimHash Hamming distance <= 6 of 64 bits: ~95%+ identical text)
* duplicate <title>, meta description and <h1> values
Usage: pip install httpx beautifulsoup4
python3 duplicate_content_checker.py https://example.com/sitemap.xml --limit 300
python3 duplicate_content_checker.py urls.txt
"""
import argparse, asyncio, hashlib, re, sys
from collections import defaultdict
from itertools import combinations
import httpx
from bs4 import BeautifulSoup
MAX_HAMMING = 6
MIN_WORDS = 50
WORD = re.compile(r"[a-z0-9]+(?:['-][a-z0-9]+)*")
def simhash(text):
tokens = WORD.findall(text.lower())[:8000]
shingles = [" ".join(tokens[i:i + 4]) for i in range(max(1, len(tokens) - 3))]
bits = [0] * 64
for sh in shingles:
h = int.from_bytes(hashlib.blake2b(sh.encode(), digest_size=8).digest(), "big")
for i in range(64):
bits[i] += 1 if (h >> i) & 1 else -1
return sum(1 << i for i in range(64) if bits[i] > 0)
def main_text(soup):
for tag in soup(["script", "style", "noscript", "nav", "header", "footer", "aside", "form"]):
tag.decompose()
root = soup.find("main") or soup.find(attrs={"role": "main"}) or soup.body or soup
return re.sub(r"\s+", " ", root.get_text(" ")).strip()
async def load_urls(source, client, limit):
if source.startswith("http"):
xml = (await client.get(source)).text
urls = re.findall(r"<loc>\s*([^<\s]+)\s*</loc>", xml)
children = [u for u in urls if u.endswith(".xml")]
for child in children[:20]: # sitemap index
urls += re.findall(r"<loc>\s*([^<\s]+)\s*</loc>", (await client.get(child)).text)
urls = [u for u in urls if not u.endswith(".xml")]
else:
urls = [l.strip() for l in open(source) if l.strip()]
return list(dict.fromkeys(urls))[:limit]
async def main(source, limit):
headers = {"User-Agent": "Mozilla/5.0 (compatible; dup-audit/1.0)"}
async with httpx.AsyncClient(timeout=20, headers=headers, follow_redirects=True) as client:
urls = await load_urls(source, client, limit)
sem = asyncio.Semaphore(6)
async def fetch(url):
async with sem:
try:
r = await client.get(url)
except httpx.HTTPError:
return None
if r.status_code != 200 or "html" not in r.headers.get("content-type", ""):
return None
soup = BeautifulSoup(r.text, "html.parser")
robots = soup.find("meta", attrs={"name": "robots"})
if robots and "noindex" in robots.get("content", "").lower():
return None # search engines ignore it anyway
meta = soup.find("meta", attrs={"name": "description"})
h1 = soup.find("h1")
page = {"url": url,
"title": (soup.title.string or "").strip().lower() if soup.title else "",
"meta": (meta.get("content", "") if meta else "").strip().lower(),
"h1": h1.get_text(" ", strip=True).lower() if h1 else ""}
text = main_text(soup)
page["words"] = len(WORD.findall(text.lower()))
page["exact"] = hashlib.md5(text.lower().encode()).hexdigest()
page["simhash"] = simhash(text)
return page
pages = [p for p in await asyncio.gather(*(fetch(u) for u in urls)) if p]
print(f"Analysed {len(pages)} indexable HTML pages\n")
body = [p for p in pages if p["words"] >= MIN_WORDS]
groups = defaultdict(list)
for p in body:
groups[p["exact"]].append(p["url"])
exact = [g for g in groups.values() if len(g) > 1]
print(f"== Exact duplicate groups: {len(exact)}")
for g in exact:
print(" " + "\n ".join(g) + "\n")
near = []
for a, b in combinations(body, 2): # O(n^2): fine up to a few thousand pages
if a["exact"] == b["exact"]:
continue
d = bin(a["simhash"] ^ b["simhash"]).count("1")
if d <= MAX_HAMMING:
near.append((d, a["url"], b["url"]))
print(f"== Near-duplicate pairs: {len(near)}")
for d, a, b in sorted(near)[:50]:
print(f" {round((1 - d / 64) * 100, 1)}% similar {a} <-> {b}")
for field in ("title", "meta", "h1"):
dupes = defaultdict(list)
for p in pages:
if p[field]:
dupes[p[field]].append(p["url"])
dupes = {k: v for k, v in dupes.items() if len(v) > 1}
missing = sum(1 for p in pages if not p[field])
print(f"\n== Duplicate {field}: {len(dupes)} value(s) shared by {sum(map(len, dupes.values()))} pages; missing on {missing}")
for value, us in sorted(dupes.items(), key=lambda kv: -len(kv[1]))[:10]:
print(f" [{len(us)}] \"{value[:70]}\" e.g. {us[0]}")
if __name__ == "__main__":
ap = argparse.ArgumentParser()
ap.add_argument("source", help="sitemap URL or file of URLs")
ap.add_argument("--limit", type=int, default=300)
a = ap.parse_args()
asyncio.run(main(a.source, a.limit))Two notes before you trust the output. It reads raw HTML, so on a client-rendered single-page app the main content may be empty; use a rendering crawler there. It also doesn't apply canonical tags. If two URLs both point their canonical at the same page, that's already handled, so drop that pair from your to-do list.
How to Fix Each Type of Duplicate
| Duplicate type | Right fix | Avoid |
|---|---|---|
/index.html, www vs bare, http vs https | 301 to the canonical form at the server | Fixing it only with a canonical tag |
| Tracking and sort parameters | rel=canonical to the clean URL; never link to parameter URLs internally | noindex on pages that should pass link signals |
| Thin tag/archive pages | Merge tags; noindex low-value archives; add unique intro text to the ones you keep (see our thin content guide) | Leaving hundreds of one-post tags indexable |
| Two pages on the same topic | Merge into one stronger page and 301 the weaker URL | Keeping both and hoping Google picks right |
| Duplicate titles and metas | Write unique ones from each page's actual content | Auto-appending "Page 2" and calling it unique |
| Print, AMP and staging copies | rel=canonical to the main page, or block/remove them | Leaving staging hosts publicly indexable |
For the canonical details, our guide to canonical tags and duplicate content covers self-referencing canonicals and cross-domain cases. When two pages compete for the same query rather than sharing text, that's keyword cannibalization, a related problem with a different fix. For parameter-heavy sites, see how to handle URL parameters.
# β
Collapse /index.html and the www host into one canonical URL (Nginx)
server {
server_name www.example.com;
return 301 https://example.com$request_uri;
}
location = /index.html { return 301 /; }How BugViso Detects Duplicate Content
BugViso's Duplicate Content Detection runs across every page in a multi-page crawl. Each page is rendered in Chromium, and the engine picks the element that holds the page's main content: the declared <main> landmark, the skip-link target, or a dominant <article>, falling back to <body>. Navigation and footer boilerplate therefore don't inflate similarity. That text is fingerprinted twice: an exact hash for byte-identical duplicates and a 64-bit SimHash over 4-word shingles for near-duplicates within 6 bits, using banded lookup so large crawls stay fast.
Before comparing, it does what search engines do. noindex pages are excluded, and URLs that declare the same rel=canonical target are merged into one (65 such variants in the dataset above). The report then lists exact-duplicate groups, near-duplicate pairs with a similarity percentage, and pages sharing a title, meta description or H1, each with a suggested fix in the remediation playbook. You can run a BugViso crawl on your site to see all three lists.
Traps That Produce False Positives
- Hashing the whole page. Short pages on a shared template look 90% identical when you include the menu and footer. Compare main content only.
- Ignoring canonicals and
noindex. These are resolved duplicates, not problems. - Pagination.
/blog/page/2and/blog/page/3share a title by design. Since Google retiredrel=prev/next, focus on whether the series is crawlable, as explained in our pagination SEO guide. - Legal pages served behind a consent or login wall. If every URL returns the same interstitial, every page looks like an exact duplicate. Check what your crawler actually received.
- Localised versions. Proper
hreflangalternates in different languages aren't duplicates. Same-language regional copies (en-US vs en-GB) can be, so make sure hreflang ties them together.
FAQ
Is duplicate content a Google penalty?
No, not for ordinary duplication within your own site. Google filters duplicates and shows one canonical version. You lose control over which URL ranks and split your signals, but there's no penalty unless the copying is deceptive or manipulative.
How similar do two pages have to be to count as duplicates?
There's no official percentage. In our benchmark, a 6-bit SimHash threshold flagged pages that were roughly 95% or more identical. That catches copies, parameter variants and templated pages with little unique text, while leaving genuinely different pages on the same topic alone.
What is the best free duplicate content checker?
For your own site, a crawler that compares main content and respects canonical and noindex will beat any copy-paste comparison tool. The script in this article does that from your sitemap. A free BugViso scan adds rendered pages and a full site crawl.
Do duplicate meta descriptions hurt SEO?
They don't hurt rankings directly. They hurt click-through rate, because Google often rewrites them, and they're a strong sign of templated, thin pages. In our data, 26.3% of sites shared at least one meta description across their top-level pages.
Should I use a canonical tag or a 301 redirect for duplicates?
Use a 301 when the duplicate URL has no reason to exist, such as /index.html or the www vs bare host. Use rel=canonical when both URLs must keep working for users (filters, tracking parameters, print versions) but only one should be indexed.
Conclusion
Most duplicate content isn't plagiarism. It's /index.html, a stray hostname, a parameter or a near-empty tag page, and in our data 42% of sites had duplicate titles, metas or H1s on their most visible pages. A checker that compares main content and respects canonicals finds the real problems without the noise, which is exactly how a BugViso site crawl reports them.
See where your site stands
Run a free BugViso audit for SEO, speed, accessibility and AI search readiness β with fixes you can ship today.