Duplicate Content Checker: Find Exact & Near-Duplicate Pages

How a duplicate content checker finds exact and near-duplicate pages with SimHash, what it misses, and data from 1,332 pages on 57 sites. Free Python script.

BugViso

15 min read

A duplicate content checker compares the main text of every page on your site and reports two things. Exact duplicates are pages whose body text is identical, found by hashing it. Near-duplicates are pages that are almost the same, found with a fingerprint such as SimHash. A good checker also flags pages that share the same <title>, meta description or H1, and it ignores pages that are already consolidated with rel=canonical or excluded with noindex.

That last part decides whether the report is useful or noise. When we crawled 1,332 pages across 57 websites, canonical tags had already merged 65 URL variants on 25 of the sites. A checker that ignored those would have reported dozens of "duplicates" that search engines had already resolved.

This guide explains how exact and near-duplicate detection actually work, how sensitive SimHash really is (we benchmarked it), what duplicate content looks like on real sites, and how to run the same check yourself with a copy-paste Python script.


What Counts as Duplicate Content (and the Penalty Myth)

Duplicate content is substantial blocks of content that match or closely resemble other content on the same site or across sites. In most cases it isn't a penalty. Google's documentation on consolidating duplicate URLs describes what actually happens: Google groups the duplicates, picks one canonical URL to show in results, and filters out the rest.

The cost is in what you lose control of:

  • Google may pick a different canonical than you intended, which shows up in Search Console as "Duplicate, Google chose different canonical than user".
  • Signals split across URLs: internal links and backlinks point at several versions instead of one.
  • Crawl requests are wasted on URLs that will never rank.
  • The wrong page ranks. On a store, that might be a filtered category URL instead of the clean one.

Penalties are reserved for deliberately scraped or mass-generated copies meant to manipulate rankings. Your /index.html duplicate is a hygiene problem, not spam.


How Duplicate Content Checkers Work

Exact duplicates: hash the body

Normalise the page's main content (lowercase, collapse whitespace), hash it, and group pages with the same hash. That's fast and has no false positives, but a single changed character breaks the match. So a date in the sidebar or a rotating "related posts" block is enough to hide an exact duplicate.

Near-duplicates: SimHash

SimHash (Charikar's algorithm) turns a document into a 64-bit fingerprint where similar documents differ in only a few bits:

  1. Split the text into overlapping shingles, for example 4-word sequences.
  2. Hash each shingle to 64 bits.
  3. For each bit position, add +1 if the shingle's hash has a 1 there and βˆ’1 if it has a 0.
  4. The fingerprint has a 1 wherever the total is positive.

Two pages are near-duplicates when the Hamming distance between their fingerprints (the number of differing bits) is at or below a threshold. BugViso uses ≀ 6 of 64 bits. At scale, fingerprints are split into four 16-bit bands and only pages sharing a band are compared, so a 10,000-page crawl doesn't need 50 million pairwise checks.

Element duplicates: titles, metas, H1s

The cheapest and often most useful check: group pages by their exact <title>, meta description and <h1>. Duplicate titles are the most visible symptom of templated or thin pages, and they're what Google shows in results.

πŸ’‘ Fingerprint the main content, not the whole page. Navigation, footers and cookie text are identical on every URL. If you hash the full <body>, every pair of short pages looks like a near-duplicate.


How Sensitive Is SimHash? A Benchmark

Thresholds are easy to quote and hard to reason about, so we tested one. We took 12 real long-form articles (900 words each) and created controlled variants, then measured the Hamming distance from the original using 64-bit SimHash over 4-word shingles, the same settings as the script below.

Variant of the original articleMedian distanceRangeFlagged at ≀6 bits
Identical copy00–012 / 12
Same article + a 40-word boilerplate footer51–99 / 12
2% of the text rewritten (one block)3.51–1310 / 12
5% of the text rewritten64–117 / 12
One word swapped every 40 words ("city page" template)8.55–133 / 12
10% of the text rewritten9.53–123 / 12
25% of the text rewritten14.57–230 / 12
50% of the text rewritten2111–280 / 12
A different article from the same site32.526–370 / 12

Three practical conclusions:

  1. A 6-bit threshold catches pages that are about 95% or more identical. It reliably flags copies, boilerplate-only variants and light edits. It does not flag pages with a quarter of their text changed, and it shouldn't.
  2. Templated "city pages" sit right on the edge. Swapping one word (the city name) every 40 words pushed the median distance to 8.5, so most of them escape a strict check. If you run location or programmatic pages, test them with a looser threshold (around 10 bits) as a separate audit.
  3. Unrelated pages land around 32 bits, which is what you'd expect for random 64-bit fingerprints. Anything well below that has real overlap.

What Duplicate Content Looks Like on Real Sites

We sampled 57 websites from the Tranco top-sites list (ranks 1,001–50,000). For each, we rendered the homepage and up to 25 pages linked from it in Chromium, and ran BugViso's duplicate-content analysis on the main-content text: 1,332 pages in total. Sites where most pages redirected off-site, to a consent wall for example, were excluded.

Finding (57 sites, ~25 top-level pages each)SitesShare
β‰₯1 duplicate <title>1322.8%
β‰₯1 duplicate meta description1526.3%
β‰₯1 duplicate <h1>1017.5%
Any title / meta / H1 duplicate2442.1%
Near-duplicate body content (≀6 bits)58.8%
Exact duplicate body content23.5%
URL variants consolidated by rel=canonical25 sites65 variants

At the page level, 5.0% of pages shared their title with another page, 8.4% shared a meta description and 7.1% shared an H1. 18.7% had no meta description and 18.5% had no H1.

These are only the pages one click from the homepage, the best-maintained part of any site. Duplication usually grows deeper in, in tag archives, filtered listings and paginated series.

The duplicates we found fell into five recognisable patterns:

PatternExample (paths only)Type
Default document duplicate/ and /index.htmlExact
Same path on two hostnames/about-us.html on www. and the bare domainExact
Tracking or session parameters/preferences/edit?ref_=… vs /preferences/edit?…ReturnUrl=…Exact
Tag pages with almost no unique content/tag/…?lang=… Γ— 5Near (1–2 bits)
Two listing pages showing the same items/category/game-guides/ vs /trending-guides/Near (5 bits)

Look at the distances. Most near-duplicate pairs were 1–2 bits apart: templates with a few words changed, not "similar topics". That's the problem a near-duplicate check is built for.


Run Your Own Duplicate Content Check

This script reads your sitemap (or a list of URLs), extracts the main content, and reports exact duplicates, near-duplicates and duplicate titles/metas/H1s. It skips noindex pages. Install two packages and point it at your sitemap.

python
#!/usr/bin/env python3
"""Duplicate content checker: exact + near-duplicate pages, duplicate titles/metas/H1s.

Reads URLs from a sitemap (or a text file with one URL per line), extracts the main
content text, fingerprints it with a 64-bit SimHash over 4-word shingles, and reports:
  * exact duplicates  (identical normalised body text)
  * near-duplicates   (SimHash Hamming distance <= 6 of 64 bits: ~95%+ identical text)
  * duplicate <title>, meta description and <h1> values

Usage:  pip install httpx beautifulsoup4
        python3 duplicate_content_checker.py https://example.com/sitemap.xml --limit 300
        python3 duplicate_content_checker.py urls.txt
"""
import argparse, asyncio, hashlib, re, sys
from collections import defaultdict
from itertools import combinations

import httpx
from bs4 import BeautifulSoup

MAX_HAMMING = 6
MIN_WORDS = 50
WORD = re.compile(r"[a-z0-9]+(?:['-][a-z0-9]+)*")


def simhash(text):
    tokens = WORD.findall(text.lower())[:8000]
    shingles = [" ".join(tokens[i:i + 4]) for i in range(max(1, len(tokens) - 3))]
    bits = [0] * 64
    for sh in shingles:
        h = int.from_bytes(hashlib.blake2b(sh.encode(), digest_size=8).digest(), "big")
        for i in range(64):
            bits[i] += 1 if (h >> i) & 1 else -1
    return sum(1 << i for i in range(64) if bits[i] > 0)


def main_text(soup):
    for tag in soup(["script", "style", "noscript", "nav", "header", "footer", "aside", "form"]):
        tag.decompose()
    root = soup.find("main") or soup.find(attrs={"role": "main"}) or soup.body or soup
    return re.sub(r"\s+", " ", root.get_text(" ")).strip()


async def load_urls(source, client, limit):
    if source.startswith("http"):
        xml = (await client.get(source)).text
        urls = re.findall(r"<loc>\s*([^<\s]+)\s*</loc>", xml)
        children = [u for u in urls if u.endswith(".xml")]
        for child in children[:20]:                     # sitemap index
            urls += re.findall(r"<loc>\s*([^<\s]+)\s*</loc>", (await client.get(child)).text)
        urls = [u for u in urls if not u.endswith(".xml")]
    else:
        urls = [l.strip() for l in open(source) if l.strip()]
    return list(dict.fromkeys(urls))[:limit]


async def main(source, limit):
    headers = {"User-Agent": "Mozilla/5.0 (compatible; dup-audit/1.0)"}
    async with httpx.AsyncClient(timeout=20, headers=headers, follow_redirects=True) as client:
        urls = await load_urls(source, client, limit)
        sem = asyncio.Semaphore(6)

        async def fetch(url):
            async with sem:
                try:
                    r = await client.get(url)
                except httpx.HTTPError:
                    return None
                if r.status_code != 200 or "html" not in r.headers.get("content-type", ""):
                    return None
                soup = BeautifulSoup(r.text, "html.parser")
                robots = soup.find("meta", attrs={"name": "robots"})
                if robots and "noindex" in robots.get("content", "").lower():
                    return None                         # search engines ignore it anyway
                meta = soup.find("meta", attrs={"name": "description"})
                h1 = soup.find("h1")
                page = {"url": url,
                        "title": (soup.title.string or "").strip().lower() if soup.title else "",
                        "meta": (meta.get("content", "") if meta else "").strip().lower(),
                        "h1": h1.get_text(" ", strip=True).lower() if h1 else ""}
                text = main_text(soup)
                page["words"] = len(WORD.findall(text.lower()))
                page["exact"] = hashlib.md5(text.lower().encode()).hexdigest()
                page["simhash"] = simhash(text)
                return page

        pages = [p for p in await asyncio.gather(*(fetch(u) for u in urls)) if p]

    print(f"Analysed {len(pages)} indexable HTML pages\n")
    body = [p for p in pages if p["words"] >= MIN_WORDS]

    groups = defaultdict(list)
    for p in body:
        groups[p["exact"]].append(p["url"])
    exact = [g for g in groups.values() if len(g) > 1]
    print(f"== Exact duplicate groups: {len(exact)}")
    for g in exact:
        print("  " + "\n  ".join(g) + "\n")

    near = []
    for a, b in combinations(body, 2):              # O(n^2): fine up to a few thousand pages
        if a["exact"] == b["exact"]:
            continue
        d = bin(a["simhash"] ^ b["simhash"]).count("1")
        if d <= MAX_HAMMING:
            near.append((d, a["url"], b["url"]))
    print(f"== Near-duplicate pairs: {len(near)}")
    for d, a, b in sorted(near)[:50]:
        print(f"  {round((1 - d / 64) * 100, 1)}% similar  {a}  <->  {b}")

    for field in ("title", "meta", "h1"):
        dupes = defaultdict(list)
        for p in pages:
            if p[field]:
                dupes[p[field]].append(p["url"])
        dupes = {k: v for k, v in dupes.items() if len(v) > 1}
        missing = sum(1 for p in pages if not p[field])
        print(f"\n== Duplicate {field}: {len(dupes)} value(s) shared by {sum(map(len, dupes.values()))} pages; missing on {missing}")
        for value, us in sorted(dupes.items(), key=lambda kv: -len(kv[1]))[:10]:
            print(f"  [{len(us)}] \"{value[:70]}\"  e.g. {us[0]}")


if __name__ == "__main__":
    ap = argparse.ArgumentParser()
    ap.add_argument("source", help="sitemap URL or file of URLs")
    ap.add_argument("--limit", type=int, default=300)
    a = ap.parse_args()
    asyncio.run(main(a.source, a.limit))

Two notes before you trust the output. It reads raw HTML, so on a client-rendered single-page app the main content may be empty; use a rendering crawler there. It also doesn't apply canonical tags. If two URLs both point their canonical at the same page, that's already handled, so drop that pair from your to-do list.


How to Fix Each Type of Duplicate

Duplicate typeRight fixAvoid
/index.html, www vs bare, http vs https301 to the canonical form at the serverFixing it only with a canonical tag
Tracking and sort parametersrel=canonical to the clean URL; never link to parameter URLs internallynoindex on pages that should pass link signals
Thin tag/archive pagesMerge tags; noindex low-value archives; add unique intro text to the ones you keep (see our thin content guide)Leaving hundreds of one-post tags indexable
Two pages on the same topicMerge into one stronger page and 301 the weaker URLKeeping both and hoping Google picks right
Duplicate titles and metasWrite unique ones from each page's actual contentAuto-appending "Page 2" and calling it unique
Print, AMP and staging copiesrel=canonical to the main page, or block/remove themLeaving staging hosts publicly indexable

For the canonical details, our guide to canonical tags and duplicate content covers self-referencing canonicals and cross-domain cases. When two pages compete for the same query rather than sharing text, that's keyword cannibalization, a related problem with a different fix. For parameter-heavy sites, see how to handle URL parameters.

nginx
# βœ… Collapse /index.html and the www host into one canonical URL (Nginx)
server {
    server_name www.example.com;
    return 301 https://example.com$request_uri;
}
location = /index.html { return 301 /; }

How BugViso Detects Duplicate Content

BugViso's Duplicate Content Detection runs across every page in a multi-page crawl. Each page is rendered in Chromium, and the engine picks the element that holds the page's main content: the declared <main> landmark, the skip-link target, or a dominant <article>, falling back to <body>. Navigation and footer boilerplate therefore don't inflate similarity. That text is fingerprinted twice: an exact hash for byte-identical duplicates and a 64-bit SimHash over 4-word shingles for near-duplicates within 6 bits, using banded lookup so large crawls stay fast.

Before comparing, it does what search engines do. noindex pages are excluded, and URLs that declare the same rel=canonical target are merged into one (65 such variants in the dataset above). The report then lists exact-duplicate groups, near-duplicate pairs with a similarity percentage, and pages sharing a title, meta description or H1, each with a suggested fix in the remediation playbook. You can run a BugViso crawl on your site to see all three lists.


Traps That Produce False Positives

  • Hashing the whole page. Short pages on a shared template look 90% identical when you include the menu and footer. Compare main content only.
  • Ignoring canonicals and noindex. These are resolved duplicates, not problems.
  • Pagination. /blog/page/2 and /blog/page/3 share a title by design. Since Google retired rel=prev/next, focus on whether the series is crawlable, as explained in our pagination SEO guide.
  • Legal pages served behind a consent or login wall. If every URL returns the same interstitial, every page looks like an exact duplicate. Check what your crawler actually received.
  • Localised versions. Proper hreflang alternates in different languages aren't duplicates. Same-language regional copies (en-US vs en-GB) can be, so make sure hreflang ties them together.

FAQ

Is duplicate content a Google penalty?

No, not for ordinary duplication within your own site. Google filters duplicates and shows one canonical version. You lose control over which URL ranks and split your signals, but there's no penalty unless the copying is deceptive or manipulative.

How similar do two pages have to be to count as duplicates?

There's no official percentage. In our benchmark, a 6-bit SimHash threshold flagged pages that were roughly 95% or more identical. That catches copies, parameter variants and templated pages with little unique text, while leaving genuinely different pages on the same topic alone.

What is the best free duplicate content checker?

For your own site, a crawler that compares main content and respects canonical and noindex will beat any copy-paste comparison tool. The script in this article does that from your sitemap. A free BugViso scan adds rendered pages and a full site crawl.

Do duplicate meta descriptions hurt SEO?

They don't hurt rankings directly. They hurt click-through rate, because Google often rewrites them, and they're a strong sign of templated, thin pages. In our data, 26.3% of sites shared at least one meta description across their top-level pages.

Should I use a canonical tag or a 301 redirect for duplicates?

Use a 301 when the duplicate URL has no reason to exist, such as /index.html or the www vs bare host. Use rel=canonical when both URLs must keep working for users (filters, tracking parameters, print versions) but only one should be indexed.


Conclusion

Most duplicate content isn't plagiarism. It's /index.html, a stray hostname, a parameter or a near-empty tag page, and in our data 42% of sites had duplicate titles, metas or H1s on their most visible pages. A checker that compares main content and respects canonicals finds the real problems without the noise, which is exactly how a BugViso site crawl reports them.

Found this useful? Share it.

See where your site stands

Run a free BugViso audit for SEO, speed, accessibility and AI search readiness β€” with fixes you can ship today.