Programmatic SEO Without Thin or Duplicate Pages (QA Guide)
Programmatic SEO duplicate content, measured: 497 template pages from 42 families. Why SimHash misses template-heavy pages, quality thresholds and a QA script.
Programmatic SEO pages become duplicate content when the template outweighs the data: hundreds of URLs share the same intro, the same benefits list and the same FAQ, with only a name swapped in. Google's spam policies call the extreme version scaled content abuse — many pages "generated for the primary purpose of manipulating search rankings and not helping users". The fix is to measure, before launch, how much of each page is actually unique, and to hold back pages that don't clear the bar.
We measured real template families to see where the line falls. On 10 October 2026 we sampled sitemaps from 900 random sites and found 45 programmatic page families — integrations, connectors, locations, destinations, glossaries, comparisons, templates, directories — on 33 sites. In the 42 families with real HTML content (497 pages), only 4.4% of page pairs were near-duplicates by SimHash, yet 16.3% of pages had less than 20% unique text and 26.4% had fewer than 300 words of main content. Three more families served pages that were empty without JavaScript.
The lesson: a near-duplicate check alone misses most weak programmatic pages. This guide shows the data, the thresholds we'd use, how to build templates that pass, and a QA script that runs both checks on a batch before you publish.
Why Programmatic Pages End Up "Crawled – Currently Not Indexed"
Google doesn't penalize templates; most good sites use them. Its guidance on creating helpful content asks whether a page adds original value, which is a per-page question. What it does is decide, page by page, whether a URL adds enough to be worth indexing. Programmatic batches fail that test in predictable ways:
- Template-heavy pages. The data block is 80 words; the shared copy around it is 1,200. Each page is long, and none is distinctive.
- Thin data. Hundreds of city pages with one sentence of local information.
- Near-duplicates. Variants that differ by a single attribute ("blue" vs "red", "Boston" vs "Cambridge").
- Empty without JavaScript. Content injected client-side after load, so the HTML Google first sees is the same shell on every URL.
- Weak internal linking. Pages reachable only from the XML sitemap, with no links from hub pages, look unimportant.
The first two are the most common, and they're the ones a SimHash duplicate check doesn't catch.
Data: 497 Programmatic Pages Measured
| Item | Detail |
|---|---|
| Sites | 900 random reachable homepages from the Tranco top 200,000 (seed 20261010); sitemaps read from robots.txt or /sitemap.xml |
| Family | URLs sharing the same path except the last segment, ≥40 URLs, with a path segment that is a programmatic marker (integrations, connectors, templates, glossary, definition, locations, destination, countries, cities, directory, vs, compare, statistics, jobs, tools…) |
| Sample | 12 random pages per family, max 2 families per site → 45 families, 33 sites |
| Text | Raw HTML (no JavaScript) with <header>, <nav>, <footer>, <aside>, forms and scripts removed |
| Metrics | Main-content words; unique share = fraction of a page's 4-word shingles that appear on no other sampled page in its family; SimHash (BugViso's implementation) pairwise Hamming distance, ≤6 of 64 bits = near-duplicate |
| Measurement (42 content families, 497 pages) | Result |
|---|---|
| Median family size (URLs in sitemap) | 176 |
| Median main-content words per page | 729 |
| Pages under 300 words | 26.4% |
| Median unique share per page | 69.7% |
| Pages with < 20% unique text | 16.3% |
| Pages with < 10% unique text (template-only) | 10.9% |
| Page pairs that are near-duplicates (SimHash ≤ 6 bits) | 4.4% |
| Families with at least one near-duplicate pair | 14.3% |
| Families whose median page is < 25% unique | 21.4% |
| Families served as JavaScript shells (pages < 50 words in raw HTML) | 3 of 45 (excluded above) |
By family type:
| Type | Families | Median words | Median unique share | Near-duplicate pairs |
|---|---|---|---|---|
| Locations / destinations | 14 | 688 | 61.7% | 8.2% |
| Integrations / connectors | 6 | 820 | 43.8% | 0.0% |
| Glossary / definitions | 6 | 940 | 88.6% | 0.0% |
| Templates | 2 | 723 | 73.6% | 0.0% |
| Comparisons / alternatives | 2 | 2,146 | 94.3% | 0.0% |
| Directories, statistics, jobs and other | 12 | 286–938 | 35–89% | 0–8% |
Two patterns stand out. Integration pages were long and had zero near-duplicate pairs — yet the median page was only 43.8% unique: a shared description of the core product, shared setup steps and a shared FAQ around a short paragraph about the partner. And location pages were the most likely to be outright near-duplicates (8.2% of pairs). Glossaries and comparison pages, where the data is the content, did best.
Unique share is measured against the 11 other sampled pages, so it slightly overstates uniqueness: against the full family, shared phrasing would appear even more often.
Two Checks, Not One
SimHash (and similar near-duplicate fingerprints) answers "are these two pages essentially the same document?" It's the right tool for variant URLs, parameter duplicates and copy-pasted pages, and it's what crawler duplicate reports use. But a page that's 60% shared template and 40% unique data won't be within 6 bits of anything — it passes, while still being mostly boilerplate.
Unique share answers the programmatic question: "what does this page say that its siblings don't?" Combine them:
| Check | Threshold we'd use | What failing means |
|---|---|---|
| Near-duplicate | SimHash ≤ 6 of 64 bits to any sibling | Merge, canonicalize or drop one of the pair |
| Template-only | < 10% unique 4-word shingles | Don't publish; the data isn't there yet |
| Low unique | < 25% unique | Add data or cut the shared copy before launch |
| Thin | < 300 main-content words and low unique | Expand with real data or noindex |
The thresholds are starting points, not Google numbers — Google publishes none. Calibrate them on your best-performing existing pages: measure what share of their text is unique and hold new batches to the same bar.
Building Templates That Pass
1. Make the data the majority of the page
Count words: shared copy vs page-specific data. If the template contributes most of the text, cut it — move the generic product pitch to one hub page and link to it.
❌ 1,400 words: 1,100 shared (product intro, generic benefits, FAQ) + 300 specific
✅ 700 words: 200 shared (one-paragraph context, CTA) + 500 specific2. Use more data fields, not more adjectives
For integration pages: what syncs (objects and fields), direction, frequency, setup steps that differ per partner, limitations, pricing tier required, a real screenshot. For location pages: local facts the page can only have because of its data — counts, prices, opening hours, nearby entities, reviews tied to that place.
3. Set a publish gate
Generate pages only when the data meets a minimum: at least N specific fields filled, at least M words of specific text. Pages below the gate stay unpublished (or noindex) until the data exists.
4. Server-render the content
Three of the 45 families we found delivered the same near-empty HTML shell on every URL, with the content filled in by JavaScript. Google can render it later; many crawlers never do. Server-render or statically generate programmatic pages.
5. Link them like real pages
Create hub pages (by category, region, use case) that link to every child, link siblings to each other where it helps users ("other CRM integrations"), and keep every page within a few clicks of the homepage. Our internal link audit and orphan pages data study shows how often sitemap-only pages occur.
6. Launch in batches and watch indexing
Publish a first batch of your strongest pages, watch Search Console's Page indexing report for "Crawled – currently not indexed" on that pattern, and only scale once the first batch is indexing. Our diagnostic for pages Google isn't indexing covers the report.
QA a Batch Before Launch: Script
This script runs both checks — unique share and BugViso-compatible SimHash — on a list of URLs or a sitemap prefix, plus thin-page and duplicate title/H1 detection. Point it at staging before launch, or at a live family to find the pages to fix first.
#!/usr/bin/env python3
"""pseo_qa.py: quality-check a batch of programmatic pages before (or after) launch.
Usage:
python3 pseo_qa.py urls.txt # one URL per line (staging or live)
python3 pseo_qa.py --sitemap https://example.com/sitemap.xml --prefix /integrations/ --sample 40
Standard library only. For each page: main-content words (header/nav/footer/aside/scripts removed),
the share of its 4-word shingles that no other page in the batch has, and a 64-bit SimHash
(the same blake2b/4-gram/Charikar scheme BugViso uses; <= 6 differing bits = near-duplicate).
Flags thin pages, template-only pages, near-duplicate pairs and duplicate titles/H1s.
"""
import argparse, hashlib, itertools, random, re, ssl, sys, urllib.request
from collections import defaultdict
UA = "Mozilla/5.0 (compatible; pseo-qa/1.0)"
CTX = ssl.create_default_context()
STRIP = re.compile(r"<(script|style|noscript|svg|header|nav|footer|aside|form)\b.*?</\1>", re.I | re.S)
WORD = re.compile(r"[a-z0-9]+(?:['-][a-z0-9]+)*")
def get(url):
req = urllib.request.Request(url, headers={"User-Agent": UA})
with urllib.request.urlopen(req, timeout=20, context=CTX) as r:
return r.read(3_000_000).decode("utf-8", "replace")
def page(url):
html = get(url)
title = re.search(r"<title[^>]*>(.*?)</title>", html, re.I | re.S)
h1 = re.search(r"<h1\b[^>]*>(.*?)</h1>", html, re.I | re.S)
body = html.split("<body", 1)[-1]
text = re.sub(r"&[a-z#0-9]+;", " ", re.sub(r"<[^>]+>", " ", STRIP.sub(" ", body)))
clean = lambda m: re.sub(r"\s+", " ", re.sub(r"<[^>]+>", " ", m.group(1))).strip() if m else ""
return {"url": url, "title": clean(title), "h1": clean(h1), "tokens": WORD.findall(text.lower())}
def shingles(tokens, k=4):
return {" ".join(tokens[i:i + k]) for i in range(max(0, len(tokens) - k + 1))}
def simhash(tokens):
bits = [0] * 64
for sh in [" ".join(tokens[i:i + 4]) for i in range(max(0, len(tokens) - 3))][:8000]:
h = int.from_bytes(hashlib.blake2b(sh.encode(), digest_size=8).digest(), "big")
for i in range(64):
bits[i] += 1 if (h >> i) & 1 else -1
return sum(1 << i for i in range(64) if bits[i] > 0)
def from_sitemap(url, prefix, n):
locs, queue = [], [url]
while queue and len(locs) < 50000:
xml = get(queue.pop(0))
found = re.findall(r"<loc>\s*([^<\s]+)\s*</loc>", xml)
(queue.extend if "<sitemapindex" in xml[:2000] else locs.extend)(found)
locs = [u for u in locs if prefix in u]
return random.Random(0).sample(locs, min(n, len(locs)))
def main():
ap = argparse.ArgumentParser()
ap.add_argument("urls", nargs="?")
ap.add_argument("--sitemap"); ap.add_argument("--prefix", default="/"); ap.add_argument("--sample", type=int, default=30)
ap.add_argument("--min-words", type=int, default=300); ap.add_argument("--min-unique", type=float, default=0.25)
a = ap.parse_args()
urls = from_sitemap(a.sitemap, a.prefix, a.sample) if a.sitemap else [l.strip() for l in open(a.urls) if l.strip()]
pages = []
for u in urls:
try:
pages.append(page(u))
except Exception as e:
print(f"? {u}: {e}")
if len(pages) < 2:
sys.exit("need at least two pages")
sh = [shingles(p["tokens"]) for p in pages]
for i, p in enumerate(pages):
others = set().union(*(sh[j] for j in range(len(sh)) if j != i))
p["unique"] = len(sh[i] - others) / len(sh[i]) if sh[i] else 0.0
p["hash"] = simhash(p["tokens"])
near = [(a_["url"], b_["url"], (a_["hash"] ^ b_["hash"]).bit_count())
for a_, b_ in itertools.combinations(pages, 2) if (a_["hash"] ^ b_["hash"]).bit_count() <= 6]
dup_titles = {t: us for t, us in _group(pages, "title").items() if len(us) > 1}
dup_h1 = {t: us for t, us in _group(pages, "h1").items() if len(us) > 1}
print(f"{len(pages)} pages checked\n")
print(f"{'words':>6} {'unique':>7} verdict url")
for p in sorted(pages, key=lambda p: p["unique"]):
w = len(p["tokens"])
verdict = ("TEMPLATE-ONLY" if p["unique"] < 0.10 else "low unique" if p["unique"] < a.min_unique else
"thin" if w < a.min_words else "ok")
print(f"{w:>6} {p['unique']:>7.0%} {verdict:15s} {p['url']}")
print(f"\nnear-duplicate pairs (SimHash ≤ 6 of 64 bits): {len(near)}")
for x, y, d in near[:15]:
print(f" {d} bits {x}\n {y}")
print(f"duplicate titles: {sum(len(v) for v in dup_titles.values())} pages in {len(dup_titles)} groups; "
f"duplicate H1s: {sum(len(v) for v in dup_h1.values())} pages in {len(dup_h1)} groups")
bad = sum(1 for p in pages if p["unique"] < a.min_unique or len(p["tokens"]) < a.min_words)
print(f"\n{bad} of {len(pages)} pages below the bar (unique < {a.min_unique:.0%} or < {a.min_words} words)")
return 1 if bad or near else 0
def _group(pages, key):
g = defaultdict(list)
for p in pages:
if p[key]:
g[p[key].lower()].append(p["url"])
return g
if __name__ == "__main__":
sys.exit(main())Real output for 15 integration pages from a CRM vendor in our sample (10 October 2026, domain replaced):
15 pages checked
words unique verdict url
1385 13% low unique https://crm-vendor.example/integrations/homeonline/
1396 20% low unique https://crm-vendor.example/integrations/sarv/
1532 22% low unique https://crm-vendor.example/integrations/google-meet/
1464 22% low unique https://crm-vendor.example/integrations/whatsapp/
1608 25% low unique https://crm-vendor.example/integrations/quickr/
1552 29% ok https://crm-vendor.example/integrations/iim-jobs/
1390 29% ok https://crm-vendor.example/integrations/cashfree/
1590 30% ok https://crm-vendor.example/integrations/monster/
1508 31% ok https://crm-vendor.example/integrations/aweber/
1578 33% ok https://crm-vendor.example/integrations/shopify-integration/
1389 33% ok https://crm-vendor.example/integrations/gmail/
1482 37% ok https://crm-vendor.example/integrations/get-response/
1554 37% ok https://crm-vendor.example/integrations/snov-io/
1366 38% ok https://crm-vendor.example/integrations/zapier-integration/
1703 42% ok https://crm-vendor.example/integrations/zoho-forms/
near-duplicate pairs (SimHash ≤ 6 of 64 bits): 0
duplicate titles: 0 pages in 0 groups; duplicate H1s: 0 pages in 0 groups
5 of 15 pages below the bar (unique < 25% or < 300 words)This is the pattern from the data in one batch: every page is long (1,366–1,703 words) and SimHash finds no near-duplicates, but no page is more than 42% unique and five are under 25%. A crawler's duplicate report would call this family clean. The fix is editorial — trim the shared copy and add partner-specific detail — not technical.
How BugViso Checks Programmatic Pages
A BugViso multi-page crawl runs its duplicate content engine across every crawled page: exact duplicates (identical main-content hash), near-duplicates (64-bit SimHash over 4-word shingles of the main content, ≤6 differing bits — the same scheme as the script) and duplicate titles, meta descriptions and H1s, with canonicalized and noindex variants consolidated so they aren't double-counted. The AI readiness engine flags thin main content under 300 words, and the internal link graph reports orphan pages and click depth, so you can see whether a programmatic family is linked from anywhere but the sitemap.
What a crawl doesn't compute is the unique-share percentage per page, and crawl size is capped by plan (75 to 300 pages), so for a 2,000-page batch crawl a representative section and run the script on a sample for the unique-share check. You can crawl a section of your site with BugViso; the duplicate checks are covered in our duplicate content checker guide and on the advanced SEO intelligence page.
Traps and Edge Cases
- Translated programmatic pages aren't duplicates of each other, but each language version must still be unique within its own language family — and needs correct hreflang.
- Canonicalizing a whole family to the hub hides the pages from search entirely. Only canonicalize true variants.
- AI-written filler doesn't raise unique share in a useful way. It can make text technically unique while adding nothing a user needs — exactly what Google's scaled content abuse policy describes.
- Parameter variants multiply families. Sort and filter parameters on programmatic listings create near-duplicates; keep them out of the sitemap and canonical to the clean URL.
FAQ
Is programmatic SEO considered duplicate content?
Not inherently. It becomes a problem when pages share most of their text and differ only by a name or attribute. In our sample, 16.3% of programmatic pages had less than 20% unique text.
How unique does a programmatic page need to be?
Google publishes no percentage. We'd hold back pages under 10% unique (template-only) and improve pages under 25% before launch, then calibrate against your best-performing existing pages.
Why do my programmatic pages say "Crawled – currently not indexed"?
Usually because Google judged them too similar to siblings or too thin to be worth indexing. Check unique share, word count, whether content is in the raw HTML, and whether the pages are linked from hub pages.
Can a duplicate content checker find thin programmatic pages?
Only partly. SimHash-style checkers catch near-identical pages (4.4% of pairs in our data) but not template-heavy pages that are long and slightly different. Pair it with a unique-share check.
Conclusion
Programmatic pages earn indexing when their data outweighs their template: measure unique share as well as near-duplication, gate publishing on real data, server-render, and link every page from a hub — then use a BugViso crawl to catch duplicates, thin pages and orphans after launch.
See where your site stands
Run a free BugViso audit for SEO, speed, accessibility and AI search readiness — with fixes you can ship today.