Shopify Duplicate Content: Fix Collection URLs and Variants
Shopify duplicate content, measured on 64 stores: what canonicals already handle, why 25% of themes still link collection-path URLs, and the Liquid fixes.
Shopify duplicate content comes from four URL patterns every store has: the same product at /products/<handle> and at /collections/<collection>/products/<handle>, ?variant= URLs for each option, tag-filtered collection pages (/collections/<collection>/<tag>), and sorted or paginated collection URLs. Shopify's default canonical tags already consolidate most of them. What you have to fix is your theme: product grids that link to the collection-path URLs, and tag pages that stay indexable.
We measured how stores actually behave. On 10 October 2026 we checked 64 Shopify stores found in a random sample of the Tranco top 200,000. Canonicals were in good shape: 49 of 50 collection-path product URLs either redirected or canonicalized to /products/, and 95.5% of ?variant= URLs canonicalized to the clean product URL. But 25.0% of collection grids still linked to the collection-path version of their products, 8 of 10 tag-filtered collection pages we found were indexable and self-canonical, and 11% of stores had replaced Shopify's default robots.txt rules.
This guide explains each pattern, what Liquid edits can fix and what Shopify won't let you change, and gives you a checker script.
The Four Duplicate Patterns
| Pattern | Example | Shopify's default handling | Your job |
|---|---|---|---|
| Collection-path product URLs | /collections/shirts/products/oxford-blue | Canonical points to /products/oxford-blue | Stop linking to them from grids |
| Variant URLs | /products/oxford-blue?variant=4123… | Canonical drops ?variant= | Usually nothing |
| Tag-filtered collections | /collections/shirts/linen | Self-canonical, indexable | Decide: index as a landing page or noindex |
| Sort and pagination | ?sort_by=price-ascending, ?page=2 | Sort canonicalized; robots.txt disallows sort and filter combinations | Keep the defaults |
Shopify also creates /collections/all, /collections/vendors?q= and /collections/types?q=. They're legitimate pages, but they shouldn't compete with your real collections.
Data: 64 Stores
| Item | Detail |
|---|---|
| Sample | Tranco 94GG2 ranks 1,001–200,000, 6,000 random domains (seed 20261010), 3,416 reachable homepages, 64 identified as Shopify (Shopify CDN assets or headers) |
| Per store | Homepage → first linked collection → product links in its grid → one product, its collection-path URL, a ?variant= URL (when the product had several variants), a tag-filtered collection page (when linked), ?page=2, ?sort_by=, robots.txt |
| Method | Plain HTTP requests 1.5 s apart, no JavaScript, 10 October 2026 |
| Measurement | Result |
|---|---|
| Product page canonical is self-referencing | 56 of 56 (100%) |
Collection-path product URL answers 200 / redirects to /products/ / 404 | 35 / 15 / 6 |
| …of the 35 served at the collection path, canonical → product URL | 34 (97.1%) |
Collection grids linking via /collections/…/products/ | 14 of 56 (25.0%); 3 linked only that way |
?variant= URL canonicalizes to the clean product URL | 21 of 22 (95.5%) |
Variant URLs with the same <title> as the product | 22 of 22 |
| Tag-filtered collection pages: indexable and self-canonical | 8 of 10 |
?page=2 keeps its own canonical (recommended) | 51 of 61 (83.6%); 6 pointed to page 1 |
?sort_by= canonicalized to the clean collection URL | 59 of 60 (98.3%) |
| robots.txt still contains Shopify's default sort/filter rules | 56 of 63 (88.9%) |
| Homepage declares hreflang (usually Shopify Markets) | 26 of 64 (40.6%) |
Two conclusions follow. First, canonicals aren't the problem on most stores (Google treats them as a strong hint, as its guide to consolidating duplicate URLs explains): the platform and most themes get them right. Second, internal links are. A quarter of collection grids send crawlers and link equity to non-canonical URLs, so Google keeps crawling and consolidating thousands of duplicates that your own theme creates.
Fix 1: Link Product Grids to /products/
Older and some current themes build product links with Shopify's within filter, which produces the collection path:
{% comment %} ❌ Produces /collections/shirts/products/oxford-blue {% endcomment %}
<a href="{{ product.url | within: collection }}">{{ product.title }}</a>
{% comment %} ✅ Produces /products/oxford-blue (the canonical URL) {% endcomment %}
<a href="{{ product.url }}">{{ product.title }}</a>Search your theme for within: collection — typically in snippets/product-card.liquid, card-product.liquid or product-grid-item.liquid — and remove the filter. Check quick-view, "recently viewed" and recommendation snippets too. Mixed grids like the one in the script output below come from a second snippet that wasn't updated.
What you lose: the breadcrumb on the product page can no longer infer which collection the visitor came from. If you need that, use a breadcrumb based on the product's primary collection (a metafield) rather than the URL.
Fix 2: Decide What Tag Pages Are For
Tag-filtered pages (/collections/shirts/linen) are created for every tag you use for filtering. 8 of the 10 we found were indexable and canonicalized to themselves. That's fine for a few deliberate landing pages ("linen shirts" may be a real search), and wasteful for the hundreds of combinations nobody searches.
{% comment %} In layout/theme.liquid <head>: noindex tag-filtered collection pages, keep links followable {% endcomment %}
{% if template contains 'collection' and current_tags %}
<meta name="robots" content="noindex, follow">
{% endif %}For the few tag pages that deserve to rank, give them a unique title, a description and some copy, or better, turn them into real collections (automated collections can use the same tag rule) so they get a clean URL.
Fix 3: Keep Pagination Self-Canonical
Google's ecommerce pagination guidance is for each page to carry its own canonical (?page=2 canonical to ?page=2), so products deep in the list remain discoverable. 83.6% of stores did this. The 6 stores canonicalizing page 2 to page 1 tell Google to ignore every product that only appears from page 2 onward. Our guide to pagination SEO without rel=prev/next covers the details.
{% comment %} ✅ Shopify's canonical_url object already includes ?page=N; don't override it {% endcomment %}
<link rel="canonical" href="{{ canonical_url }}">Fix 4: Don't Throw Away the Default robots.txt Rules
Since 2021 Shopify has let you edit robots.txt.liquid (see Shopify's robots.txt documentation). 7 of 63 stores' robots.txt no longer contained the default rules that disallow sort_by and +-combined tag URLs, usually because someone replaced the file instead of extending it. Add rules inside the default loop, so Shopify's defaults stay in place:
{% comment %} templates/robots.txt.liquid: extend the defaults, don't replace them {% endcomment %}
{% for group in robots.default_groups %}
{{- group.user_agent }}
{%- for rule in group.rules -%}
{{ rule }}
{%- endfor -%}
{%- if group.user_agent.value == '*' -%}
{{ 'Disallow: /collections/*?filter*' }}
{%- endif -%}
{%- if group.sitemap != blank -%}
{{ group.sitemap }}
{%- endif -%}
{% endfor %}What Shopify Won't Let You Change
Be clear with clients about the limits, so nobody spends a sprint fighting the platform:
- URL prefixes are fixed:
/products/,/collections/,/pages/,/blogs/<blog>/. You can't remove them or nest products under categories. - The collection-path route always resolves. You can stop linking to it, and it canonicalizes, but you can't delete it.
- Variant URLs exist for every option. Rely on the canonical, and don't try to block
?variant=in robots.txt, because that stops Google seeing the canonical. /collections/alland thevendors/typeslistings exist on every store. Link to them deliberately or not at all.
If URL structure really matters for the business (for example multi-level category paths), that's an architecture decision such as headless Shopify, not a theme fix.
Check a Store: Script
This script runs the same checks as our study on one store and one collection, pacing requests 1.5 seconds apart because Shopify rate-limits bursts. Standard library only.
#!/usr/bin/env python3
"""shopify_duplicates.py: check a Shopify store for the duplicate-URL patterns Shopify creates.
Usage:
python3 shopify_duplicates.py https://store.example.com [--collection /collections/shirts]
Standard library only; waits 1.5 s between requests (Shopify rate-limits bursts). Checks:
- collection grids linking to /collections/<c>/products/<p> instead of /products/<p>
- that the collection-path product URL canonicalises to /products/<p>
- ?variant= URLs canonicalising to the clean product URL
- tag-filtered collection pages (/collections/<c>/<tag>): canonical / noindex
- ?page=2 and ?sort_by= canonicals
- robots.txt still has Shopify's default sort/filter rules
"""
import argparse, json, re, ssl, sys, time, urllib.error, urllib.request
from urllib.parse import urljoin, urlsplit
UA = "Mozilla/5.0 (compatible; shopify-duplicates/1.0)"
CTX = ssl.create_default_context()
def get(url):
time.sleep(1.5)
req = urllib.request.Request(url, headers={"User-Agent": UA})
try:
with urllib.request.urlopen(req, timeout=20, context=CTX) as r:
return r.status, r.geturl(), r.read(3_000_000).decode("utf-8", "replace")
except urllib.error.HTTPError as e:
return e.code, url, ""
except Exception as e:
return type(e).__name__, url, ""
def canonical(base, html):
m = re.search(r"<link[^>]+rel=[\"']canonical[\"'][^>]*href=[\"']([^\"']+)", html, re.I) or \
re.search(r"<link[^>]+href=[\"']([^\"']+)[\"'][^>]*rel=[\"']canonical", html, re.I)
return urljoin(base, m.group(1)) if m else None
def path(u):
return urlsplit(u).path.rstrip("/") if u else None
def line(ok, label, detail):
print(f" {'✓' if ok else ('✗' if ok is False else '•')} {label:42s} {detail}")
return ok is not False
def main():
ap = argparse.ArgumentParser()
ap.add_argument("store")
ap.add_argument("--collection", help="collection path to test (auto-detected if omitted)")
a = ap.parse_args()
base = a.store.rstrip("/")
ok = True
st, home, html = get(base + "/")
links = {urljoin(home, h).split("?")[0] for h in re.findall(r"href=[\"']([^\"'#]+)", html)}
col = base + a.collection if a.collection else next(
(l for l in sorted(links) if re.search(r"/collections/[^/]+/?$", urlsplit(l).path) and "/collections/all" not in l),
base + "/collections/all")
print(f"\n{base} (collection: {path(col)})")
st, cf, ch = get(col)
grid = [urljoin(cf, h) for h in re.findall(r"href=[\"']([^\"'#]+)", ch)]
via = {l.split("?")[0] for l in grid if re.search(r"/collections/[^/]+/products/", l)}
direct = {l.split("?")[0] for l in grid if re.search(r"^(/[a-z]{2}(-[a-z]{2})?)?/products/", urlsplit(l).path)}
ok &= line(not via, "grid links to /products/<handle>",
f"{len(direct)} direct, {len(via)} via /collections/…/products/")
handle = next((re.search(r"/products/([^/?#]+)", l).group(1) for l in sorted(via | direct)), None)
if not handle:
print(" • no product links found on the collection page")
return 1
prod = f"{base}/products/{handle}"
st, pf, ph = get(prod)
ok &= line(path(canonical(pf, ph)) == path(pf), "product canonical is self", canonical(pf, ph) or "missing")
st, f2, h2 = get(f"{base}{path(col)}/products/{handle}")
c2 = canonical(f2, h2)
ok &= line(path(c2) == path(prod), "/collections/…/products/ canonicalises to /products/", f"HTTP {st}, canonical {c2}")
try:
variants = [v["id"] for v in json.loads(get(f"{prod}.js")[2]).get("variants", [])]
except ValueError:
variants = []
if len(variants) > 1:
st, f3, h3 = get(f"{prod}?variant={variants[1]}")
c3 = canonical(f3, h3) or ""
ok &= line(path(c3) == path(prod) and "variant=" not in c3, "?variant= canonicalises to product", c3 or "missing")
else:
line(None, "?variant= URLs", f"{len(variants)} variant(s): nothing to test")
tags = sorted({l for l in links | {urljoin(cf, h).split("?")[0] for h in re.findall(r"href=[\"']([^\"'#]+)", ch)}
if re.search(rf"{re.escape(path(col))}/[^/]+$", urlsplit(l).path) and "/products/" not in l})
if tags:
st, f4, h4 = get(tags[0])
c4 = canonical(f4, h4)
robots = re.search(r"<meta[^>]+name=[\"']robots[\"'][^>]*content=[\"']([^\"']+)", h4, re.I)
noidx = bool(robots and "noindex" in robots.group(1).lower())
line(noidx or path(c4) == path(col) or None, "tag page is consolidated or noindex",
f"{path(tags[0])}: canonical {path(c4)}{', noindex' if noidx else ''}")
else:
line(None, "tag-filtered collection pages", "none linked")
st, f5, h5 = get(col + "?page=2")
c5 = canonical(f5, h5) or ""
ok &= line("page=2" in c5, "?page=2 keeps its own canonical", c5 or "missing")
st, f6, h6 = get(col + "?sort_by=price-ascending")
c6 = canonical(f6, h6) or ""
ok &= line(bool(c6) and "sort_by" not in c6, "?sort_by= canonicalises to the clean URL", c6 or "missing")
st, _, rob = get(base + "/robots.txt")
ok &= line("/collections/*sort_by*" in rob, "robots.txt keeps Shopify's default rules", f"HTTP {st}")
return 0 if ok else 1
if __name__ == "__main__":
sys.exit(main())Real output for an apparel store from our sample (10 October 2026, domain replaced, long URLs shortened):
https://apparel-store.example (collection: /collections/Amaze-Puff-F26)
✗ grid links to /products/<handle> 43 direct, 20 via /collections/…/products/
✓ product canonical is self https://apparel-store.example/products/bot-n-hombre-peakfreak-ii-mid-outdry-…
✓ /collections/…/products/ canonicalises to /products/ HTTP 200, canonical https://apparel-store.example/products/bot-n-hombre-peakfreak-ii-mid-outdry-…
✓ ?variant= canonicalises to product https://apparel-store.example/products/bot-n-hombre-peakfreak-ii-mid-outdry-…
• tag-filtered collection pages none linked
✓ ?page=2 keeps its own canonical https://apparel-store.example/collections/amaze-puff-f26?page=2
✓ ?sort_by= canonicalises to the clean URL https://apparel-store.example/collections/amaze-puff-f26
✓ robots.txt keeps Shopify's default rules HTTP 200This store gets the hard parts right. Every canonical consolidates correctly, pagination is self-canonical and the robots.txt defaults are intact. But the collection grid mixes 43 direct product links with 20 collection-path links, which is the Fix 1 problem, probably from a second card snippet. It's a one-line Liquid change that removes thousands of crawlable duplicates.
How BugViso Finds Shopify Duplicates
A BugViso multi-page crawl reports what the theme actually produces:
- Duplicate content engine: exact and near-duplicate pages (64-bit SimHash over main content), plus duplicate titles, meta descriptions and H1s, with canonicalized variants consolidated so a correctly canonicalized variant isn't double-counted.
- Canonical analysis on every audited page: missing, relative or cross-URL canonicals, protocol mismatches, and subpages canonicalizing to the homepage.
- Crawl discovery starts from homepage links and then the sitemap tree, so collection-path URLs that your grids link to appear in the crawl inventory (and its CSV export) next to their
/products/twins. - Structured data validation for
ProductandOffer(requiredname,image,offers,price,priceCurrency), which Shopify themes output with very different completeness.
Crawls are capped by plan (75 to 300 pages), so on large catalogues crawl one collection at a time. You can crawl your Shopify store with BugViso. The checks are described in our duplicate content checker guide, and on the advanced SEO intelligence page.
Traps and Edge Cases
- Markets subfolders (
/en-gb/,/fr/) multiply every URL by the number of markets. Shopify adds hreflang automatically. Check it with our hreflang errors validator. - Apps that create pages (wishlists, filters, size guides) can add their own parameter URLs. Check the crawl for unknown parameters after installing an app.
- Product handles changed after launch create redirects automatically (Shopify's URL redirects). Keep internal links pointing at the new handle.
- Duplicate product listings (the same item listed twice to target two keywords) are true duplicates that no canonical fixes. Merge them.
- For a full store audit, combine this with our ecommerce website audit guide, and see how to speed up a Shopify store for the performance side.
FAQ
Does Shopify create duplicate content?
Yes. Every product is reachable through collection paths and variant parameters, and tag filters create extra collection URLs. Shopify's canonical tags consolidate most of it. The main remaining issue in our sample was themes linking to collection-path URLs (25% of grids).
Should I remove "within: collection" from my Shopify theme?
Yes, for product grids. It makes internal links point at non-canonical URLs. Use {{ product.url }} instead, and build breadcrumbs from a primary-collection metafield if you need them.
Should Shopify tag pages be indexed?
Only deliberate ones with unique copy. Add noindex, follow to tag-filtered collection pages by default, and turn tags people actually search for into real collections.
Can I change Shopify's /products/ and /collections/ URL structure?
No. The prefixes are fixed on standard Shopify. Focus on internal links, canonicals and indexing directives, which you do control.
Conclusion
Shopify's canonicals handle most duplicates for you, so spend your effort on what the theme controls: link grids to /products/, noindex tag combinations, keep pagination self-canonical and extend (don't replace) robots.txt. A BugViso crawl shows which of those your store still leaks.
See where your site stands
Run a free BugViso audit for SEO, speed, accessibility and AI search readiness — with fixes you can ship today.