Hreflang Errors: 7 Mistakes That Break International SEO
Hreflang errors found on 661 multilingual homepages: missing return links, en-UK codes, redirecting alternates and canonical conflicts, plus a validator.
Hreflang errors are mismatches between what your hreflang annotations claim and what the URLs actually do: codes that aren't valid ISO values (en-UK), alternates that redirect or return 404, pages that canonicalize somewhere else, missing self-references, and alternates that don't link back. Google ignores annotations it can't confirm, so one broken link in a cluster can quietly stop the right language version from showing in the right country.
We measured how common each mistake is. On 10 October 2026 we validated the hreflang clusters of 661 multilingual homepages — every site declaring hreflang in a random sample of 3,416 reachable homepages from the Tranco top 200,000 — fetching up to ten alternates per site. 39.9% had at least one error. The most common: an alternate that wasn't a direct 200 (16.6% of sites), a homepage that canonicalized to a different URL than the one carrying the annotations (13.6%), alternates that canonicalized elsewhere (10.4%), no self-reference (8.9%) and missing return links (8.3%).
This guide walks through the seven mistakes with their real frequencies and fixes, then gives you the validator we used.
How Google Reads Hreflang (the Rules That Matter)
Google's localized versions documentation sets the rules:
- The value is an ISO 639-1 language code, optionally followed by an ISO 3166-1 alpha-2 region (
en-GB,pt-BR), optionally with an ISO 15924 script (zh-Hant). A region on its own isn't allowed. x-defaultis a reserved value for the fallback page.- Annotations must be bidirectional: "If two pages don't both point to each other, the tags will be ignored."
- You can declare hreflang in HTML
<link>tags, HTTPLinkheaders or the XML sitemap; the methods are equivalent, and using several "has no benefit".
Everything else follows from those rules plus canonicalization (Google's guide to consolidating duplicate URLs): an alternate URL should be the indexable, canonical, 200-status version of that page, because that's the URL Google would show.
<!-- ✅ A complete cluster on https://example.com/en-gb/pricing (repeated identically on every member) -->
<link rel="alternate" hreflang="en-gb" href="https://example.com/en-gb/pricing">
<link rel="alternate" hreflang="en-us" href="https://example.com/en-us/pricing">
<link rel="alternate" hreflang="de" href="https://example.com/de/preise">
<link rel="alternate" hreflang="x-default" href="https://example.com/pricing">
<link rel="canonical" href="https://example.com/en-gb/pricing">What We Found on 661 Multilingual Homepages
| Item | Detail |
|---|---|
| Sample | Tranco 94GG2 ranks 1,001–200,000; 6,000 random domains (seed 20261010) → 3,416 reachable HTML homepages → 688 (20.1%) declare hreflang → 661 audited (rest errored) |
| Method | HTML <link> and HTTP Link header annotations on the homepage; up to 10 alternates fetched without following redirects; each alternate's status, robots meta, canonical and its own annotations checked |
| Median cluster size | 5 hreflang entries |
| Error | Sites | Share of 661 |
|---|---|---|
| Alternate isn't a direct 200 (redirect, 4xx, 5xx, unreachable) | 110 | 16.6% |
| …of which redirects (301/302/307/308) | 56 | 8.5% |
| …of which 4xx | 26 | 3.9% |
| Annotated page canonicalizes to a different URL | 90 | 13.6% |
| An alternate canonicalizes elsewhere | 69 | 10.4% |
| No self-referencing entry | 59 | 8.9% |
| Missing return link from an alternate | 55 | 8.3% |
| Invalid language/region code | 37 | 5.6% |
| Same code points to several URLs | 16 | 2.4% |
Alternate is noindex | 5 | 0.8% |
| At least one error | 264 | 39.9% |
No x-default (recommendation, not an error) | 248 | 37.5% |
At the level of individual links the picture is better: of 2,548 alternates that returned 200, 95.4% linked back. Most clusters are mostly right — the errors concentrate in a handful of URLs per site (median 2 errors per affected site), which is exactly why they go unnoticed.
Two caveats. The random sample includes some gambling doorway networks, whose hreflang points across throwaway domains and inflates the self-reference and return-link failures. And a small share of canonical mismatches (5 of 90) differ only by a trailing slash — still a real inconsistency, because Google treats /en and /en/ as different URLs.
The 7 Mistakes and How to Fix Them
1. Alternates that redirect or 404
16.6% of sites listed at least one alternate that didn't answer 200 directly. The patterns we saw: alternates written with http:// that redirect to HTTPS; URLs with an explicit :443 port; old paths left after a locale was renamed; a development server URL on port 3000 published in production; internal CMS paths (/content/<site>/language-masters/en) leaking into the output; and alternates pointing at a store's internal myshopify.com hostname instead of its real domain.
<!-- ❌ Redirecting / internal alternates -->
<link rel="alternate" hreflang="en" href="http://www.example.com/en">
<link rel="alternate" hreflang="fr" href="https://www.example.com:443/fr/home.html">
<!-- ✅ The final, canonical URL every time -->
<link rel="alternate" hreflang="en" href="https://www.example.com/en/">
<link rel="alternate" hreflang="fr" href="https://www.example.com/fr/">Generate hreflang from the same function that generates canonicals, so they can't disagree.
2. Canonical conflicts
13.6% of annotated homepages declared a canonical pointing to a different URL — typically the root canonicalizing to /en/ or /home, or HTTPS pages canonicalizing to http://. When a page canonicalizes away, Google consolidates on the canonical target and the annotations on the non-canonical URL don't count. Another 10.4% listed alternates that themselves canonicalized elsewhere. Rule: each URL in the cluster must be its own canonical, and must not be blocked with noindex (see Google's block indexing documentation for how the two signals interact).
3. Missing self-reference
8.9% of clusters didn't include the page itself. The self-reference isn't optional decoration: it's how the page states its own language. Render the full cluster identically on every member, including the current page.
4. Missing return links
8.3% of sites had an alternate that didn't link back. The usual causes: one locale on a different template or CMS, a newly added language not yet in the older pages' lists, or regional domains managed by different teams. Because Google ignores one-way annotations, the missing direction silently disables the pair.
5. Invalid codes (en-UK and friends)
5.6% of sites used at least one invalid code. The most frequent:
| Code used | Problem | Use instead |
|---|---|---|
es-419 (7 sites) | UN M.49 numeric region; Google requires ISO 3166-1 alpha-2 | One code per country (es-MX, es-AR, …) or plain es |
fil, fil-ph | Three-letter code; Google's format is ISO 639-1 | tl / tl-PH, or check whether plain en-PH suits the page |
ua | Region code used as a language | uk (Ukrainian) / uk-UA |
cz | Region code used as a language | cs / cs-CZ |
in-id | Deprecated code for Indonesian | id / id-ID |
ja-JA, en-JA | JA isn't a country | ja-JP, en-JP |
en-UK | UK isn't in ISO 3166-1 | en-GB |
Case doesn't matter (en-gb = en-GB), but hyphens do: en_US is invalid. The Library of Congress maintains the ISO 639 language code list if you need to look a code up.
6. Duplicate codes
2.4% of sites assigned the same code to two different URLs — often x-default and en both pointing at different homepages, or two regional stores both claiming en. Each code should map to exactly one URL in the cluster.
7. Missing x-default (and mixed methods)
37.5% of clusters had no x-default. It's not required, but without it users whose language isn't listed get whichever version Google picks. Point it at your language selector or your main international version. Separately, 1.2% declared hreflang in both HTML and HTTP headers; that isn't wrong, but Google says it adds nothing, and two copies tend to drift apart.
Validate a Cluster: Script
This validator applies Google's rules to any URL: it reads HTML and HTTP-header hreflang, checks codes against the full ISO 639-1 and 3166-1 lists, and fetches each alternate without following redirects to check status, noindex, canonical and return links. It's the exact code we ran across the 661 sites.
#!/usr/bin/env python3
"""hreflang_check.py: validate a page's hreflang cluster the way Google reads it.
Usage:
python3 hreflang_check.py https://example.com/en/pricing [--max 25]
Standard library only. Reads hreflang from the HTML <head> and the HTTP Link header, then checks:
1. codes: ISO 639-1 language [+ ISO 15924 script] [+ ISO 3166-1 alpha-2 region], e.g. en-GB not en-UK
2. a self-referencing entry 3. an x-default entry
4. one URL per language code 5. every alternate answers 200 directly (no redirect, no 4xx)
6. alternates are indexable (no noindex) and canonical to themselves
7. return links: each alternate lists this page back
Exit code 1 if any error was found.
"""
import argparse, re, ssl, sys, urllib.error, urllib.request
from urllib.parse import urljoin, urlsplit
LANGS = set("aa ab ae af ak am an ar as av ay az ba be bg bi bm bn bo br bs ca ce ch co cr cs cu cv cy da de dv dz ee el en eo es et eu fa ff fi fj fo fr fy ga gd gl gn gu gv ha he hi ho hr ht hu hy hz ia id ie ig ii ik io is it iu ja jv ka kg ki kj kk kl km kn ko kr ks ku kv kw ky la lb lg li ln lo lt lu lv mg mh mi mk ml mn mr ms mt my na nb nd ne ng nl nn no nr nv ny oc oj om or os pa pi pl ps pt qu rm rn ro ru rw sa sc sd se sg sh si sk sl sm sn so sq sr ss st su sv sw ta te tg th ti tk tl tn to tr ts tt tw ty ug uk ur uz ve vi vo wa wo xh yi yo za zh zu".split())
REGIONS = set("AD AE AF AG AI AL AM AO AQ AR AS AT AU AW AX AZ BA BB BD BE BF BG BH BI BJ BL BM BN BO BQ BR BS BT BV BW BY BZ CA CC CD CF CG CH CI CK CL CM CN CO CR CU CV CW CX CY CZ DE DJ DK DM DO DZ EC EE EG EH ER ES ET FI FJ FK FM FO FR GA GB GD GE GF GG GH GI GL GM GN GP GQ GR GS GT GU GW GY HK HM HN HR HT HU ID IE IL IM IN IO IQ IR IS IT JE JM JO JP KE KG KH KI KM KN KP KR KW KY KZ LA LB LC LI LK LR LS LT LU LV LY MA MC MD ME MF MG MH MK ML MM MN MO MP MQ MR MS MT MU MV MW MX MY MZ NA NC NE NF NG NI NL NO NP NR NU NZ OM PA PE PF PG PH PK PL PM PN PR PS PT PW PY QA RE RO RS RU RW SA SB SC SD SE SG SH SI SJ SK SL SM SN SO SR SS ST SV SX SY SZ TC TD TF TG TH TJ TK TL TM TN TO TR TT TV TW TZ UA UG UM US UY UZ VA VC VE VG VI VN VU WF WS YE YT ZA ZM ZW".split())
UA = "Mozilla/5.0 (compatible; hreflang-check/1.0)"
LINK_TAG = re.compile(r"<link\b[^>]*>", re.I)
ATTR = lambda name: re.compile(r"\b" + name + r"\s*=\s*[\"']?([^\"'\s>]+)", re.I)
CTX = ssl.create_default_context()
class NoRedirect(urllib.request.HTTPRedirectHandler):
def redirect_request(self, *a, **k):
return None
def fetch(url):
"""Return (status, location, html, link_header) without following redirects."""
opener = urllib.request.build_opener(NoRedirect, urllib.request.HTTPSHandler(context=CTX))
try:
with opener.open(urllib.request.Request(url, headers={"User-Agent": UA}), timeout=20) as r:
return r.status, None, r.read(1_500_000).decode("utf-8", "replace"), r.headers.get("Link", "")
except urllib.error.HTTPError as e:
return e.code, e.headers.get("Location"), "", e.headers.get("Link", "") if e.headers else ""
except Exception as e:
return type(e).__name__, None, "", ""
def norm(u):
p = urlsplit(u)
return f"{p.scheme.lower()}://{(p.hostname or '').lower()}{p.path or '/'}" + (f"?{p.query}" if p.query else "")
def parse(base, html, link_header):
head = html.split("</head>", 1)[0]
tags, canonical, robots = [], None, ""
for t in LINK_TAG.findall(head):
rel = (ATTR("rel").search(t) or [None, ""])[1].lower()
href = ATTR("href").search(t)
if rel == "alternate" and ATTR("hreflang").search(t) and href:
tags.append(("html", ATTR("hreflang").search(t).group(1), urljoin(base, href.group(1))))
elif rel == "canonical" and href:
canonical = urljoin(base, href.group(1))
for m in re.finditer(r"<([^>]+)>\s*;[^,]*hreflang=\"?([^\";,]+)", link_header or ""):
tags.append(("header", m.group(2), urljoin(base, m.group(1))))
meta = re.search(r"<meta[^>]+name=[\"']robots[\"'][^>]*content=[\"']([^\"']+)", head, re.I)
return tags, canonical, (meta.group(1).lower() if meta else "")
def code_problem(code):
c = code.strip()
if c.lower() == "x-default":
return None
if "_" in c:
return f"'{c}' uses an underscore (use '{c.replace('_', '-')}')"
parts = c.split("-")
if parts[0].lower() not in LANGS:
return f"'{c}': '{parts[0]}' is not an ISO 639-1 language code"
rest = parts[1:]
if rest and len(rest[0]) == 4 and rest[0].isalpha(): # ISO 15924 script, e.g. zh-Hant
rest = rest[1:]
if len(rest) > 1 or (rest and not (len(rest[0]) == 2 and rest[0].isalpha())):
return f"'{c}' is not language[-Script][-REGION] (Google doesn't accept numeric regions like 419)"
if rest and rest[0].upper() not in REGIONS:
hint = " (the UK is 'GB')" if rest[0].upper() == "UK" else ""
return f"'{c}': '{rest[0]}' is not an ISO 3166-1 region{hint}"
return None
def audit(url, max_alt=25):
res = {"url": url, "errors": [], "warnings": [], "alternates": []}
status, loc, html, link = fetch(url)
if status != 200:
res["errors"].append(f"page itself returned {status}" + (f" → {loc}" if loc else ""))
return res
tags, canonical, robots = parse(url, html, link)
res["tags"] = len(tags)
if not tags:
res["warnings"].append("no hreflang annotations found in <head> or Link header (sitemap annotations aren't checked)")
return res
methods = {m for m, _, _ in tags}
if len(methods) > 1:
res["warnings"].append("hreflang declared in both HTML and the HTTP header: Google says there's no benefit, and two copies can drift apart")
for _, code, _ in tags:
p = code_problem(code)
if p:
res["errors"].append(f"invalid code {p}")
me = {norm(url)} | ({norm(canonical)} if canonical else set())
if canonical and norm(canonical) != norm(url):
res["errors"].append(f"page canonicalises to {canonical}, so its hreflang cluster is ignored on this URL")
if not any(norm(h) in me for _, _, h in tags):
res["errors"].append("no self-referencing hreflang entry")
if not any(c.lower() == "x-default" for _, c, _ in tags):
res["warnings"].append("no x-default (recommended fallback for unmatched languages)")
by_code = {}
for _, c, h in tags:
by_code.setdefault(c.lower(), set()).add(norm(h))
for c, urls in by_code.items():
if len(urls) > 1:
res["errors"].append(f"'{c}' points to {len(urls)} different URLs")
targets = sorted({h for _, _, h in tags if norm(h) not in me})[:max_alt]
for t in targets:
st, loc, ahtml, alink = fetch(t)
row = {"url": t, "status": st}
if st != 200:
res["errors"].append(f"alternate {t} returned {st}" + (f" → {loc}" if loc else "") + " (must be a direct 200)")
else:
atags, acanon, arobots = parse(t, ahtml, alink)
if "noindex" in arobots:
res["errors"].append(f"alternate {t} is noindex")
if acanon and norm(acanon) != norm(t):
res["errors"].append(f"alternate {t} canonicalises elsewhere ({acanon})")
if not any(norm(h) in me for _, _, h in atags):
res["errors"].append(f"no return link: {t} does not list this page")
row["return"] = False
else:
row["return"] = True
res["alternates"].append(row)
return res
if __name__ == "__main__":
ap = argparse.ArgumentParser()
ap.add_argument("url")
ap.add_argument("--max", type=int, default=25, help="max alternates to fetch")
a = ap.parse_args()
r = audit(a.url, a.max)
print(f"{r['url']}: {r.get('tags', 0)} hreflang entries, {len(r['alternates'])} alternates fetched")
for e in r["errors"]:
print(f" ✗ {e}")
for w in r["warnings"]:
print(f" ⚠ {w}")
if not r["errors"] and not r["warnings"]:
print(" ✓ no problems found")
sys.exit(1 if r["errors"] else 0)Real output for a corporate group's homepage (10 October 2026, domain replaced):
https://www.corporate-group.example/: 2 hreflang entries, 2 alternates fetched
✗ no self-referencing hreflang entry
✗ alternate https://www.corporate-group.example:443/en canonicalises elsewhere (https://www.corporate-group.example)
✗ no return link: https://www.corporate-group.example:443/en does not list this page
✗ alternate https://www.corporate-group.example:443/home.html returned 301 → https://www.corporate-group.example (must be a direct 200)
⚠ no x-default (recommended fallback for unmatched languages)Two annotations, four errors: the URLs carry an explicit :443 port, one redirects, the English alternate canonicalizes back to the root, and the root doesn't list itself. Each fix is a template change. For large sites, run the script over a list of templates (homepage, category, product, article) rather than every URL — hreflang errors are almost always template-level.
How BugViso Checks Hreflang
BugViso's Advanced SEO Intelligence engine reads the hreflang annotations on each audited page and flags malformed language codes (underscores and values that don't fit the language-region pattern) and a missing self-referencing entry. During a multi-page crawl it builds the hreflang map of every crawled page and verifies reciprocity, reporting annotations whose return tag is missing on the target page. Canonical analysis on the same pages catches protocol mismatches, canonicals pointing at the other www/apex host and pages canonicalizing to the homepage — the conflicts behind mistake #2.
Two things to know. The code check validates the shape of a value, so a well-formed but non-existent region such as en-UK passes; use the script above for ISO validity. And reciprocity is only checked between pages the crawl reached, so set the crawl to include each locale's section. You can run a free hreflang check with a BugViso scan; the checks are listed on the advanced SEO intelligence page. For the canonical side, see canonical tags and duplicate content.
Validation Checklist
- Every code is ISO 639-1 [+ script] [+ ISO 3166-1 alpha-2 region]; no underscores, no
UK, no numeric regions. - Every page lists itself.
- Every listed URL returns 200 directly — no redirects, no 4xx, no internal hostnames or ports.
- Every listed URL is indexable and canonical to itself.
- Every alternate lists the original page back.
- One URL per code; an
x-defaultexists. - One declaration method (HTML, header or sitemap) generated from the same source as canonicals.
- Re-check after every locale launch, CMS migration or domain change — the moments when clusters break. Our website migration SEO checklist covers those launches.
FAQ
What are the most common hreflang errors?
In our sample of 661 multilingual homepages: alternates that redirect or return errors (16.6% of sites), annotated pages that canonicalize elsewhere (13.6%), alternates that canonicalize elsewhere (10.4%), missing self-references (8.9%) and missing return links (8.3%).
Is en-UK a valid hreflang code?
No. The United Kingdom's ISO 3166-1 code is GB, so use en-GB. Google requires ISO 3166-1 alpha-2 region codes.
Do I need x-default?
It's recommended, not required. Without it, visitors whose language you don't target get whichever version Google chooses. In our data 37.5% of clusters had none.
Should hreflang URLs be canonical?
Yes. Each URL in a cluster should return 200, be indexable and declare itself as canonical. A canonical pointing elsewhere tells Google the annotated URL isn't the one to index.
HTML tags, HTTP headers or sitemap — which is best?
They're equivalent for Google. Use the one you can generate most reliably from the same source as your canonicals; HTML tags for normal pages, the sitemap for very large clusters, headers for non-HTML files like PDFs.
Conclusion
Hreflang fails quietly: four in ten multilingual homepages in our sample had at least one error, usually a redirecting alternate or a canonical conflict, so generate hreflang from the same source as your canonicals, validate clusters per template, and let a BugViso crawl check reciprocity across your locales.
See where your site stands
Run a free BugViso audit for SEO, speed, accessibility and AI search readiness — with fixes you can ship today.