Hreflang Errors: 7 Mistakes That Break International SEO

Hreflang errors found on 661 multilingual homepages: missing return links, en-UK codes, redirecting alternates and canonical conflicts, plus a validator.

BugViso

17 min read

Hreflang errors are mismatches between what your hreflang annotations claim and what the URLs actually do: codes that aren't valid ISO values (en-UK), alternates that redirect or return 404, pages that canonicalize somewhere else, missing self-references, and alternates that don't link back. Google ignores annotations it can't confirm, so one broken link in a cluster can quietly stop the right language version from showing in the right country.

We measured how common each mistake is. On 10 October 2026 we validated the hreflang clusters of 661 multilingual homepages — every site declaring hreflang in a random sample of 3,416 reachable homepages from the Tranco top 200,000 — fetching up to ten alternates per site. 39.9% had at least one error. The most common: an alternate that wasn't a direct 200 (16.6% of sites), a homepage that canonicalized to a different URL than the one carrying the annotations (13.6%), alternates that canonicalized elsewhere (10.4%), no self-reference (8.9%) and missing return links (8.3%).

This guide walks through the seven mistakes with their real frequencies and fixes, then gives you the validator we used.


How Google Reads Hreflang (the Rules That Matter)

Google's localized versions documentation sets the rules:

  • The value is an ISO 639-1 language code, optionally followed by an ISO 3166-1 alpha-2 region (en-GB, pt-BR), optionally with an ISO 15924 script (zh-Hant). A region on its own isn't allowed.
  • x-default is a reserved value for the fallback page.
  • Annotations must be bidirectional: "If two pages don't both point to each other, the tags will be ignored."
  • You can declare hreflang in HTML <link> tags, HTTP Link headers or the XML sitemap; the methods are equivalent, and using several "has no benefit".

Everything else follows from those rules plus canonicalization (Google's guide to consolidating duplicate URLs): an alternate URL should be the indexable, canonical, 200-status version of that page, because that's the URL Google would show.

html
<!-- ✅ A complete cluster on https://example.com/en-gb/pricing (repeated identically on every member) -->
<link rel="alternate" hreflang="en-gb" href="https://example.com/en-gb/pricing">
<link rel="alternate" hreflang="en-us" href="https://example.com/en-us/pricing">
<link rel="alternate" hreflang="de"    href="https://example.com/de/preise">
<link rel="alternate" hreflang="x-default" href="https://example.com/pricing">
<link rel="canonical" href="https://example.com/en-gb/pricing">

What We Found on 661 Multilingual Homepages

ItemDetail
SampleTranco 94GG2 ranks 1,001–200,000; 6,000 random domains (seed 20261010) → 3,416 reachable HTML homepages → 688 (20.1%) declare hreflang → 661 audited (rest errored)
MethodHTML <link> and HTTP Link header annotations on the homepage; up to 10 alternates fetched without following redirects; each alternate's status, robots meta, canonical and its own annotations checked
Median cluster size5 hreflang entries
ErrorSitesShare of 661
Alternate isn't a direct 200 (redirect, 4xx, 5xx, unreachable)11016.6%
…of which redirects (301/302/307/308)568.5%
…of which 4xx263.9%
Annotated page canonicalizes to a different URL9013.6%
An alternate canonicalizes elsewhere6910.4%
No self-referencing entry598.9%
Missing return link from an alternate558.3%
Invalid language/region code375.6%
Same code points to several URLs162.4%
Alternate is noindex50.8%
At least one error26439.9%
No x-default (recommendation, not an error)24837.5%

At the level of individual links the picture is better: of 2,548 alternates that returned 200, 95.4% linked back. Most clusters are mostly right — the errors concentrate in a handful of URLs per site (median 2 errors per affected site), which is exactly why they go unnoticed.

Two caveats. The random sample includes some gambling doorway networks, whose hreflang points across throwaway domains and inflates the self-reference and return-link failures. And a small share of canonical mismatches (5 of 90) differ only by a trailing slash — still a real inconsistency, because Google treats /en and /en/ as different URLs.


The 7 Mistakes and How to Fix Them

1. Alternates that redirect or 404

16.6% of sites listed at least one alternate that didn't answer 200 directly. The patterns we saw: alternates written with http:// that redirect to HTTPS; URLs with an explicit :443 port; old paths left after a locale was renamed; a development server URL on port 3000 published in production; internal CMS paths (/content/<site>/language-masters/en) leaking into the output; and alternates pointing at a store's internal myshopify.com hostname instead of its real domain.

html
<!-- ❌ Redirecting / internal alternates -->
<link rel="alternate" hreflang="en" href="http://www.example.com/en">
<link rel="alternate" hreflang="fr" href="https://www.example.com:443/fr/home.html">
<!-- ✅ The final, canonical URL every time -->
<link rel="alternate" hreflang="en" href="https://www.example.com/en/">
<link rel="alternate" hreflang="fr" href="https://www.example.com/fr/">

Generate hreflang from the same function that generates canonicals, so they can't disagree.

2. Canonical conflicts

13.6% of annotated homepages declared a canonical pointing to a different URL — typically the root canonicalizing to /en/ or /home, or HTTPS pages canonicalizing to http://. When a page canonicalizes away, Google consolidates on the canonical target and the annotations on the non-canonical URL don't count. Another 10.4% listed alternates that themselves canonicalized elsewhere. Rule: each URL in the cluster must be its own canonical, and must not be blocked with noindex (see Google's block indexing documentation for how the two signals interact).

3. Missing self-reference

8.9% of clusters didn't include the page itself. The self-reference isn't optional decoration: it's how the page states its own language. Render the full cluster identically on every member, including the current page.

8.3% of sites had an alternate that didn't link back. The usual causes: one locale on a different template or CMS, a newly added language not yet in the older pages' lists, or regional domains managed by different teams. Because Google ignores one-way annotations, the missing direction silently disables the pair.

5. Invalid codes (en-UK and friends)

5.6% of sites used at least one invalid code. The most frequent:

Code usedProblemUse instead
es-419 (7 sites)UN M.49 numeric region; Google requires ISO 3166-1 alpha-2One code per country (es-MX, es-AR, …) or plain es
fil, fil-phThree-letter code; Google's format is ISO 639-1tl / tl-PH, or check whether plain en-PH suits the page
uaRegion code used as a languageuk (Ukrainian) / uk-UA
czRegion code used as a languagecs / cs-CZ
in-idDeprecated code for Indonesianid / id-ID
ja-JA, en-JAJA isn't a countryja-JP, en-JP
en-UKUK isn't in ISO 3166-1en-GB

Case doesn't matter (en-gb = en-GB), but hyphens do: en_US is invalid. The Library of Congress maintains the ISO 639 language code list if you need to look a code up.

6. Duplicate codes

2.4% of sites assigned the same code to two different URLs — often x-default and en both pointing at different homepages, or two regional stores both claiming en. Each code should map to exactly one URL in the cluster.

7. Missing x-default (and mixed methods)

37.5% of clusters had no x-default. It's not required, but without it users whose language isn't listed get whichever version Google picks. Point it at your language selector or your main international version. Separately, 1.2% declared hreflang in both HTML and HTTP headers; that isn't wrong, but Google says it adds nothing, and two copies tend to drift apart.


Validate a Cluster: Script

This validator applies Google's rules to any URL: it reads HTML and HTTP-header hreflang, checks codes against the full ISO 639-1 and 3166-1 lists, and fetches each alternate without following redirects to check status, noindex, canonical and return links. It's the exact code we ran across the 661 sites.

python
#!/usr/bin/env python3
"""hreflang_check.py: validate a page's hreflang cluster the way Google reads it.

Usage:
    python3 hreflang_check.py https://example.com/en/pricing [--max 25]

Standard library only. Reads hreflang from the HTML <head> and the HTTP Link header, then checks:
  1. codes: ISO 639-1 language [+ ISO 15924 script] [+ ISO 3166-1 alpha-2 region], e.g. en-GB not en-UK
  2. a self-referencing entry          3. an x-default entry
  4. one URL per language code        5. every alternate answers 200 directly (no redirect, no 4xx)
  6. alternates are indexable (no noindex) and canonical to themselves
  7. return links: each alternate lists this page back
Exit code 1 if any error was found.
"""
import argparse, re, ssl, sys, urllib.error, urllib.request
from urllib.parse import urljoin, urlsplit

LANGS = set("aa ab ae af ak am an ar as av ay az ba be bg bi bm bn bo br bs ca ce ch co cr cs cu cv cy da de dv dz ee el en eo es et eu fa ff fi fj fo fr fy ga gd gl gn gu gv ha he hi ho hr ht hu hy hz ia id ie ig ii ik io is it iu ja jv ka kg ki kj kk kl km kn ko kr ks ku kv kw ky la lb lg li ln lo lt lu lv mg mh mi mk ml mn mr ms mt my na nb nd ne ng nl nn no nr nv ny oc oj om or os pa pi pl ps pt qu rm rn ro ru rw sa sc sd se sg sh si sk sl sm sn so sq sr ss st su sv sw ta te tg th ti tk tl tn to tr ts tt tw ty ug uk ur uz ve vi vo wa wo xh yi yo za zh zu".split())
REGIONS = set("AD AE AF AG AI AL AM AO AQ AR AS AT AU AW AX AZ BA BB BD BE BF BG BH BI BJ BL BM BN BO BQ BR BS BT BV BW BY BZ CA CC CD CF CG CH CI CK CL CM CN CO CR CU CV CW CX CY CZ DE DJ DK DM DO DZ EC EE EG EH ER ES ET FI FJ FK FM FO FR GA GB GD GE GF GG GH GI GL GM GN GP GQ GR GS GT GU GW GY HK HM HN HR HT HU ID IE IL IM IN IO IQ IR IS IT JE JM JO JP KE KG KH KI KM KN KP KR KW KY KZ LA LB LC LI LK LR LS LT LU LV LY MA MC MD ME MF MG MH MK ML MM MN MO MP MQ MR MS MT MU MV MW MX MY MZ NA NC NE NF NG NI NL NO NP NR NU NZ OM PA PE PF PG PH PK PL PM PN PR PS PT PW PY QA RE RO RS RU RW SA SB SC SD SE SG SH SI SJ SK SL SM SN SO SR SS ST SV SX SY SZ TC TD TF TG TH TJ TK TL TM TN TO TR TT TV TW TZ UA UG UM US UY UZ VA VC VE VG VI VN VU WF WS YE YT ZA ZM ZW".split())
UA = "Mozilla/5.0 (compatible; hreflang-check/1.0)"
LINK_TAG = re.compile(r"<link\b[^>]*>", re.I)
ATTR = lambda name: re.compile(r"\b" + name + r"\s*=\s*[\"']?([^\"'\s>]+)", re.I)
CTX = ssl.create_default_context()


class NoRedirect(urllib.request.HTTPRedirectHandler):
    def redirect_request(self, *a, **k):
        return None


def fetch(url):
    """Return (status, location, html, link_header) without following redirects."""
    opener = urllib.request.build_opener(NoRedirect, urllib.request.HTTPSHandler(context=CTX))
    try:
        with opener.open(urllib.request.Request(url, headers={"User-Agent": UA}), timeout=20) as r:
            return r.status, None, r.read(1_500_000).decode("utf-8", "replace"), r.headers.get("Link", "")
    except urllib.error.HTTPError as e:
        return e.code, e.headers.get("Location"), "", e.headers.get("Link", "") if e.headers else ""
    except Exception as e:
        return type(e).__name__, None, "", ""


def norm(u):
    p = urlsplit(u)
    return f"{p.scheme.lower()}://{(p.hostname or '').lower()}{p.path or '/'}" + (f"?{p.query}" if p.query else "")


def parse(base, html, link_header):
    head = html.split("</head>", 1)[0]
    tags, canonical, robots = [], None, ""
    for t in LINK_TAG.findall(head):
        rel = (ATTR("rel").search(t) or [None, ""])[1].lower()
        href = ATTR("href").search(t)
        if rel == "alternate" and ATTR("hreflang").search(t) and href:
            tags.append(("html", ATTR("hreflang").search(t).group(1), urljoin(base, href.group(1))))
        elif rel == "canonical" and href:
            canonical = urljoin(base, href.group(1))
    for m in re.finditer(r"<([^>]+)>\s*;[^,]*hreflang=\"?([^\";,]+)", link_header or ""):
        tags.append(("header", m.group(2), urljoin(base, m.group(1))))
    meta = re.search(r"<meta[^>]+name=[\"']robots[\"'][^>]*content=[\"']([^\"']+)", head, re.I)
    return tags, canonical, (meta.group(1).lower() if meta else "")


def code_problem(code):
    c = code.strip()
    if c.lower() == "x-default":
        return None
    if "_" in c:
        return f"'{c}' uses an underscore (use '{c.replace('_', '-')}')"
    parts = c.split("-")
    if parts[0].lower() not in LANGS:
        return f"'{c}': '{parts[0]}' is not an ISO 639-1 language code"
    rest = parts[1:]
    if rest and len(rest[0]) == 4 and rest[0].isalpha():   # ISO 15924 script, e.g. zh-Hant
        rest = rest[1:]
    if len(rest) > 1 or (rest and not (len(rest[0]) == 2 and rest[0].isalpha())):
        return f"'{c}' is not language[-Script][-REGION] (Google doesn't accept numeric regions like 419)"
    if rest and rest[0].upper() not in REGIONS:
        hint = " (the UK is 'GB')" if rest[0].upper() == "UK" else ""
        return f"'{c}': '{rest[0]}' is not an ISO 3166-1 region{hint}"
    return None


def audit(url, max_alt=25):
    res = {"url": url, "errors": [], "warnings": [], "alternates": []}
    status, loc, html, link = fetch(url)
    if status != 200:
        res["errors"].append(f"page itself returned {status}" + (f" → {loc}" if loc else ""))
        return res
    tags, canonical, robots = parse(url, html, link)
    res["tags"] = len(tags)
    if not tags:
        res["warnings"].append("no hreflang annotations found in <head> or Link header (sitemap annotations aren't checked)")
        return res
    methods = {m for m, _, _ in tags}
    if len(methods) > 1:
        res["warnings"].append("hreflang declared in both HTML and the HTTP header: Google says there's no benefit, and two copies can drift apart")
    for _, code, _ in tags:
        p = code_problem(code)
        if p:
            res["errors"].append(f"invalid code {p}")
    me = {norm(url)} | ({norm(canonical)} if canonical else set())
    if canonical and norm(canonical) != norm(url):
        res["errors"].append(f"page canonicalises to {canonical}, so its hreflang cluster is ignored on this URL")
    if not any(norm(h) in me for _, _, h in tags):
        res["errors"].append("no self-referencing hreflang entry")
    if not any(c.lower() == "x-default" for _, c, _ in tags):
        res["warnings"].append("no x-default (recommended fallback for unmatched languages)")
    by_code = {}
    for _, c, h in tags:
        by_code.setdefault(c.lower(), set()).add(norm(h))
    for c, urls in by_code.items():
        if len(urls) > 1:
            res["errors"].append(f"'{c}' points to {len(urls)} different URLs")
    targets = sorted({h for _, _, h in tags if norm(h) not in me})[:max_alt]
    for t in targets:
        st, loc, ahtml, alink = fetch(t)
        row = {"url": t, "status": st}
        if st != 200:
            res["errors"].append(f"alternate {t} returned {st}" + (f" → {loc}" if loc else "") + " (must be a direct 200)")
        else:
            atags, acanon, arobots = parse(t, ahtml, alink)
            if "noindex" in arobots:
                res["errors"].append(f"alternate {t} is noindex")
            if acanon and norm(acanon) != norm(t):
                res["errors"].append(f"alternate {t} canonicalises elsewhere ({acanon})")
            if not any(norm(h) in me for _, _, h in atags):
                res["errors"].append(f"no return link: {t} does not list this page")
                row["return"] = False
            else:
                row["return"] = True
        res["alternates"].append(row)
    return res


if __name__ == "__main__":
    ap = argparse.ArgumentParser()
    ap.add_argument("url")
    ap.add_argument("--max", type=int, default=25, help="max alternates to fetch")
    a = ap.parse_args()
    r = audit(a.url, a.max)
    print(f"{r['url']}: {r.get('tags', 0)} hreflang entries, {len(r['alternates'])} alternates fetched")
    for e in r["errors"]:
        print(f"  ✗ {e}")
    for w in r["warnings"]:
        print(f"  ⚠ {w}")
    if not r["errors"] and not r["warnings"]:
        print("  ✓ no problems found")
    sys.exit(1 if r["errors"] else 0)

Real output for a corporate group's homepage (10 October 2026, domain replaced):

text
https://www.corporate-group.example/: 2 hreflang entries, 2 alternates fetched
  ✗ no self-referencing hreflang entry
  ✗ alternate https://www.corporate-group.example:443/en canonicalises elsewhere (https://www.corporate-group.example)
  ✗ no return link: https://www.corporate-group.example:443/en does not list this page
  ✗ alternate https://www.corporate-group.example:443/home.html returned 301 → https://www.corporate-group.example (must be a direct 200)
  ⚠ no x-default (recommended fallback for unmatched languages)

Two annotations, four errors: the URLs carry an explicit :443 port, one redirects, the English alternate canonicalizes back to the root, and the root doesn't list itself. Each fix is a template change. For large sites, run the script over a list of templates (homepage, category, product, article) rather than every URL — hreflang errors are almost always template-level.


How BugViso Checks Hreflang

BugViso's Advanced SEO Intelligence engine reads the hreflang annotations on each audited page and flags malformed language codes (underscores and values that don't fit the language-region pattern) and a missing self-referencing entry. During a multi-page crawl it builds the hreflang map of every crawled page and verifies reciprocity, reporting annotations whose return tag is missing on the target page. Canonical analysis on the same pages catches protocol mismatches, canonicals pointing at the other www/apex host and pages canonicalizing to the homepage — the conflicts behind mistake #2.

Two things to know. The code check validates the shape of a value, so a well-formed but non-existent region such as en-UK passes; use the script above for ISO validity. And reciprocity is only checked between pages the crawl reached, so set the crawl to include each locale's section. You can run a free hreflang check with a BugViso scan; the checks are listed on the advanced SEO intelligence page. For the canonical side, see canonical tags and duplicate content.


Validation Checklist

  1. Every code is ISO 639-1 [+ script] [+ ISO 3166-1 alpha-2 region]; no underscores, no UK, no numeric regions.
  2. Every page lists itself.
  3. Every listed URL returns 200 directly — no redirects, no 4xx, no internal hostnames or ports.
  4. Every listed URL is indexable and canonical to itself.
  5. Every alternate lists the original page back.
  6. One URL per code; an x-default exists.
  7. One declaration method (HTML, header or sitemap) generated from the same source as canonicals.
  8. Re-check after every locale launch, CMS migration or domain change — the moments when clusters break. Our website migration SEO checklist covers those launches.

FAQ

What are the most common hreflang errors?

In our sample of 661 multilingual homepages: alternates that redirect or return errors (16.6% of sites), annotated pages that canonicalize elsewhere (13.6%), alternates that canonicalize elsewhere (10.4%), missing self-references (8.9%) and missing return links (8.3%).

Is en-UK a valid hreflang code?

No. The United Kingdom's ISO 3166-1 code is GB, so use en-GB. Google requires ISO 3166-1 alpha-2 region codes.

Do I need x-default?

It's recommended, not required. Without it, visitors whose language you don't target get whichever version Google chooses. In our data 37.5% of clusters had none.

Should hreflang URLs be canonical?

Yes. Each URL in a cluster should return 200, be indexable and declare itself as canonical. A canonical pointing elsewhere tells Google the annotated URL isn't the one to index.

HTML tags, HTTP headers or sitemap — which is best?

They're equivalent for Google. Use the one you can generate most reliably from the same source as your canonicals; HTML tags for normal pages, the sitemap for very large clusters, headers for non-HTML files like PDFs.


Conclusion

Hreflang fails quietly: four in ten multilingual homepages in our sample had at least one error, usually a redirecting alternate or a canonical conflict, so generate hreflang from the same source as your canonicals, validate clusters per template, and let a BugViso crawl check reciprocity across your locales.

Found this useful? Share it.

See where your site stands

Run a free BugViso audit for SEO, speed, accessibility and AI search readiness — with fixes you can ship today.