How to Fix Broken Internal Links (Data From 219 Sites)
How to fix broken internal links: find every 404 and redirecting href, update the source link instead of stacking redirects, and verify. Data from 219 sites.
To fix broken internal links, crawl your site, list every internal <a href> that returns a 4xx or 5xx status, and edit the link at its source so it points to the correct live URL. Add a 301 redirect only for URLs that other websites link to. Then re-crawl to confirm that every internal link returns a direct 200, with no redirect hop in between.
That last step is the one most teams skip. On 3 October 2026 we checked 17,435 unique internal links on 219 homepages drawn at random from the Tranco top-sites list. Hard 404s were rare. Internal links that redirect were everywhere: more than half of the homepages pointed their own visitors and crawlers at URLs that bounce somewhere else.
This guide covers the full repair loop: how a broken internal link damages crawling, how to reproduce one with curl, the six root causes we keep finding, a copy-paste Python crawler that maps each broken target back to the page that links to it, and the edge cases that make a checker report false positives.
What Broken Internal Links Actually Cost You
A broken internal link is an <a href> on your own domain whose target returns an error, usually 404 Not Found or 410 Gone, sometimes 500/502 from a crashed backend. Each one wastes a crawl request, passes no link equity, and leaves a visitor on a dead end that you created yourself.
Google discovers pages by following links, as its documentation on making links crawlable explains. A broken internal link is therefore a broken edge in your own discovery graph. If the dead URL was the only path to a page, that page becomes an orphan. You can read about why orphans fall out of the index in our guide to finding and fixing orphan pages.
Internal redirects are the quieter version of the same problem. A link to /pricing that 301s to /pricing/ still works for a human. But every crawler request now costs two round trips, and the redirect becomes permanent technical debt the moment someone changes the destination again.
๐ก Rule of thumb: You control both ends of an internal link. Any internal link that doesn't return a direct 200 is a bug in your own HTML, not a "redirect strategy".
What We Found Checking 17,435 Internal Links
We drew a random sample of 420 domains from Tranco list 94GG2 (ranks 1,001โ50,000), loaded each homepage in headless Chromium, and kept the 219 that served a real page. Bot challenges, DNS failures and empty responses were excluded. Every unique same-site link on each homepage was then requested once: HEAD first, GET as a fallback.
| Metric (219 homepages) | Result |
|---|---|
| Unique internal links checked | 17,435 (median 74 per homepage) |
| Homepages linking to โฅ1 broken internal URL | 38 (17.4%) |
| Internal links that were broken (4xx/5xx) | 139 (0.8%) |
| Share of broken links that were plain 404 | 81% (112 of 139) |
| Homepages with โฅ1 internal link that redirects | 124 (56.6%) |
| Internal links that redirect before resolving | 1,200 (6.9%) |
| Redirected internal links passing through 2+ hops | 189 (15.8%) of redirects |
| Internal links that redirect into an error page | 8 |
Homepages using href="#" placeholders | 88 (40.2%) |
Homepages using javascript: hrefs | 39 (17.8%) |
| Links we could not verify (401/403/429/503, timeouts) | 2,243 (12.9%) |
Three findings matter more than the headline 404 rate.
Redirecting links outnumber broken ones about 9 to 1. Most of them come from trailing-slash mismatches, http:// links left over from an HTTPS migration, and old slugs that were redirected instead of updated.
One in eight links couldn't be verified at all. Bot protection answered our checker with 403, 429 or 503. A checker that counts these as "broken" floods you with false positives, which is why we report them separately.
Broken links cluster. The median affected homepage had exactly one broken link, but one site had 38. That pattern usually means a template or menu is generating them, and a single fix in the template clears them all.
โ ๏ธ Limit of this data: we checked links found on homepages only, from one location, without logging in. Deeper pages (blog archives, old product pages) typically have more broken links than homepages, so treat these numbers as a floor.
How to Reproduce and Diagnose a Broken Internal Link
Before you change anything, confirm the status yourself. Browsers hide redirects and soft errors, so test with curl.
# Does the target return a direct 200, a redirect, or an error?
curl -sI https://example.com/pricing | head -n 1
HTTP/2 301
# Follow it and print every hop with its status code
curl -sIL https://example.com/pricing | grep -iE "^(HTTP|location)"
HTTP/2 301
location: https://example.com/pricing/
HTTP/2 200A 301 โ 200 like this is not broken, but it is a link worth updating. A response that ends in 404 or 410 is broken. A 200 that shows "Page not found" in the body is a soft 404; our breakdown of soft 404s versus real 404s shows how Google classifies those.
To find which page links to the dead URL, search the rendered DOM, not the source. JavaScript frameworks often build links after load:
// Paste into the DevTools console on the page you suspect
[...document.querySelectorAll('a[href]')]
.map(a => ({ text: a.innerText.trim().slice(0, 40), href: a.href }))
.filter(l => l.href.includes('/old-slug'))If the link isn't on the page you expected, it's probably in a shared component: the header, footer, mega-menu or a "related posts" widget. That is how one site in our sample ended up with 38 broken links from a single template.
The 6 Root Causes (and the Fix for Each)
1. A slug changed and links were never updated
The page moved from /blog/seo-tips-2024 to /blog/seo-tips. Someone added a redirect, and dozens of internal links still point at the old URL.
<!-- โ Broken habit: internal link relies on a redirect forever -->
<a href="/blog/seo-tips-2024">Read our SEO tips</a>
<!-- โ
Fixed: link straight to the canonical URL -->
<a href="/blog/seo-tips">Read our SEO tips</a>Keep the 301 for backlinks and bookmarks, but update every link you control. Our guide to flattening redirect chains and loops covers the server side of that clean-up.
2. Trailing-slash and case mismatches
Running our own checker on a large open-source CMS site turned up eleven internal links that redirected only because they lacked a trailing slash:
== REDIRECTED (update the href): 11
301 -> 200 (1 hop) https://example.org/documentation
linked from https://example.org/documentation/
301 -> 200 (1 hop) https://example.org/news
linked from https://example.org/news/The fix is a single convention, enforced in code. In a React or Next.js codebase, route every internal link through one helper:
// โ
One place decides what an internal URL looks like
export function internalHref(path: string): string {
const clean = path.toLowerCase().replace(/\/+$/, '') // no trailing slash
return clean === '' ? '/' : clean
}
// <Link href={internalHref('/Pricing/')}> renders href="/pricing"3. Hard-coded http:// or old-domain URLs
After an HTTPS or domain migration, CMS content still contains absolute URLs like http://old-domain.com/about. Each one is at least one redirect hop, and two if the old domain also redirects.
-- WordPress: rewrite stored absolute URLs in post content (back up first)
UPDATE wp_posts
SET post_content = REPLACE(post_content, 'http://old-domain.com', 'https://new-domain.com')
WHERE post_content LIKE '%http://old-domain.com%';4. Placeholder href="#" and javascript: links
40.2% of homepages in our sample shipped href="#", and 17.8% used javascript: URLs. Neither is a crawlable link. If the element opens a menu or a modal, it should be a <button>.
<!-- โ Not a link: no destination, confuses crawlers and screen readers -->
<a href="#" onclick="openMenu()">Products</a>
<!-- โ
Fixed: a button for actions, a real href for navigation -->
<button type="button" aria-expanded="false" aria-controls="products-menu">Products</button>
<a href="/products">All products</a>Google only follows <a> elements with an href. Our article on JavaScript links and Googlebot discovery explains why onclick navigation is invisible to crawlers.
5. Deleted pages still linked from navigation
A product was discontinued and its page deleted, but the footer, sitemap and three blog posts still link to it. Decide on the replacement first:
| Situation | Do this | Status |
|---|---|---|
| A direct replacement exists | Update links; 301 the old URL to the replacement | 301 |
| Gone for good, no equivalent | Remove the links; return 410 | 410 |
| Temporarily unavailable | Keep the page and say so on it | 200 |
| Never existed (a typo in the href) | Fix the href | โ |
Never 301 a deleted page to your homepage. Google treats irrelevant redirects like that as soft 404s anyway.
6. Links that only work when logged in
Our checker flagged a /favorites/ link on a large public site as a 404. Signed in, it works; signed out, a crawler gets an error. Links to account-only areas should be shown to logged-in users only, or point to a login page that returns 200.
A Copy-Paste Script to Find Every Broken Internal Link
This Python script crawls up to --max-pages pages on your site and checks each unique internal target once. It then prints every broken or redirecting target alongside the pages that link to it. It treats 401/403/429/503 as unverified, not broken, which matches the rule we used for the dataset above.
#!/usr/bin/env python3
"""Find broken and redirected internal links on a website.
Usage: pip install httpx
python3 find_broken_internal_links.py https://example.com --max-pages 200
"""
import argparse, asyncio, sys
from collections import defaultdict, deque
from html.parser import HTMLParser
from urllib.parse import urljoin, urlsplit, urlunsplit
import httpx
UNKNOWN = {401, 403, 429, 503, 999} # bot blocks / rate limits: re-check by hand
class Links(HTMLParser):
def __init__(self):
super().__init__()
self.hrefs = []
def handle_starttag(self, tag, attrs):
if tag == "a":
href = dict(attrs).get("href")
if href:
self.hrefs.append(href.strip())
def normalise(url):
p = urlsplit(url)
return urlunsplit((p.scheme, p.netloc.lower(), p.path or "/", p.query, ""))
async def main(start, max_pages, concurrency):
host = urlsplit(start).netloc.lower()
sources = defaultdict(set)
seen, queue, crawled = {normalise(start)}, deque([normalise(start)]), 0
sem = asyncio.Semaphore(concurrency)
headers = {"User-Agent": "Mozilla/5.0 (compatible; link-audit/1.0)"}
async with httpx.AsyncClient(timeout=15, headers=headers) as client:
async def check(url):
async with sem:
try:
r = await client.head(url)
if r.status_code in (400, 403, 405, 501): # HEAD not supported
r = await client.get(url)
first = r.status_code
if 300 <= first < 400:
f = await client.get(url, follow_redirects=True)
return first, f.status_code, len(f.history)
return first, first, 0
except httpx.HTTPError as exc:
return type(exc).__name__, None, 0
while queue and crawled < max_pages:
page = queue.popleft()
try:
r = await client.get(page, follow_redirects=True)
except httpx.HTTPError:
continue
crawled += 1
if "text/html" not in r.headers.get("content-type", ""):
continue
parser = Links()
parser.feed(r.text)
for href in parser.hrefs:
if href.startswith(("#", "mailto:", "tel:", "javascript:")):
continue
target = normalise(urljoin(str(r.url), href))
if urlsplit(target).netloc.lower() != host:
continue
sources[target].add(page)
if target not in seen:
seen.add(target)
queue.append(target)
targets = list(sources)
status = dict(zip(targets, await asyncio.gather(*(check(t) for t in targets))))
broken = {t: s for t, s in status.items()
if s[1] is None or (isinstance(s[1], int) and s[1] >= 400 and s[1] not in UNKNOWN)}
redirected = {t: s for t, s in status.items()
if isinstance(s[0], int) and 300 <= s[0] < 400 and t not in broken}
unknown = {t: s for t, s in status.items() if s[1] in UNKNOWN}
print(f"Crawled {crawled} pages, checked {len(targets)} unique internal URLs\n")
for label, group in (("BROKEN", broken), ("REDIRECTED (update the href)", redirected),
("UNVERIFIED (blocked/rate-limited)", unknown)):
print(f"== {label}: {len(group)}")
for t, (first, final, hops) in sorted(group.items()):
print(f" {first} -> {final} ({hops} hop{'s' if hops != 1 else ''}) {t}")
for src in sorted(sources[t])[:3]:
print(f" linked from {src}")
print()
return 1 if broken else 0
if __name__ == "__main__":
ap = argparse.ArgumentParser()
ap.add_argument("url")
ap.add_argument("--max-pages", type=int, default=200)
ap.add_argument("--concurrency", type=int, default=8)
a = ap.parse_args()
sys.exit(asyncio.run(main(a.url, a.max_pages, a.concurrency)))The script exits with code 1 when it finds broken links. You can therefore run it in a pre-deploy job and fail the build. Keep --concurrency low (4โ8) on shared hosting so your own WAF doesn't start answering with 429.
It parses raw HTML, so links injected by client-side JavaScript won't appear. For a React or Vue single-page app, use a rendering crawler. The section below shows one.
Step-by-Step Remediation Workflow
- Export the list. Run the script (or a full-site crawl) and save the BROKEN and REDIRECTED groups with their source pages.
- Fix templates first. Sort by how many source pages share a target. A target linked from 40 pages is almost always in a header, footer or sidebar component: one edit fixes all 40.
- Decide each dead target using the table above: update, 301, 410 or remove.
- Update hrefs, not just redirects. For each REDIRECTED link, replace the href with the final URL the chain resolves to.
- Check the sitemap. Your XML sitemap should list only canonical 200 URLs. Our XML sitemap guide covers the rules.
- Re-crawl and diff. The run is clean when BROKEN is empty and REDIRECTED contains only links you deliberately keep (for example,
/loginredirecting to an SSO provider).
๐ก Prioritise by traffic, not count. One broken link in the main navigation costs more than twenty in a 2019 blog post nobody reads. Fix navigation, high-traffic pages and conversion paths first.
How BugViso Catches Broken Internal Links Automatically
BugViso's Concurrent Link & Image Validation engine collects every href and image src on each page after JavaScript has rendered in Chromium, so it also sees links that a raw-HTML crawler misses. It sorts them into internal and external and requests them concurrently. Each one is classified as a 404, a server error, a redirect loop or a healthy response.
On a multi-page crawl, BugViso discovers URLs from your sitemap.xml (including sitemaps declared in robots.txt) with a same-host breadth-first fallback. It runs the same link validation on every page it audits. The internal link graph built from that crawl also surfaces orphan pages and click depth, so you can see when a broken link was a page's only inbound path.
Rate-limit responses (429) and bot-protection challenges are not counted as broken links, the same rule we applied to the dataset above. Findings land in the web report and in the PDF, CSV and JSON exports, so you can hand developers a list of source page โ broken target pairs. With scheduled weekly or monthly scans, a link broken by next month's content update shows up in the next report instead of in Search Console weeks later.
You can run a free BugViso scan on any URL to see the link results for your own site.
Traps and Edge Cases That Cause False Positives
HEADnot supported. Some servers answerHEADwith 405 or 403 but serveGETfine. Always retry withGETbefore calling a link broken.- Rate limiting during the crawl. A checker firing 50 parallel requests trips your own WAF. Every 429 it records is a false positive, so throttle it.
- Geo and language redirects. A link that 302s to
/en-gb/from London may go straight to 200 from New York. Check redirect chains from the region your audience is in. - Links added by JavaScript. "Related products" widgets and infinite-scroll pagination often build hrefs client-side. A static crawler misses both their links and their breakages.
- Fragment-only links.
href="/guide#step-3"returns 200 even if the#step-3anchor no longer exists. Status codes can't catch a missing anchor; checkdocument.getElementByIdinstead. - Case-sensitive servers.
/Aboutand/aboutare different URLs on Linux hosting. One may work and the other 404.
FAQ
Do broken internal links hurt SEO rankings directly?
Not as a sitewide penalty. Google doesn't demote a site for having some 404s. The cost is indirect: wasted crawl requests, lost link equity to the target, and pages that become harder to discover. Broken links in navigation also hurt engagement, which affects conversions even when rankings stay put.
Should I redirect or update a broken internal link?
Update it. Redirects exist for URLs you don't control, such as backlinks, bookmarks and old emails. For links in your own HTML, edit the href to point directly at the live URL so every request gets a single 200 response.
How often should I check for broken internal links?
After every deploy that changes URLs or navigation, and on a monthly schedule otherwise. Sites that publish often, or remove products, should check weekly. Content changes break links far more often than code changes do.
Is a 301 internal redirect a broken link?
No, but it's worth fixing. It works for users, yet each one adds a network round trip and can turn into a chain when the destination moves again. In our sample, 15.8% of redirecting internal links already went through two or more hops.
Why does my checker show 403 errors on pages that load fine in my browser?
Bot protection, usually a CDN or WAF rule that blocks non-browser user agents or high request rates. Treat 403/429/503 as "unverified" and re-check those URLs manually or from an allow-listed IP.
Conclusion
Broken internal links are rare on any single page, yet more than half of the sites we sampled were sending their own crawlers through redirects they could have removed with a one-line href change. Fix the links you control at the source, keep 301s for everyone else, and re-crawl until every internal link returns a direct 200, which is exactly what a free BugViso audit checks on every page it crawls.
See where your site stands
Run a free BugViso audit for SEO, speed, accessibility and AI search readiness โ with fixes you can ship today.