Website Structure Audit: How to Audit Site Architecture
How to run a website structure audit: click depth, orphan pages, internal link distribution, URL hierarchy and sitemap hygiene, with a script and report template.
Quick answer: A website structure audit checks how your pages connect: how many clicks each page is from the homepage (click depth), which pages have no internal links (orphans), where internal links concentrate, whether URLs follow a logical hierarchy, and whether your XML sitemap matches the pages you actually link to. Important pages should be reachable within three clicks, linked from relevant hubs, and listed in the sitemap at their final URL.
Search engines discover and weigh pages mostly through internal links. A page buried six clicks deep, or linked from nowhere, gets crawled less and inherits little authority, however good its content. A structure audit finds those pages and the template decisions that created them.
This guide covers what to check, a step-by-step process, a script that does the crawling for you and a report template.
What a Website Structure Audit Checks
| Check | What you're looking for | Why it matters |
|---|---|---|
| Click depth | Pages more than 3 clicks from the homepage | Deep pages are crawled less often and look less important |
| Orphan pages | Pages in the sitemap or analytics that no page links to | Crawlers can only find them through the sitemap; they get no internal link equity |
| Internal link distribution | Which pages receive the most links, and which receive almost none | Links signal importance; money pages should be well linked |
| URL hierarchy | Folders that reflect sections (/blog/, /products/category/item) | Clear sections help users and crawlers understand grouping |
| Navigation and breadcrumbs | Crawlable <a href> links in menus, footers and breadcrumbs | JavaScript-only menus hide structure from crawlers |
| Sitemap vs crawl | URLs in one but not the other; sitemap URLs that redirect or 404 | The sitemap should list exactly your indexable, final URLs |
Google's guidance on making links crawlable is the baseline: links need to be <a> elements with an href to count.
What We Found Auditing Real Site Structures
We crawled up to 250 pages on each of 80 sites drawn at random from our study sample (Tranco ranks 1,001–50,000), following raw-HTML links from the homepage, and compared the result with each site's XML sitemap.
| Finding | Result |
|---|---|
Sites with a sitemap findable via robots.txt or /sitemap.xml | 38 of 76 that completed (50%) |
| Median sitemap size | 1,412 URLs |
Sitemaps where every sampled URL returned 200 | 22 of 36 (61%) |
| Sitemaps listing URLs that redirect | 9 of 36 (25%) |
Sitemaps listing URLs that return 404/410 | 5 of 36 (14%) |
| Internal links on a typical page (median) | 75 |
And from our separate raw-vs-rendered study of 219 homepages, 9.2% exposed fewer than half of their links without JavaScript. Those sites' structure is largely invisible to crawlers that don't render, including most AI crawlers.
The pattern: sitemaps drift. Pages get redirected or removed, but the sitemap generator keeps listing the old URLs. A structure audit is where that drift gets caught.
How to Run a Website Structure Audit, Step by Step
Step 1: Crawl from the homepage, following only links
Start at the homepage and follow internal <a href> links breadth-first. Record the click depth of each page as the first time it's reached. Don't seed the crawl from the sitemap for this step, or you'll hide orphans.
Step 2: Check click depth
List pages by depth. Anything important at depth 4 or more needs a shorter path: a link from a hub page, a category page or the main navigation.
| Depth | Typical pages | Action if important pages land here |
|---|---|---|
| 1 | Main navigation targets | — |
| 2 | Category and hub pages, recent posts | — |
| 3 | Product and article pages | Acceptable |
| 4+ | Paginated archives, old content | Add links from hubs or related-content blocks |
Step 3: Find orphan pages
Compare the crawl with the sitemap (and with landing pages from analytics or Search Console). Sitemap URLs the crawl never reached are orphan candidates. For each one, decide: link it from a relevant page, merge it, or remove it. Our guide to orphan pages covers the fixes.
Step 4: Review internal link distribution
Count inbound internal links per page. Your most important commercial pages should be among the most linked. If your most-linked pages are the privacy policy and a tag archive, your templates are distributing links badly. See our internal linking audit checklist.
Step 5: Check URL hierarchy
Folders should match sections. Watch for the same content under two paths (/blog/post and /news/post), parameters creating duplicate pages, and inconsistent trailing slashes.
Step 6: Check navigation, footers and breadcrumbs are crawlable
View the page source (not DevTools) and confirm menu links are real <a href> elements. Breadcrumbs should be links, ideally with BreadcrumbList structured data; see breadcrumb navigation for SEO.
Step 7: Reconcile the sitemap
The sitemap should list every indexable page at its final URL, and nothing else: no redirects, no 404s, no noindex pages, no parameter duplicates. Google's sitemap overview explains the format.
Script: Audit Your Site Structure
This standard-library Python script runs steps 1, 2, 3 and 7. It crawls raw HTML only, so links that exist only after JavaScript runs won't be followed (which is exactly what non-rendering crawlers see).
"""Website structure audit: click depth, orphan pages and sitemap hygiene.
Usage: python3 structure_audit.py https://example.com/ [max_pages]
Crawls same-host <a href> links breadth-first from the homepage (raw HTML, no
JavaScript), then compares the result with the sitemap:
- click depth of every page reached
- sitemap URLs never reached by links (orphan candidates)
- linked pages missing from the sitemap
- sitemap URLs that redirect or error
Standard library only.
"""
import html as htmllib, re, sys, urllib.error, urllib.request
from collections import Counter, deque
from urllib.parse import urljoin, urlsplit, urlunsplit
UA = {"User-Agent": "Mozilla/5.0 (structure-audit)"}
SKIP = re.compile(r"\.(jpe?g|png|gif|webp|avif|svg|pdf|zip|mp4|css|js|xml|json|ico)$", re.I)
class NoRedirect(urllib.request.HTTPRedirectHandler):
def redirect_request(self, *args, **kwargs):
return None
def fetch(url, follow=True):
opener = urllib.request.build_opener() if follow else urllib.request.build_opener(NoRedirect)
try:
r = opener.open(urllib.request.Request(url, headers=UA), timeout=15)
ctype = r.headers.get("Content-Type", "")
body = r.read(3_000_000).decode("utf-8", "replace") if "html" in ctype or "xml" in ctype else ""
return r.status, r.geturl(), body
except urllib.error.HTTPError as e:
return e.code, url, ""
except Exception: # noqa: BLE001
return None, url, ""
def norm(url):
s = urlsplit(url)
path = s.path.rstrip("/") or "/"
return urlunsplit((s.scheme, s.netloc.lower().removeprefix("www."), path, s.query, ""))
def crawl(home, max_pages):
host = urlsplit(home).netloc.lower().removeprefix("www.")
depth, queue, pages = {norm(home): 0}, deque([home]), 0
while queue and pages < max_pages:
url = queue.popleft()
status, final, html = fetch(url)
pages += 1
if status != 200:
continue
for href in re.findall(r'<a\s[^>]*href=["\']([^"\'#]+)', html, re.I):
link = urljoin(final, htmllib.unescape(href))
s = urlsplit(link)
if s.scheme in ("http", "https") and s.netloc.lower().removeprefix("www.") == host and not SKIP.search(s.path):
key = norm(link)
if key not in depth:
depth[key] = depth[norm(url)] + 1
queue.append(link)
return depth, not queue
def sitemap(home):
base = "{0.scheme}://{0.netloc}".format(urlsplit(home))
_, _, robots = fetch(base + "/robots.txt")
maps = re.findall(r"(?im)^sitemap:\s*(\S+)", robots) or [base + "/sitemap.xml"]
urls = []
for sm in maps[:5]:
_, _, xml = fetch(sm)
locs = re.findall(r"<loc>\s*([^<\s]+)\s*</loc>", xml)
if "<sitemapindex" in xml:
for child in locs[:10]:
urls += re.findall(r"<loc>\s*([^<\s]+)\s*</loc>", fetch(child)[2])
else:
urls += locs
return list(dict.fromkeys(u.replace("&", "&") for u in urls))
def main(home, max_pages=500):
depth, complete = crawl(home, max_pages)
listed = sitemap(home)
listed_keys = {norm(u) for u in listed}
print(f"Crawled {min(len(depth), max_pages)} pages ({'whole site' if complete else 'stopped at cap'}), sitemap lists {len(listed)} URLs\n")
print("Click depth from homepage:")
for d, n in sorted(Counter(min(v, 5) for v in depth.values()).items()):
print(f" {d if d < 5 else '5+'} clicks: {n} pages")
orphans = [u for u in listed if norm(u) not in depth]
print(f"\nSitemap URLs not reached by internal links: {len(orphans)}" + ("" if complete else " (crawl hit the cap; some may just be deeper)"))
for u in orphans[:10]:
print(" ", u)
missing = [u for u in depth if u not in listed_keys]
print(f"\nLinked pages missing from the sitemap: {len(missing)}")
for u in missing[:5]:
print(" ", u)
codes = Counter()
bad = []
for u in listed[:50]:
code = fetch(u, follow=False)[0]
codes[code] += 1
if code != 200:
bad.append((code, u))
print(f"\nSitemap hygiene (first {min(50, len(listed))} URLs): {dict(codes)}")
for code, u in bad[:10]:
print(f" {code} {u}")
if __name__ == "__main__":
main(sys.argv[1], int(sys.argv[2]) if len(sys.argv) > 2 else 500)Example output:
Crawled 257 pages (whole site), sitemap lists 256 URLs
Click depth from homepage:
0 clicks: 1 pages
1 clicks: 8 pages
2 clicks: 242 pages
3 clicks: 5 pages
4 clicks: 1 pages
Sitemap URLs not reached by internal links: 4
https://example.com/terms
https://example.com/privacy
https://example.com/refund
https://example.com/contact-sales
Linked pages missing from the sitemap: 5
https://example.com/register?plan=pro&interval=monthly
Sitemap hygiene (first 50 URLs): {200: 50}In this example, four legal and sales pages are only linked from a JavaScript-rendered footer, so a raw-HTML crawl never reaches them; the "missing from sitemap" URLs are parameter links that correctly stay out of the sitemap. The script surfaces candidates; you decide which are real problems.
Website Structure Audit Report Template
# Website Structure Audit: {site} — {date}
## Summary
- Pages crawled: {n} | In sitemap: {n} | Crawl complete: {yes/no}
- Pages deeper than 3 clicks: {n} ({%})
- Orphan candidates: {n} | Sitemap URLs not returning 200: {n}
## Click depth
| Depth | Pages | Important pages at this depth |
|---|---|---|
## Orphan candidates
| URL | In sitemap | Traffic (last 3 months) | Action: link / merge / remove |
|---|---|---|---|
## Internal link distribution
| Page | Inbound internal links | Should it be higher or lower? |
|---|---|---|
## Sitemap issues
| URL | Status | Fix |
|---|---|---|
## Recommendations (ordered)
1. ...How BugViso Audits Site Structure
A BugViso audit crawls your site through a headless browser, so links injected by JavaScript are followed too, and builds an internal link graph from every crawled page. It reports:
- Orphan pages with no inbound internal links among the crawled pages (marked as unconfirmed when the crawl stopped before covering the whole site).
- Click depth from the homepage, flagging pages three or more clicks deep.
- Pages with excessive outbound links.
- The most internally linked pages, so you can see where link equity concentrates.
Confirmed orphans cost points in the health score (2 per orphan, capped at 8), alongside broken links, redirects and canonical problems found in the same crawl.
More detail is on the structured data and duplicate content checks feature page.
Common Structure Audit Mistakes
- Seeding the crawl from the sitemap. Every sitemap URL then looks linked, and orphans disappear.
- Auditing the rendered DOM only. Links that need JavaScript are invisible to many crawlers. Compare raw and rendered.
- Counting footer links as structure. A link from every page's footer says little about topical relationships; contextual links from relevant pages matter more.
- Ignoring pagination. Deep pagination pushes old content far from the homepage; add category hubs or "related posts" links.
- Leaving redirects in the sitemap. One in four sitemaps we sampled listed redirecting URLs.
FAQ
What is a website structure audit?
A review of how a site's pages are organised and linked: click depth, orphan pages, internal link distribution, URL hierarchy, navigation and sitemap accuracy.
What is a good click depth for SEO?
Keep important pages within three clicks of the homepage. Deeper pages are crawled less often and usually receive fewer internal links.
How do I find orphan pages?
Crawl the site by following links from the homepage, then compare the pages you reached with your sitemap and analytics landing pages. URLs that appear in the sitemap or analytics but not in the crawl are orphans.
What should a website structure audit report include?
Pages crawled, click-depth distribution, orphan candidates with recommended actions, internal link distribution, sitemap issues and an ordered list of fixes. The template above covers each.
How often should I audit site structure?
After every redesign, migration or navigation change, and at least twice a year for sites that publish regularly.
The Takeaway
Structure decides which of your pages crawlers find and value, and a BugViso crawl maps click depth and orphan pages across your site in one pass.
See where your site stands
Run a free BugViso audit for SEO, speed, accessibility and AI search readiness — with fixes you can ship today.