Visualize Internal Link Graphs with Python and NetworkX

Learn how to build, analyze, and visualize your website internal link graph using Python, NetworkX, and PyVis to detect orphan clusters and equity bottlenecks.

BugViso

14 min read

Spreadsheets and flat tabular reports are ill-suited for analyzing website architecture. When an SEO audit outputs 20,000 rows of crawl data, reviewing columns of source and target URLs provides zero intuitive insight into information architecture, topical clustering, or authority bottlenecks.

A website is fundamentally a directed mathematical graph ($G = (V, E)$), where each unique URL represents a vertex (node) and each internal hyperlink represents a directed edge pointing from source to destination. By modeling your website as a formal network graph, you can apply graph theory algorithms to discover structural anomalies that tabular audits completely miss: isolated orphan clusters, high-betweenness bottlenecks, and unbalanced authority distribution.

This guide provides an end-to-end engineering blueprint for building, analyzing, and visualizing your website's internal link graph using Python, networkx, and pyvis.


The Graph Theory Toolkit for Technical Architecture

Before writing code, we must understand the core network metrics that describe the health of an internal link topology:

Diagram
┌─────────────────────────────────────────────────────────────┐
│                 Network Graph Metrics Map                   │
├──────────────────────────────┬──────────────────────────────┤
│ Metric                       │ Architectural Significance   │
├──────────────────────────────┼──────────────────────────────┤
│ In-Degree Centrality         │ Total incoming internal links│
│ Out-Degree Centrality        │ Total outbound links on page │
│ Betweenness Centrality       │ Critical informational bridge│
│ PageRank Centrality          │ Stationary authority weight  │
│ Strongly Connected (SCC)     │ Mutual reachability groups   │
│ Weakly Connected (WCC)       │ Structural islands / orphans │
└──────────────────────────────┴──────────────────────────────┘

In-Degree vs. Out-Degree

  • In-Degree ($k_{in}$): The count of directed edges pointing to a node. High in-degree indicates that a page is widely referenced across your site (e.g., your homepage, terms page, or top-level category hub).
  • Out-Degree ($k_{out}$): The count of outgoing links originating from a node. Extremely high out-degree (e.g., uncurated footer menus with 250+ links) dilutes the equity transferred along each edge.

Eigenvector Centrality vs. PageRank Centrality

While both metrics quantify node influence by considering the importance of neighboring nodes, their mathematical behavior differs significantly:

  • Eigenvector Centrality: Solves for the principal eigenvector of the adjacency matrix: $\lambda x = A x$. It assumes that connections to high-scoring nodes contribute more to the score of the node. However, on directed acyclic graphs (DAGs) or graphs with dead ends, eigenvector centrality fails to converge or collapses to zero for non-cyclic subgraphs.
  • PageRank Centrality: Resolves this limitation by incorporating the damping factor ($d = 0.85$) and random teleportation matrix $\frac{1 - d}{N} \mathbf{E}$. This guarantees convergence across all directed web graphs, regardless of whether cycles or sinkholes exist.

Betweenness Centrality: Detecting Bottlenecks

Betweenness centrality measures the fraction of all shortest paths between all pairs of nodes in the graph that pass through a specific node $v$:

$$C_B(v) = \sum_{s \neq v \neq t} \frac{\sigma_{st}(v)}{\sigma_{st}}$$

Where $\sigma_{st}$ is the total number of shortest paths from node $s$ to node $t$, and $\sigma_{st}(v)$ is the number of those paths that pass through $v$.

In technical SEO, a page with high betweenness centrality acts as a critical gateway. If this bridge page breaks (returns a 404 or 500 error) or is removed during a CMS migration, entire sections of your website become disconnected from the main crawl graph. For procedures on detecting broken pathways before they affect crawlers, review our guide on broken link checker tools for large sites.

Louvain Community Detection: Evaluating Topical Silos

To assess whether your internal links respect topical clustering, you can apply the Louvain modularity algorithm. Modularity ($Q$) measures the density of links inside communities compared to links between communities:

$$Q = \frac{1}{2m} \sum_{i, j} \left[ A_{ij} - \frac{k_i k_j}{2m} \right] \delta(c_i, c_j)$$

Where:

  • $A_{ij}$ is the edge weight between nodes $i$ and $j$.
  • $k_i, k_j$ are the degrees of the nodes.
  • $m$ is the total number of edges.
  • $\delta(c_i, c_j)$ is 1 if nodes $i$ and $j$ belong to the same community, 0 otherwise.

When executed on your crawl graph, Louvain clustering should partition your site along clear semantic boundaries (e.g., all database articles cluster together; all frontend articles cluster together). If the algorithm groups unrelated topics into a single community, your site suffers from excessive, unfocused cross-linking.

Connected Components: Finding Structural Islands

  • Strongly Connected Component (SCC): A maximal subgraph where every node is reachable from every other node along directed paths. A healthy topic cluster should form an SCC.
  • Weakly Connected Component (WCC): A subgraph where nodes are connected if edge directions are ignored. If your graph contains multiple disjoint WCCs, you have completely isolated "islands" that search crawlers cannot discover through standard link traversal.

Data Pipeline: Extracting Directed Edges from Crawl Data

To build our graph, we need a dataset containing directed link pairs: a source URL and a target URL.

The following Python script uses BeautifulSoup and requests to crawl an internal domain, extract internal hyperlinks, strip fragments and tracking query parameters, and export a clean edge list to CSV:

python
#!/usr/bin/env python3
"""
crawl_link_edges.py - High-throughput crawler to extract internal link edges.
Saves clean directed edges (source, target) for graph analysis.
"""

import csv
import sys
from urllib.parse import urljoin, urlparse, urldefrag
from collections import deque
import requests
from bs4 import BeautifulSoup

def is_internal_url(url: str, base_domain: str) -> bool:
    parsed = urlparse(url)
    return parsed.netloc == base_domain and parsed.scheme in ["http", "https"]

def clean_url(url: str) -> str:
    # Strip URL fragments (#section) and normalize trailing slashes
    clean, _ = urldefrag(url)
    parsed = urlparse(clean)
    path = parsed.path.rstrip("/") if parsed.path != "/" else "/"
    return f"{parsed.scheme}://{parsed.netloc}{path}"

def crawl_edges(start_url: str, max_pages: int = 500, output_csv: str = "crawl_edges.csv"):
    parsed_start = urlparse(start_url)
    base_domain = parsed_start.netloc
    start_url_clean = clean_url(start_url)
    
    visited = set()
    queue = deque([start_url_clean])
    edges = set()
    
    session = requests.Session()
    session.headers.update({"User-Agent": "GraphCrawler/1.0"})
    
    print(f"[*] Starting crawl on: {start_url_clean}")
    print(f"[*] Target max pages: {max_pages}")
    
    while queue and len(visited) < max_pages:
        current_url = queue.popleft()
        
        if current_url in visited:
            continue
            
        visited.add(current_url)
        print(f"[{len(visited)}/{max_pages}] Crawling: {current_url}")
        
        try:
            resp = session.get(current_url, timeout=5)
            if "text/html" not in resp.headers.get("Content-Type", ""):
                continue
                
            soup = BeautifulSoup(resp.text, "html.parser")
            for a_tag in soup.find_all("a", href=True):
                raw_href = a_tag["href"].strip()
                abs_url = urljoin(current_url, raw_href)
                target_url = clean_url(abs_url)
                
                if is_internal_url(target_url, base_domain):
                    # Record directed edge: source -> target
                    edges.add((current_url, target_url))
                    
                    if target_url not in visited and target_url not in queue:
                        queue.append(target_url)
                        
        except Exception as err:
            print(f"[!] Error fetching {current_url}: {err}")
            
    # Export edges to CSV
    print(f"\n[*] Crawl complete. Exporting {len(edges)} unique edges to {output_csv}...")
    with open(output_csv, "w", newline="", encoding="utf-8") as f:
        writer = csv.writer(f)
        writer.writerow(["source", "target"])
        for src, tgt in sorted(edges):
            writer.writerow([src, tgt])
            
    print(f"[+] Successfully generated edge list: {output_csv}")

if __name__ == "__main__":
    if len(sys.argv) < 2:
        print("Usage: python3 crawl_link_edges.py <https://example.com> [max_pages]")
        sys.exit(1)
        
    pages = int(sys.argv[2]) if len(sys.argv) > 2 else 250
    crawl_edges(sys.argv[1], max_pages=pages)

Run the edge extraction crawler:

bash
python3 scripts/crawl_link_edges.py https://example.com 300

Building the Directed Graph and Computing Metrics with NetworkX

Once the edge list is exported, we load the dataset into networkx to instantiate a directed graph (nx.DiGraph()) and compute our topological metrics:

python
#!/usr/bin/env python3
"""
analyze_link_graph.py - Analyzes internal link network topology.
Computes In-Degree, Out-Degree, Betweenness, and PageRank Centrality using NetworkX.
"""

import sys
import pandas as pd
import networkx as nx

def analyze_graph(edge_csv: str):
    print(f"[*] Ingesting edge list from: {edge_csv}")
    df = pd.read_csv(edge_csv)
    
    # Instantiate directed graph
    G = nx.DiGraph()
    for _, row in df.iterrows():
        G.add_edge(row['source'], row['target'])
        
    num_nodes = G.number_of_nodes()
    num_edges = G.number_of_edges()
    
    print(f"\n" + "="*60)
    print(f"NETWORK GRAPH TOPOLOGY OVERVIEW:")
    print(f"="*60)
    print(f"Total Unique Nodes (Pages): {num_nodes}")
    print(f"Total Directed Edges (Links): {num_edges}")
    print(f"Graph Density: {nx.density(G):.5f}")
    
    # 1. Compute PageRank
    print("[*] Computing PageRank centrality (damping=0.85)...")
    pagerank = nx.pagerank(G, alpha=0.85)
    
    # 2. Compute Betweenness Centrality
    print("[*] Computing Betweenness centrality...")
    betweenness = nx.betweenness_centrality(G)
    
    # 3. Degrees
    in_degrees = dict(G.in_degree())
    out_degrees = dict(G.out_degree())
    
    # Compile Node Metrics DataFrame
    metrics = []
    for node in G.nodes():
        metrics.append({
            "url": node,
            "in_degree": in_degrees.get(node, 0),
            "out_degree": out_degrees.get(node, 0),
            "betweenness": betweenness.get(node, 0.0),
            "pagerank": pagerank.get(node, 0.0)
        })
        
    metrics_df = pd.DataFrame(metrics)
    
    # 4. Connected Components Analysis
    print("\n" + "="*60)
    print("CONNECTED COMPONENTS & ORPHAN CLUSTER AUDIT:")
    print("="*60)
    wccs = list(nx.weakly_connected_components(G))
    print(f"Total Weakly Connected Components: {len(wccs)}")
    
    if len(wccs) > 1:
        print("[!] STRUCTURAL ALERT: Disconnected islands detected!")
        for idx, comp in enumerate(wccs[1:], start=2):
            print(f"  Component #{idx} ({len(comp)} nodes):")
            for url in list(comp)[:5]:
                print(f"    - {url}")
            if len(comp) > 5:
                print(f"    ... and {len(comp) - 5} more.")
    else:
        print("[+] SUCCESS: The entire crawled network is weakly connected.")
        
    # Top Authority Nodes
    print("\n" + "="*60)
    print("TOP 10 NODES BY INTERNAL PAGERANK:")
    print("="*60)
    top_pr = metrics_df.sort_values(by="pagerank", ascending=False).head(10)
    for idx, row in top_pr.reset_index().iterrows():
        print(f"{idx+1:02d}. [PR: {row['pagerank']:.5f} | In: {row['in_degree']:3d}] {row['url']}")
        
    # Top Bottleneck Bridges
    print("\n" + "="*60)
    print("TOP 5 STRUCTURAL BRIDGES (BETWEENNESS CENTRALITY):")
    print("="*60)
    top_between = metrics_df.sort_values(by="betweenness", ascending=False).head(5)
    for idx, row in top_between.reset_index().iterrows():
        print(f"{idx+1:02d}. [Betweenness: {row['betweenness']:.5f}] {row['url']}")
        
    metrics_df.to_csv("network_metrics_report.csv", index=False)
    print(f"\n[+] Full metrics export saved to network_metrics_report.csv")
    return G, metrics_df

if __name__ == "__main__":
    csv_input = sys.argv[1] if len(sys.argv) > 1 else "crawl_edges.csv"
    analyze_graph(csv_input)

Run the graph analysis:

bash
python3 scripts/analyze_link_graph.py crawl_edges.csv

Interactive Visualization: Generating Force-Directed Graphs with PyVis

While numerical metrics reveal structural statistics, visual representation is necessary to communicate architectural realities to non-technical stakeholders and cross-functional engineering teams.

Using pyvis, we can export an interactive, physics-driven HTML visualization that renders directly in any modern browser. In the script below, we color nodes based on their topic path and scale their physical radius proportionally to their calculated PageRank:

python
#!/usr/bin/env python3
"""
visualize_pyvis.py - Generates an interactive force-directed graph with PyVis.
Nodes are sized by PageRank and colored by path category.
"""

import sys
from urllib.parse import urlparse
import pandas as pd
import networkx as nx
from pyvis.network import Network

def assign_color_by_path(url: str) -> str:
    path = urlparse(url).path
    if path == "" or path == "/":
        return "#FF4444"  # Homepage: Vibrant Red
    elif path.startswith("/blog"):
        return "#3B82F6"  # Blog: Blue
    elif path.startswith("/docs") or path.startswith("/guides"):
        return "#10B981"  # Documentation: Emerald Green
    elif path.startswith("/product") or path.startswith("/features"):
        return "#8B5CF6"  # Commercial: Purple
    else:
        return "#9CA3AF"  # General Utility: Gray

def generate_interactive_visualization(edge_csv: str, output_html: str = "link_graph.html"):
    df = pd.read_csv(edge_csv)
    
    G = nx.DiGraph()
    for _, row in df.iterrows():
        G.add_edge(row['source'], row['target'])
        
    print(f"[*] Calculating PageRank for visual node scaling...")
    pr = nx.pagerank(G, alpha=0.85)
    max_pr = max(pr.values()) if pr else 1.0
    
    # Initialize PyVis Network
    net = Network(
        height="900px",
        width="100%",
        directed=True,
        bgcolor="#111827",
        font_color="#F9FAFB"
    )
    
    # Configure physics simulation engine for clean cluster repulsion
    net.set_options("""
    {
      "nodes": {
        "borderWidth": 1,
        "borderWidthSelected": 3
      },
      "edges": {
        "color": { "inherit": true },
        "smooth": { "type": "continuous" },
        "arrows": { "to": { "enabled": true, "scaleFactor": 0.5 } }
      },
      "physics": {
        "forceAtlas2Based": {
          "gravitationalConstant": -50,
          "centralGravity": 0.01,
          "springLength": 100,
          "springConstant": 0.08
        },
        "maxVelocity": 50,
        "solver": "forceAtlas2Based",
        "timestep": 0.35,
        "stabilization": { "iterations": 150 }
      }
    }
    """)
    
    print(f"[*] Adding {G.number_of_nodes()} nodes and {G.number_of_edges()} edges to PyVis...")
    
    for node in G.nodes():
        node_pr = pr.get(node, 0.0)
        # Normalize radius between 10px and 45px
        node_size = 10 + (node_pr / max_pr) * 35
        node_color = assign_color_by_path(node)
        
        # Tooltip content
        in_deg = G.in_degree(node)
        out_deg = G.out_degree(node)
        title_tooltip = f"<b>{node}</b><br>PageRank: {node_pr:.5f}<br>In-Links: {in_deg}<br>Out-Links: {out_deg}"
        
        parsed = urlparse(node)
        label = parsed.path if parsed.path else "/"
        if len(label) > 25:
            label = label[:22] + "..."
            
        net.add_node(
            node,
            label=label,
            title=title_tooltip,
            value=node_size,
            color=node_color
        )
        
    for src, tgt in G.edges():
        net.add_edge(src, tgt)
        
    print(f"[*] Generating standalone interactive HTML: {output_html}")
    net.write_html(output_html)
    print(f"[+] Interactive visualization ready! Open '{output_html}' in any web browser.")

if __name__ == "__main__":
    input_csv = sys.argv[1] if len(sys.argv) > 1 else "crawl_edges.csv"
    generate_interactive_visualization(input_csv)

Generate the visualization:

bash
python3 scripts/visualize_pyvis.py crawl_edges.csv

When you open link_graph.html in your browser, the physics simulation activates. Dense thematic clusters naturally coalesce into visual constellations, while orphaned nodes drift to the periphery.


Diagnosing Architectural Anti-Patterns Visually

When inspecting your force-directed link graph, look for these four common structural patterns:

Diagram
┌─────────────────────────────────────────────────────────────┐
│             Visual Network Topology Archetypes              │
├─────────────────────────────────────────────────────────────┤
│ 1. The Star Topology (Centralized Bottleneck):              │
│    All nodes connect exclusively to a single hub.           │
│    Risk: Leaf nodes lack lateral peer cross-links.          │
├─────────────────────────────────────────────────────────────┤
│ 2. The Barbell Topology (Divided Silos):                    │
│    Two dense clusters connected by a single fragile bridge. │
│    Risk: Breaking the bridge cuts off 50% of the site.      │
├─────────────────────────────────────────────────────────────┤
│ 3. The Satellite Island (Disconnected Subgraph):            │
│    A cluster of nodes floating free from the main body.     │
│    Risk: Search crawlers cannot access the island.          │
├─────────────────────────────────────────────────────────────┤
│ 4. The Megamenu Mesh (Over-Densified Graph):                │
│    Every node links to every other node indiscriminately.   │
│    Risk: Total dilution of PageRank across low-value pages. │
└─────────────────────────────────────────────────────────────┘

1. The Barbell Topology

If your blog cluster and your documentation cluster are only connected by a single link in the main navigation, your graph forms a barbell structure.

If an editorial change inadvertently alters that navigation link, the entire documentation cluster's incoming link equity collapses.

Architectural Fix: Inject contextual in-content link bridges between relevant blog articles and technical documentation guides, thickening the connection between clusters.

2. Satellite Islands (Weakly Connected Components)

If your visualization displays a detached cluster floating outside the gravitational pull of the main network, you have discovered an orphaned sub-domain or abandoned landing page campaign.

These pages are accessible only if a crawler already knows the URL (e.g., via XML sitemaps), but receive zero internal authority from the domain's primary backlink profile.

3. The Megamenu Mesh

If your graph appears as an impenetrable, uniform sphere where every node connects to every other node, your site suffers from boilerplate over-linking.

When every page contains 300 identical links in the header and footer, the network graph lacks topological differentiation. Search engines cannot discern which pages represent core topical hubs versus secondary utility endpoints.

For a structured review of sitewide link health, consult our step-by-step full website audit checklist.


A standard Python requests script only inspects static HTML returned by the initial HTTP response. If your website is a single-page application (SPA) that renders links dynamically via client-side JavaScript, a static scraper will miss critical edges, producing a fragmented, inaccurate graph.

To ensure all edges are captured:

  1. Pre-render HTML Server-Side: Guarantee that all <a href="..."> elements exist in the raw server-rendered response before client hydration.
  2. Execute Headless Crawls: Use headless browser frameworks (such as Playwright or Puppeteer) to allow JavaScript execution before extracting the DOM.

To understand how Googlebot parses dynamic links, consult our guide on JavaScript links crawl discovery and Googlebot onclick issues.


Building custom Python crawl scripts and NetworkX pipelines provides deep insights, but requires significant maintenance as site scale increases. Maintaining scripts, managing proxy rotations, and rendering complex single-page applications at scale can quickly strain engineering resources.

The BugViso auditing platform automates internal link graph analysis at enterprise scale. Utilizing an ultra-fast headless crawling engine powered by Lightpanda and Playwright with distributed Redis job queues, BugViso:

  1. Renders and Maps Every Dynamic Edge: Executes complete client-side JavaScript to capture dynamic DOM links, pagination triggers, and interactive navigational elements.
  2. Calculates Real-Time Centrality Scores: Computes In-Degree, Out-Degree, Betweenness, and Internal PageRank across tens of thousands of URLs automatically.
  3. Pinpoints Structural Islands and Orphan Traps: Flags disconnected subgraphs, weakly connected components, and buried pages residing at Depth 4+.
  4. Detects Equity Leaks: Correlates link graph edges with HTTP status codes to catch redirect chains and 404 dead ends that bleed link authority. To trace redirect hops, review our guide on redirect chain audits and link equity drains.

Transforming your site's internal links into an interactive directed graph replaces intuitive guesswork with empirical mathematical architecture, revealing precisely how authority flows through your digital ecosystem.

Generate your website's complete internal link graph and diagnose structural bottlenecks by launching a comprehensive crawl with BugViso.

Found this useful? Share it.

See where your site stands

Run a free BugViso audit for SEO, speed, accessibility and AI search readiness — with fixes you can ship today.