XML Sitemap Engineering: Dynamic Generation and Best Practices

Master XML sitemap best practices and dynamic generation. Build auto-generating sitemaps with accurate lastmod, sub-sitemap splitting, and validation for sites of any scale.

BugViso

15 min read

An XML sitemap is a structured file that tells search engines which URLs on your site exist, when they were last modified, and how frequently they change. For sites with under 500 pages and clean internal linking, a sitemap is a helpful supplement. For enterprise sites with 50,000+ URLs, dynamic content, and complex URL architectures, the sitemap becomes a critical indexation signal — the primary mechanism through which Googlebot discovers new pages, prioritizes recrawl frequency, and differentiates fresh content from stale. A well-engineered sitemap can reduce time-to-indexation for new content from weeks to hours.

This guide covers production-grade sitemap engineering: the XML schema spec, dynamic generation patterns, sub-sitemap splitting for scale, lastmod accuracy, error prevention, and validation workflows.

The XML Sitemap Schema Specification

The Sitemaps protocol defines a strict XML schema. Every sitemap must conform to this structure or search engines will reject it.

Basic Sitemap Structure

xml
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://example.com/</loc>
    <lastmod>2026-09-15T08:30:00+00:00</lastmod>
    <changefreq>daily</changefreq>
    <priority>1.0</priority>
  </url>
  <url>
    <loc>https://example.com/blog/technical-seo-guide/</loc>
    <lastmod>2026-09-10T14:00:00+00:00</lastmod>
    <changefreq>weekly</changefreq>
    <priority>0.8</priority>
  </url>
</urlset>

Element Reference

ElementRequiredDescriptionGoogle's Behavior
<loc>✅ YesAbsolute URL of the pageUsed for URL discovery
<lastmod>❌ OptionalISO 8601 date of last meaningful changeUsed if accurate; ignored if unreliable
<changefreq>❌ OptionalHint: always, hourly, daily, weekly, monthly, yearly, neverIgnored by Google (confirmed by Google)
<priority>❌ Optional0.0–1.0 relative importance within the siteIgnored by Google (confirmed by Google)

💡 The practical takeaway: Google ignores changefreq and priority entirely. The only elements that matter are <loc> (the URL) and <lastmod> (the modification date). Every engineering effort should focus on accurate lastmod timestamps and correct <loc> URLs — not on tuning priority values.

Dynamic Sitemap Generation: Production Patterns

Static sitemaps (manually maintained XML files) break as sites scale. Dynamic generation ensures the sitemap always reflects the current state of your content.

Pattern 1: Next.js App Router Dynamic Sitemap

typescript
// app/sitemap.ts — Next.js 14/15 built-in sitemap generation
import { MetadataRoute } from 'next';

// Database or CMS query for all published content
async function getAllPosts() {
  const posts = await db.query(`
    SELECT slug, updated_at FROM posts
    WHERE status = 'published'
    ORDER BY updated_at DESC
  `);
  return posts;
}

async function getAllProducts() {
  const products = await db.query(`
    SELECT slug, updated_at FROM products
    WHERE active = true
    ORDER BY updated_at DESC
  `);
  return products;
}

export default async function sitemap(): Promise<MetadataRoute.Sitemap> {
  const posts = await getAllPosts();
  const products = await getAllProducts();

  const staticPages: MetadataRoute.Sitemap = [
    {
      url: 'https://example.com',
      lastModified: new Date(),
      changeFrequency: 'daily',
      priority: 1,
    },
    {
      url: 'https://example.com/about',
      lastModified: new Date('2026-01-15'),
      changeFrequency: 'monthly',
      priority: 0.5,
    },
  ];

  const blogPages: MetadataRoute.Sitemap = posts.map((post) => ({
    url: `https://example.com/blog/${post.slug}`,
    lastModified: new Date(post.updated_at),
    changeFrequency: 'weekly' as const,
    priority: 0.7,
  }));

  const productPages: MetadataRoute.Sitemap = products.map((product) => ({
    url: `https://example.com/products/${product.slug}`,
    lastModified: new Date(product.updated_at),
    changeFrequency: 'daily' as const,
    priority: 0.8,
  }));

  return [...staticPages, ...blogPages, ...productPages];
}

Pattern 2: Python/Django Dynamic Sitemap

python
# sitemaps.py — Django sitemap framework
from django.contrib.sitemaps import Sitemap
from blog.models import Post
from products.models import Product

class PostSitemap(Sitemap):
    changefreq = 'weekly'
    priority = 0.7

    def items(self):
        return Post.objects.filter(
            status='published'
        ).order_by('-updated_at')

    def lastmod(self, obj):
        return obj.updated_at

    def location(self, obj):
        return f'/blog/{obj.slug}/'

class ProductSitemap(Sitemap):
    changefreq = 'daily'
    priority = 0.8

    def items(self):
        return Product.objects.filter(
            active=True
        ).order_by('-updated_at')

    def lastmod(self, obj):
        return obj.updated_at

    def location(self, obj):
        return f'/products/{obj.slug}/'

# urls.py
from django.contrib.sitemaps.views import sitemap

sitemaps = {
    'posts': PostSitemap,
    'products': ProductSitemap,
}

urlpatterns = [
    path('sitemap.xml', sitemap, {'sitemaps': sitemaps},
         name='django.contrib.sitemaps.views.sitemap'),
]

Pattern 3: Node.js/Express Programmatic Generation

typescript
// routes/sitemap.ts — Express.js dynamic sitemap
import { Router, Request, Response } from 'express';
import { SitemapStream, streamToPromise } from 'sitemap';
import { createGzip } from 'zlib';

const router = Router();

router.get('/sitemap.xml', async (req: Request, res: Response) => {
  res.header('Content-Type', 'application/xml');
  res.header('Content-Encoding', 'gzip');

  const smStream = new SitemapStream({
    hostname: 'https://example.com',
  });
  const pipeline = smStream.pipe(createGzip());

  // Static pages
  smStream.write({ url: '/', lastmod: new Date().toISOString(), priority: 1.0 });
  smStream.write({ url: '/about/', lastmod: '2026-01-15', priority: 0.5 });

  // Dynamic pages from database
  const posts = await db.getAllPublishedPosts();
  for (const post of posts) {
    smStream.write({
      url: `/blog/${post.slug}/`,
      lastmod: post.updatedAt.toISOString(),
      priority: 0.7,
    });
  }

  const products = await db.getAllActiveProducts();
  for (const product of products) {
    smStream.write({
      url: `/products/${product.slug}/`,
      lastmod: product.updatedAt.toISOString(),
      priority: 0.8,
    });
  }

  smStream.end();
  const data = await streamToPromise(pipeline);
  res.send(data);
});

export default router;

Sub-Sitemap Splitting for Scale

The Sitemaps protocol limits each sitemap file to 50,000 URLs and 50 MB uncompressed. Sites exceeding either limit must split into multiple sub-sitemaps referenced by a sitemap index file.

Sitemap Index Structure

xml
<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <sitemap>
    <loc>https://example.com/sitemaps/products-1.xml</loc>
    <lastmod>2026-09-15T08:30:00+00:00</lastmod>
  </sitemap>
  <sitemap>
    <loc>https://example.com/sitemaps/products-2.xml</loc>
    <lastmod>2026-09-14T12:00:00+00:00</lastmod>
  </sitemap>
  <sitemap>
    <loc>https://example.com/sitemaps/blog.xml</loc>
    <lastmod>2026-09-15T10:15:00+00:00</lastmod>
  </sitemap>
  <sitemap>
    <loc>https://example.com/sitemaps/categories.xml</loc>
    <lastmod>2026-09-01T00:00:00+00:00</lastmod>
  </sitemap>
</sitemapindex>

Splitting Strategy

Split by content type, not by arbitrary number. This keeps each sub-sitemap thematically coherent and makes the lastmod on the index file meaningful:

Sub-SitemapContentsTypical SizeUpdate Frequency
products-1.xml ... products-N.xmlProduct pages (batched by 10K)10,000 URLs eachDaily
blog.xmlBlog posts500–5,000 URLsWeekly
categories.xmlCategory and subcategory pages50–500 URLsMonthly
static.xmlAbout, contact, legal, features pages10–50 URLsRarely
python
# Python — Generate sub-sitemaps with batching
import math
from lxml import etree

SITEMAP_NS = 'http://www.sitemaps.org/schemas/sitemap/0.9'
BATCH_SIZE = 10_000

def generate_sub_sitemaps(products: list, output_dir: str) -> list:
    """Generate batched product sub-sitemaps, return sub-sitemap URLs."""
    sub_sitemaps = []
    total_batches = math.ceil(len(products) / BATCH_SIZE)

    for batch_num in range(total_batches):
        batch = products[batch_num * BATCH_SIZE:(batch_num + 1) * BATCH_SIZE]
        root = etree.Element('urlset', xmlns=SITEMAP_NS)

        for product in batch:
            url_elem = etree.SubElement(root, 'url')
            loc = etree.SubElement(url_elem, 'loc')
            loc.text = f'https://example.com/products/{product["slug"]}/'
            lastmod = etree.SubElement(url_elem, 'lastmod')
            lastmod.text = product['updated_at'].isoformat()

        filename = f'products-{batch_num + 1}.xml'
        filepath = f'{output_dir}/{filename}'
        tree = etree.ElementTree(root)
        tree.write(filepath, xml_declaration=True, encoding='UTF-8',
                   pretty_print=True)
        sub_sitemaps.append(f'https://example.com/sitemaps/{filename}')

    return sub_sitemaps

The lastmod Accuracy Problem

lastmod is the single most impactful sitemap element after <loc>. Google uses it to prioritize recrawl scheduling — pages with recent lastmod values get recrawled sooner. But Google only trusts lastmod if it is accurate.

What Google Says About lastmod

Google's Gary Illyes has stated: "We will use lastmod if we find it to be reliably accurate for a given site. If we detect that lastmod values don't correspond to actual content changes, we'll stop trusting them."

Common lastmod Mistakes

xml
<!-- ❌ Mistake: lastmod set to build/deploy time for all pages -->
<!-- Every page shows today's date, regardless of actual content changes -->
<url>
  <loc>https://example.com/about/</loc>
  <lastmod>2026-09-15T08:30:00+00:00</lastmod> <!-- not changed since 2025 -->
</url>

<!-- ❌ Mistake: lastmod set to current timestamp on every request -->
<!-- Dynamic generation that always returns "now" -->

<!-- ✅ Correct: lastmod reflects the actual last content edit -->
<url>
  <loc>https://example.com/about/</loc>
  <lastmod>2025-03-20T14:00:00+00:00</lastmod> <!-- real last edit date -->
</url>

Engineering Accurate lastmod

typescript
// Track actual content changes, not deployment times
interface ContentRecord {
  slug: string;
  content_hash: string;   // SHA-256 of the rendered content
  last_content_change: Date; // Updated ONLY when content_hash changes
  last_deployed: Date;     // Updated on every deployment
}

// In your sitemap generator, use last_content_change, NOT last_deployed
function getLastmod(record: ContentRecord): string {
  return record.last_content_change.toISOString();
}
python
# Django — Track content changes with a hash field
import hashlib
from django.db import models

class Post(models.Model):
    slug = models.SlugField(unique=True)
    body = models.TextField()
    content_hash = models.CharField(max_length=64, blank=True)
    content_changed_at = models.DateTimeField(auto_now_add=True)
    updated_at = models.DateTimeField(auto_now=True)

    def save(self, *args, **kwargs):
        new_hash = hashlib.sha256(self.body.encode()).hexdigest()
        if new_hash != self.content_hash:
            self.content_hash = new_hash
            self.content_changed_at = timezone.now()
        super().save(*args, **kwargs)

Sitemap Validation and Error Prevention

Invalid sitemaps are silently ignored by search engines. A syntax error in your sitemap can prevent discovery of thousands of URLs without any visible warning.

Common Sitemap Errors

ErrorCauseFix
Invalid XMLUnescaped & in URLs (e.g., ?a=1&b=2)Use &amp; in XML: ?a=1&amp;b=2
Wrong namespaceMissing or incorrect xmlns attributeUse exactly http://www.sitemaps.org/schemas/sitemap/0.9
Relative URLs<loc> contains /blog/post/ instead of full URLAlways use absolute URLs with protocol
Non-canonical URLsSitemap lists http:// but canonical is https://Match sitemap URLs to canonical URLs exactly
Noindexed URLsSitemap includes pages with noindex directiveRemove noindexed pages from sitemap
404 URLsSitemap lists URLs that return 404Validate all URLs return 200 before adding
Oversized fileSitemap exceeds 50MB or 50,000 URLsSplit into sub-sitemaps with index

Validation Script

bash
# Validate sitemap XML syntax
xmllint --noout --schema \
  https://www.sitemaps.org/schemas/sitemap/0.9/sitemap.xsd \
  sitemap.xml

# Check that all sitemap URLs return 200
grep -oP '<loc>\K[^<]+' sitemap.xml | while read url; do
  status=$(curl -s -o /dev/null -w "%{http_code}" "$url")
  if [ "$status" != "200" ]; then
    echo "ERROR: $url returned $status"
  fi
done

# Check for URLs with unescaped ampersands
grep -n '&[^a]' sitemap.xml | grep -v '&amp;' | head -10

Automated CI/CD Validation

yaml
# GitHub Actions — Validate sitemap on every deployment
name: Validate Sitemap
on:
  push:
    branches: [main]

jobs:
  validate:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Build site
        run: npm run build
      - name: Validate sitemap XML
        run: |
          xmllint --noout dist/sitemap.xml
          echo "✅ Sitemap XML is valid"
      - name: Check URL count
        run: |
          count=$(grep -c '<loc>' dist/sitemap.xml)
          echo "Sitemap contains $count URLs"
          if [ "$count" -gt 50000 ]; then
            echo "❌ ERROR: Sitemap exceeds 50,000 URL limit"
            exit 1
          fi

Registering Your Sitemap with Search Engines

robots.txt Declaration

text
# robots.txt — Declare sitemap location
User-agent: *
Allow: /

Sitemap: https://example.com/sitemap.xml

Google Search Console Submission

Navigate to Sitemaps in Google Search Console, enter your sitemap URL, and click Submit. GSC will report:

  • Status: Success, Has errors, or Couldn't fetch
  • Discovered URLs: Number of URLs Google found in the sitemap
  • Indexed URLs: Number of URLs Google has indexed (usually lower than discovered)

Ping Endpoints (Post-Publish)

bash
# Notify Google of sitemap updates after publishing new content
curl "https://www.google.com/ping?sitemap=https://example.com/sitemap.xml"

# Note: Google has deprecated the ping endpoint for most use cases
# Submit via GSC API instead for reliable notification

How BugViso Validates Sitemap-to-Crawl Alignment

BugViso's multi-page site crawl discovers URLs via sitemap.xml — including sitemaps declared in robots.txt and one level of sitemap-index recursion — and from the entry page's rendered links. This dual-source discovery exactly mirrors how Googlebot finds pages: sitemap declarations combined with link-following.

The crawl engine compares the sitemap URL set against the link-discovered URL set, surfacing the alignment gaps:

  • URLs in sitemap but not linked — sitemap-only pages that may be orphan pages with no internal link equity.
  • URLs linked but not in sitemap — discoverable via navigation but missing from the sitemap, potentially reducing their crawl priority.
  • Canonical mismatches — pages whose <link rel="canonical"> URL differs from the URL listed in the sitemap, sending conflicting indexation signals.

The Duplicate Content Detection module identifies pages that should be consolidated or excluded from the sitemap — exact duplicates, near-duplicates (SimHash analysis), and duplicate titles/meta descriptions that indicate index bloat.

Run a free BugViso scan to validate your sitemap's URL coverage, canonical alignment, and duplicate content across every crawled page.

Frequently Asked Questions

How many URLs should be in my sitemap?

Include only URLs that you want indexed. A sitemap with 10,000 curated, high-quality URLs is more valuable than one with 500,000 URLs where 90% are thin or duplicate. Google processes your sitemap more efficiently when every URL in it is indexable, unique, and returns a 200 status code.

Does Google use sitemap priority and changefreq?

No. Google's official documentation and public statements from Google engineers confirm that priority and changefreq are ignored. The only sitemap elements Google actively uses are <loc> (the URL) and <lastmod> (the modification date, when accurate).

Should I gzip my sitemap?

Yes, for sitemaps larger than 1 MB. Google supports .xml.gz compressed sitemaps. Compression reduces bandwidth and download time for both search engine crawlers and your server. Reference your gzipped sitemaps in robots.txt and sitemap index files with the .xml.gz extension.

How quickly does Google process a new sitemap submission?

Google typically fetches a newly submitted sitemap within minutes to hours. Processing the URLs within it (crawling and potentially indexing) takes longer — typically 1–14 days depending on your site's crawl budget, the volume of new URLs, and Google's assessment of your site's authority. Pages with high-quality content and strong internal links are processed fastest.

Should I include images and videos in my sitemap?

Yes, if you want them to appear in Google Image Search or Google Video Search. Use the image and video sitemap extensions. For standard web pages, the basic <urlset> schema is sufficient. Image and video sitemaps use separate namespace extensions defined at Google's sitemap documentation.

What happens if my sitemap contains 404 URLs?

Google will crawl the 404 URLs, discover they do not exist, and eventually drop them from its index. However, having many 404 URLs in your sitemap wastes crawl budget and signals poor site maintenance. Remove 404 URLs from your sitemap promptly. Validate all URLs return 200 before including them, and run automated checks on every deployment.

Conclusion

A well-engineered XML sitemap with accurate lastmod timestamps, proper sub-sitemap splitting, and rigorous validation is one of the highest-leverage technical SEO investments for sites at scale — and verifying that your sitemap's URL coverage aligns with your crawled link graph, canonical declarations, and duplicate content landscape is exactly the multi-dimensional analysis a free BugViso scan performs in a single automated pass.

Found this useful? Share it.

See where your site stands

Run a free BugViso audit for SEO, speed, accessibility and AI search readiness — with fixes you can ship today.