SaaS Technical SEO Checklist: 20 Checks for Software Sites

A SaaS technical SEO checklist with 20 checks: rendering, pricing schema, app and docs subdomains, AI crawlers. Data from 283 SaaS sites and a checker script.

BugViso

16 min read

A SaaS technical SEO checklist covers what generic checklists miss: whether marketing pages work without JavaScript, whether the pricing page is indexable and described with commerce schema, whether the app, login and docs hosts are deliberately in or out of the index, and whether AI crawlers can read the pages you want cited. Get those right first; the classic on-page items come after.

To see where SaaS sites actually stand, on 10 October 2026 we checked 283 SaaS-style homepages — sites linking to a pricing or plans page and using sign-up, free-trial, demo or log-in wording — drawn at random from the Tranco top 200,000. Rendering was mostly fine: the median homepage had 1,295 words in its raw HTML. The gaps were elsewhere: 53.1% of pricing pages showed a price but had no commerce schema, 67.5% of reachable login pages were indexable, and only 12.4% of homepages described the product as a SoftwareApplication.

Below are the 20 checks in three tiers, the data for each, and a script that runs the SaaS-specific checks against your domain.


Methodology

ItemDetail
Sampling frameTranco list 94GG2, ranks 1,001–200,000; 6,000 domains drawn at random (seed 20261010), 3,416 reachable HTML homepages
SaaS-style filterHomepage links to /pricing, /plans or similar and contains sign-up / free trial / get started / demo / log in wording → 283 sites
FetchingPlain HTTP client, no JavaScript: what non-rendering crawlers and most AI crawlers receive
Pages checkedHomepage, the linked pricing page, the linked login page (and its host's robots.txt), the linked docs page, robots.txt, llms.txt
Date10 October 2026, single vantage point

The filter is a heuristic, so the sample includes some subscription and membership sites alongside classic B2B software. That's representative of what "SaaS-style" means on the open web, but treat the percentages as indicative.


The Checklist at a Glance

#CheckTierOur sample
1Marketing pages render meaningful HTML without JavaScriptP098.9% had ≥150 words in raw HTML
2App, login and account hosts kept out of the indexP067.5% of login pages indexable
3No accidental Disallow: / or noindex on marketing hostsP07 sites blocked all crawlers in robots.txt
4Pricing page exists, returns 200, is indexableP075.3% linked a working pricing page; 0.9% noindexed it
5Pricing page canonical is self-referencingP068.5%
6One canonical host (www vs apex, trailing slash)P0see our migration data
7XML sitemap declared in robots.txtP081.0%
8Price visible in HTML (not only after JS or "contact sales")P170.4%
9Commerce schema on the pricing pageP119.7%
10SoftwareApplication (or WebApplication) schema for the productP112.4% of homepages
11Organization schema with logo and sameAsP154.1%
12Docs crawlable and linked from marketing pagesP157.8% of docs on a subdomain; 6.1% near-empty without JS
13Blog in a subfolder, not a separate subdomainP18.1% on a subdomain
14Integration, comparison and use-case pages have unique contentP1see our programmatic SEO data
15Gated content has an indexable summary pageP2not sampled
16AI crawler policy is a deliberate choiceP289% didn't name GPTBot at all
17llms.txt (optional)P231.8%
18Changelog and status pages are noindexed or consolidatedP2not sampled
19Signup and onboarding flows excluded from crawlP2not sampled
20Hreflang correct for localized marketing sitesP2see our hreflang data

Tier P0: Crawlability and Index Control

1. Marketing pages render without JavaScript

Only 1.1% of the SaaS homepages had fewer than 150 words in their raw HTML; server rendering and static generation have become the default for marketing sites. The risk has moved to sub-sections: pricing toggles, comparison tables, docs and changelogs built as client-side apps. Google renders JavaScript (with a delay, as its JavaScript SEO basics explain); many AI crawlers don't. Check each template with JavaScript disabled, and read our guide on how Google renders single-page apps for the details.

bash
# Words a non-rendering crawler sees on your pricing page
curl -s https://example.com/pricing | sed -e 's/<script.*<\/script>//g' -e 's/<[^>]*>/ /g' | wc -w

2. Keep the app out of the index

Of 151 reachable login pages, 67.5% were indexable: no noindex, and the host's robots.txt didn't block it. Most were on an app subdomain (58.9%). An indexable login page is mostly noise, but the same configuration often exposes worse: share links, workspace pages, invitation URLs and error pages that should never rank.

html
<!-- ✅ On every page of the app host (login, signup, workspace) -->
<meta name="robots" content="noindex, nofollow">
nginx
# ✅ Or as a header for the whole app host (covers JSON, PDFs and error pages too)
add_header X-Robots-Tag "noindex, nofollow" always;

Use noindex (or the header; see Google's robots meta tag reference) rather than only Disallow: / in the app host's robots.txt: a blocked URL can still be indexed from links, without Google ever seeing a noindex. 25.8% of the subdomain logins we found relied on a robots.txt block.

3. No accidental site-wide blocks

7 of the 273 sites with a robots.txt disallowed everything for all user agents. On a marketing host that's almost always a staging config that shipped. Check robots.txt and the homepage's robots meta after every deploy.

4–5. A pricing page that can rank

Pricing pages attract high-intent queries ("[product] pricing", "[product] cost"). 75.3% of sites linked a pricing page that returned 200, and just 0.9% of those noindexed it — good. But only 68.5% had a canonical pointing to itself; the rest had none or pointed elsewhere, often to the homepage or a localized variant. A pricing page that canonicalizes away won't rank for its own queries.

6. One canonical host

SaaS sites accumulate hosts: www, apex, app, docs, help, status, regional domains. Every marketing URL should resolve to one host and one trailing-slash form with a single 301. Our website migration SEO checklist has the redirect data from 219 sites.

7. Sitemap declared

81.0% of SaaS robots.txt files declared a sitemap. Include marketing pages, blog, docs and integration pages; exclude app routes, signup steps and parameter URLs.


Tier P1: Pricing, Schema and Content Architecture

8–9. Price in HTML, described with schema

70.4% of pricing pages showed a price in the raw HTML, but only 19.7% carried commerce schema (Product, SoftwareApplication or Offer). That left 53.1% with a visible price and nothing machine-readable describing it. FAQPage schema (19.7%) was as common as product schema — useful for AI answers, but not a substitute.

json
{
  "@context": "https://schema.org",
  "@type": "SoftwareApplication",
  "name": "ExampleApp",
  "applicationCategory": "BusinessApplication",
  "operatingSystem": "Web",
  "offers": [
    {"@type": "Offer", "name": "Starter", "price": "19.00", "priceCurrency": "USD",
     "url": "https://example.com/pricing#starter"},
    {"@type": "Offer", "name": "Team", "price": "49.00", "priceCurrency": "USD",
     "url": "https://example.com/pricing#team"}
  ]
}

Follow Google's software app structured data guidelines, and only mark up prices that are visible on the page and current; schema that disagrees with the page is worse than none. Don't add aggregateRating unless the ratings are real and shown — Google's structured-data policies treat invented ratings as spam.

10–11. Product and organization entities

Only 12.4% of homepages declared a SoftwareApplication/WebApplication, while 54.1% had Organization and 33.6% had no JSON-LD at all. Entity markup is how search engines and AI systems connect your brand, product name and category. Our AI search optimization guide for B2B SaaS covers entity consistency in depth.

12. Docs that can be crawled

57.8% of the docs links we followed went to a separate subdomain, and 6.1% of docs pages had under 100 words without JavaScript — client-rendered docs that non-rendering crawlers see as empty. Docs answer the long-tail "how do I…" queries that bring evaluators; make sure they render server-side, link back to marketing pages, and are in a sitemap.

13. Blog in a subfolder

Only 8.1% of SaaS blogs lived on a subdomain. Subfolders (/blog/) keep internal linking and topical signals on the main host and are simpler to maintain; if yours is on a subdomain, link it heavily from the main site or plan a migration.

14. Integration, comparison and use-case pages

These are SaaS's programmatic pages: one template, hundreds of URLs. They earn traffic only if each page says something specific. We measured how similar real template families are in our guide to programmatic SEO without thin or duplicate pages.


Tier P2: AI Crawlers, Gated Content and Housekeeping

15. Gated content

Whitepapers and reports behind a form can't rank. Publish an indexable summary page with the key findings and a form, rather than a bare form.

16. AI crawler policy

Of 273 SaaS robots.txt files, 244 (89%) didn't mention GPTBot at all — they follow the * rules by default. 8 blocked GPTBot, 7 blocked Google-Extended (Gemini training) and 10 blocked CCBot; very few blocked the search-time crawlers (OAI-SearchBot 1, PerplexityBot 1). OpenAI documents its crawlers and what each one does on its bots page. Blocking training bots while allowing search bots is a valid policy — but make it a decision, not an accident. Our guide on blocking AI training while allowing AI search has the exact rules.

17. llms.txt (optional)

31.8% of SaaS sites served a real llms.txt (a Markdown file starting with a heading — we excluded servers that answer 200 for any path). It's an emerging convention with no confirmed effect on Google rankings; publish one if it's cheap, and keep it accurate.

18–20. Housekeeping

Noindex thin changelog-per-release and status pages (or consolidate them), exclude signup and onboarding steps from crawling, and validate hreflang on localized marketing sites — see our hreflang errors guide.


Run the SaaS Checks on Your Domain

This script fetches your homepage, pricing page, login page, robots.txt and llms.txt without JavaScript and reports the SaaS-specific checks above. Standard library only.

python
#!/usr/bin/env python3
"""saas_seo_check.py: the SaaS-specific technical SEO checks, from the raw HTML a crawler sees.

Usage:
    python3 saas_seo_check.py example.com [--pricing /pricing] [--login https://app.example.com/login]

Standard library only, no JavaScript execution (that is what non-rendering crawlers and most AI
crawlers get). Checks: homepage text in raw HTML; pricing page status, indexability, canonical,
visible price and commerce schema; login/app host indexability; robots.txt sitemap + AI crawler
rules; a real llms.txt (Markdown heading, not a catch-all 200).
"""
import argparse, json, re, ssl, sys, urllib.error, urllib.request
from urllib.parse import urljoin, urlsplit

UA = "Mozilla/5.0 (compatible; saas-seo-check/1.0)"
CTX = ssl.create_default_context()
COMMERCE = {"Product", "SoftwareApplication", "WebApplication", "MobileApplication", "Offer", "AggregateOffer"}
AI_BOTS = ["GPTBot", "OAI-SearchBot", "ClaudeBot", "PerplexityBot", "Google-Extended"]


def get(url):
    req = urllib.request.Request(url, headers={"User-Agent": UA})
    try:
        with urllib.request.urlopen(req, timeout=20, context=CTX) as r:
            return r.status, r.geturl(), r.read(2_000_000).decode("utf-8", "replace"), dict(r.headers)
    except urllib.error.HTTPError as e:
        return e.code, url, "", dict(e.headers or {})
    except Exception as e:
        return type(e).__name__, url, "", {}


def words(html):
    return len(re.sub(r"<script.*?</script>|<style.*?</style>|<[^>]+>", " ", html, flags=re.S | re.I).split())


def ld_types(html):
    out = set()
    for block in re.findall(r"<script[^>]+application/ld\+json[^>]*>(.*?)</script>", html, re.I | re.S):
        try:
            stack = [json.loads(block.strip())]
        except ValueError:
            out.add("INVALID JSON-LD"); continue
        while stack:
            o = stack.pop()
            if isinstance(o, list):
                stack.extend(o)
            elif isinstance(o, dict):
                t = o.get("@type")
                out.update(t if isinstance(t, list) else [t] if t else [])
                stack.extend(v for v in o.values() if isinstance(v, (dict, list)))
    return out


def noindex(html, headers):
    m = re.search(r"<meta[^>]+name=[\"']robots[\"'][^>]*content=[\"']([^\"']+)", html, re.I)
    return "noindex" in ((m.group(1) if m else "") + headers.get("X-Robots-Tag", "")).lower()


def line(ok, label, detail=""):
    print(f"  {'✓' if ok else ('✗' if ok is False else '•')} {label:38s} {detail}")
    return ok is not False


def main():
    ap = argparse.ArgumentParser()
    ap.add_argument("domain")
    ap.add_argument("--pricing", help="pricing path or URL (auto-detected if omitted)")
    ap.add_argument("--login", help="login/app URL (auto-detected if omitted)")
    a = ap.parse_args()
    home = a.domain if a.domain.startswith("http") else f"https://{a.domain}/"
    st, home, html, hdr = get(home)
    host = urlsplit(home).hostname
    root = host.removeprefix("www.")
    links = [(urljoin(home, h), re.sub(r"<[^>]+>", " ", t).strip().lower())
             for h, t in re.findall(r"<a\b[^>]*href\s*=\s*[\"']([^\"'#]+)[\"'][^>]*>(.*?)</a>", html, re.I | re.S)]
    ok = True
    print(f"\n{home}")
    w = words(html)
    ok &= line(st == 200 and w >= 150, "homepage text in raw HTML", f"HTTP {st}, {w} words without JavaScript")
    types = ld_types(html)
    ok &= line(bool(types & {"Organization", "SoftwareApplication", "WebApplication"}) or None, "homepage schema",
               ", ".join(sorted(types)) or "no JSON-LD")
    pricing = urljoin(home, a.pricing) if a.pricing else next(
        (u for u, _ in links if re.search(r"/(pricing|plans)/?$", urlsplit(u).path, re.I)), None)
    if pricing:
        pst, pf, ph, phdr = get(pricing)
        pt = ld_types(ph)
        canon = re.search(r"<link[^>]+rel=[\"']canonical[\"'][^>]*href=[\"']([^\"']+)", ph, re.I)
        price = re.search(r"[$€£₹]\s?\d{1,4}(?:[.,]\d{2})?|\d{1,4}(?:[.,]\d{2})?\s?(?:USD|EUR|GBP)",
                          re.sub(r"<script.*?</script>", " ", ph, flags=re.S | re.I))
        ok &= line(pst == 200 and not noindex(ph, phdr), "pricing page indexable", f"{pricing} (HTTP {pst})")
        ok &= line(bool(canon) and urljoin(pf, canon.group(1)).rstrip("/") == pf.rstrip("/"), "pricing canonical is self",
                   canon.group(1) if canon else "missing")
        ok &= line(bool(price), "a price is visible in raw HTML", price.group(0) if price else "none found (rendered by JS, or 'contact sales')")
        ok &= line(bool(pt & COMMERCE) if price else None, "commerce schema on pricing",
                   ", ".join(sorted(pt & COMMERCE)) or "none (Product/SoftwareApplication + Offer)")
    else:
        line(None, "pricing page", "no /pricing or /plans link found on the homepage")
    login = a.login or next((u for u, t in links if re.search(r"\b(log ?in|sign ?in)\b", t)), None)
    if login:
        lst, lf, lh, lhdr = get(login)
        lhost = urlsplit(lf).hostname
        rst, _, rtxt, _ = get(f"https://{lhost}/robots.txt")
        blocked = bool(re.search(r"(?ims)^user-agent:\s*\*\s*$(?:(?!^user-agent:).)*?^disallow:\s*/\s*$", rtxt or ""))
        ok &= line(noindex(lh, lhdr) or (blocked and lhost != host), "login/app page kept out of the index",
                   f"{lf}: {'noindex' if noindex(lh, lhdr) else 'robots.txt blocks host' if blocked else 'indexable'}")
    rst, _, rtxt, _ = get(f"https://{host}/robots.txt")
    ok &= line(rst == 200 and bool(re.search(r"(?im)^sitemap:", rtxt)), "robots.txt declares a sitemap", f"HTTP {rst}")
    groups = re.split(r"(?im)^(?=user-agent:)", rtxt or "")
    for bot in AI_BOTS:
        g = next((g for g in groups if re.search(rf"(?im)^user-agent:\s*{re.escape(bot)}\s*$", g)), None)
        state = "unlisted (follows *)" if g is None else ("blocked" if re.search(r"(?im)^disallow:\s*/\s*$", g) else "allowed")
        line(None, f"robots.txt {bot}", state)
    lst, _, ltxt, lhdr = get(f"https://{host}/llms.txt")
    cst, _, _, chdr = get(f"https://{host}/bv-not-a-real-file.txt")
    real = lst == 200 and ltxt.lstrip().startswith("#") and not (cst == 200 and "html" not in chdr.get("Content-Type", ""))
    line(real or None, "llms.txt (optional)", "present" if real else "not present (optional; no ranking effect in Google)")
    return 0 if ok else 1


if __name__ == "__main__":
    sys.exit(main())

Real output for a large productivity SaaS (10 October 2026, name removed):

text
https://productivity-app.example/
  ✓ homepage text in raw HTML              HTTP 200, 436 words without JavaScript
  • homepage schema                        no JSON-LD
  ✓ pricing page indexable                 https://productivity-app.example/pricing (HTTP 200)
  ✓ pricing canonical is self              https://productivity-app.example/pricing
  ✓ a price is visible in raw HTML         $0
  ✗ commerce schema on pricing             none (Product/SoftwareApplication + Offer)
  ✗ login/app page kept out of the index   https://app.productivity-app.example/login: indexable
  ✓ robots.txt declares a sitemap          HTTP 200
  • robots.txt GPTBot                      unlisted (follows *)
  • robots.txt OAI-SearchBot               unlisted (follows *)
  • robots.txt ClaudeBot                   unlisted (follows *)
  • robots.txt PerplexityBot               unlisted (follows *)
  • robots.txt Google-Extended             unlisted (follows *)
  ✓ llms.txt (optional)                    present

Even a category leader shows the two most common gaps from our data: a visible price with no commerce schema, and an indexable app login. Both are an afternoon's work.


How BugViso Audits SaaS Sites

A BugViso scan covers the SaaS-specific checks alongside the standard audit:

  • Pricing schema validation: the structured-data engine detects an active pricing section on the page and flags it when there's no matching commerce schema (Product, SoftwareApplication, Offer), and validates the required properties of the schema you do have (offers, applicationCategory, price, priceCurrency).
  • Rendering: pages are read static-first and JavaScript-built pages get a Chromium render, so you can see whether content depends on JavaScript; React and Next.js hydration errors are reported separately.
  • AI search readiness (GEO): robots.txt parsing for AI crawlers, llms.txt validation, content extractability and indexability, scored 0–100.
  • Index control: noindex, X-Robots-Tag and canonical analysis, plus a host-consistency check that catches www/apex canonicals pointing at the wrong host.

What a scan doesn't do is follow your login link onto the app host — check that host with the script above. You can run a free SaaS audit with BugViso; the AI-readiness checks are detailed on the GEO and AI search page.


High-Risk Oversights

  • A staging robots.txt on production. Seven sites in our sample were blocking every crawler. Add a deploy check.
  • Pricing behind a toggle that only JavaScript renders. Make sure at least one plan's price is in the HTML.
  • Schema that contradicts the page. Update Offer prices whenever pricing changes, or generate them from the same source as the page.
  • App share links in search results. Put noindex on the whole app host via a header, not just the login page.
  • Docs on a JavaScript-only portal. Long-tail queries are often a SaaS site's best traffic; client-only docs throw it away.

FAQ

What is technical SEO for SaaS?

The same fundamentals as any site — crawlability, indexing, speed, structured data — plus SaaS-specific concerns: separating the app from marketing pages, pricing-page optimization, docs and integration pages at scale, and AI crawler access.

Should a SaaS pricing page have schema markup?

Yes, if it shows prices. Use SoftwareApplication or Product with an Offer for each plan, matching the visible prices. In our sample only 19.7% of pricing pages had any commerce schema.

Should the app subdomain be indexed?

No. Add noindex (ideally an X-Robots-Tag header on the whole app host). Relying only on a robots.txt Disallow can still leave URLs indexed from external links.

Is a blog subdomain bad for SaaS SEO?

It isn't a penalty, but subfolders are simpler and keep internal linking on one host. In our sample, only 8.1% of SaaS blogs were on a subdomain.


Conclusion

SaaS marketing sites mostly render fine now; the gaps are in index control and machine-readable pricing, so keep the app out of the index, give the pricing page commerce schema, and make your AI crawler policy deliberate — then let a BugViso scan watch the rest.

Found this useful? Share it.

See where your site stands

Run a free BugViso audit for SEO, speed, accessibility and AI search readiness — with fixes you can ship today.