SEO Audit Process Template for Agencies: A Repeatable SOP
A seven-stage SEO audit process template for agencies, from intake to monitoring, with a pre-flight script and real issue rates from 219 sampled websites.
An agency SEO audit process template is a fixed sequence of seven stages that every audit follows, whoever runs it: intake → scan → triage → client report → fix tracking → re-audit → monitoring. Each stage has a defined input, a defined output and an owner. The template exists so that audit quality doesn't depend on which strategist picked up the ticket, and so the slow, mechanical parts get automated while human judgement goes where it changes the outcome.
This page is the SOP we'd hand a new account manager. Each stage has a checklist, what to automate, what to never automate, and the artefact it produces. It includes real data on what a first-pass audit typically finds, which helps you scope honestly, and a pre-flight script that catches the five problems that wreck an audit before it starts. For the commercial side, see our guides to pricing SEO audit services and scaling audits across 50+ clients.
The SOP at a Glance
| Stage | Input | Output (artefact) | Automate | Human judgement |
|---|---|---|---|---|
| 1. Intake | Signed scope, domain | Intake file + pre-flight table | Pre-flight checks | Goals, priorities, constraints |
| 2. Scan | Intake file | Raw audit (PDF + JSON + crawl CSV) | Crawl, render, all checks | Crawl scope, blocked-site handling |
| 3. Triage | Raw audit | Prioritised issue list | Grouping, counting, scoring | Business impact, false positives |
| 4. Report | Issue list | Client PDF + 1-page summary | Branding, formatting | Narrative, top 5, the ask |
| 5. Fix tracking | Approved issues | Tickets with acceptance criteria | Ticket creation from CSV | Sequencing, ownership |
| 6. Re-audit | Closed tickets | Before/after comparison | Re-running the same scan | Verifying the fix did what it should |
| 7. Monitoring | Baseline | Monthly health report | Scheduled scans | Spotting regressions that matter |
Stage 1: Intake (Before Anyone Runs a Scan)
Most bad audits are bad from the first hour: the wrong host gets crawled, staging is mistaken for production, a bot wall blocks the crawler, or nobody asked what the client actually sells.
Intake checklist:
- Business goals and the 3–5 money pages (the pages that generate leads or revenue)
- Markets and languages (hreflang scope) and the regions that need consent banners
- Platform and who can change it (CMS admin? developer retainer? locked SaaS builder?)
- Access: Search Console (at least restricted user), analytics read access, CMS or staging if fixes are in scope
- Known history: migrations, redesigns, traffic drops and their dates
- Bot protection or WAF on the site (allowlist the crawler, or expect a blocked scan)
- The pre-flight table below, pasted into the client file
The 60-second pre-flight script
This standard-library Python script checks the things that silently invalidate an audit: four host variants landing in different places, a Disallow: / left over from staging, a missing sitemap, a noindex homepage, a missing canonical and an expiring certificate. It prints Markdown you can paste straight into the intake document.
#!/usr/bin/env python3
"""intake_preflight.py: 60-second pre-flight for a new audit client. Prints a Markdown block for the client file.
Usage:
python3 intake_preflight.py example.com > intake-example.com.md
Standard library only. Checks the four host variants (http/https x www/apex) and where each
lands, robots.txt (and whether it blocks everything), sitemaps declared or at /sitemap.xml with
URL counts, homepage status, meta robots / X-Robots-Tag, canonical, TLS days left, a CMS hint
and analytics/tag-manager IDs. Anything marked ✗ goes to the top of the audit scope.
"""
import re, socket, ssl, sys, urllib.error, urllib.request
from datetime import datetime, timezone
from urllib.parse import urljoin, urlsplit
UA = "Mozilla/5.0 (compatible; intake-preflight/1.0)"
class NoRedirect(urllib.request.HTTPRedirectHandler):
def redirect_request(self, *a, **k):
return None
def hop(url, follow=True, limit=600_000):
opener = urllib.request.build_opener() if follow else urllib.request.build_opener(NoRedirect)
req = urllib.request.Request(url, headers={"User-Agent": UA})
try:
with opener.open(req, timeout=15) as r:
return r.status, r.geturl(), dict(r.headers), r.read(limit).decode("utf-8", "replace")
except urllib.error.HTTPError as e:
return e.code, e.headers.get("Location") or url, dict(e.headers or {}), ""
except Exception as e:
return None, url, {}, f"{type(e).__name__}: {e}"
def chain(url, max_hops=6):
hops = []
for _ in range(max_hops):
status, loc, headers, _ = hop(url, follow=False)
hops.append(f"{status or 'ERR'}")
if status in (301, 302, 303, 307, 308) and headers.get("Location"):
url = urljoin(url, headers["Location"])
continue
return url, hops
return url, hops + ["…"]
def tls_days(host):
try:
with socket.create_connection((host, 443), timeout=8) as raw:
with ssl.create_default_context().wrap_socket(raw, server_hostname=host) as s:
exp = datetime.strptime(s.getpeercert()["notAfter"], "%b %d %H:%M:%S %Y %Z").replace(tzinfo=timezone.utc)
return (exp - datetime.now(timezone.utc)).days
except Exception as e:
return f"error ({type(e).__name__})"
def main(domain):
domain = domain.removeprefix("https://").removeprefix("http://").strip("/").removeprefix("www.")
rows = []
add = lambda ok, check, detail: rows.append(("✓" if ok else "✗" if ok is False else "•", check, detail))
finals = {}
for v in (f"http://{domain}/", f"http://www.{domain}/", f"https://{domain}/", f"https://www.{domain}/"):
final, hops = chain(v)
finals[v] = final
add(None, f"`{v}`", f"→ `{final}` ({' → '.join(hops)})")
targets = {f for f in finals.values()}
add(len(targets) == 1 and next(iter(targets)).startswith("https://"), "One canonical host over HTTPS",
"all variants converge" if len(targets) == 1 else f"{len(targets)} different end points")
home = sorted(targets, key=lambda u: list(finals.values()).count(u))[-1]
status, final, headers, html = hop(home)
add(status == 200, "Homepage status", f"{status} at `{final}`")
xrt = headers.get("X-Robots-Tag", "")
mrobots = re.search(r'<meta[^>]+name=["\']robots["\'][^>]*content=["\']([^"\']+)', html, re.I)
noindex = "noindex" in (xrt + (mrobots.group(1) if mrobots else "")).lower()
add(not noindex, "Homepage indexable", f"meta robots: {mrobots.group(1) if mrobots else 'none'}; X-Robots-Tag: {xrt or 'none'}")
canon = re.search(r'<link[^>]+rel=["\']canonical["\'][^>]*href=["\']([^"\']+)', html, re.I)
add(bool(canon) and canon.group(1).rstrip("/") == final.rstrip("/"), "Homepage canonical", canon.group(1) if canon else "missing")
base = f"{urlsplit(final).scheme}://{urlsplit(final).netloc}"
rs, _, _, robots = hop(base + "/robots.txt")
block_all = bool(re.search(r"(?ims)^user-agent:\s*\*\s*$(?:(?!^user-agent:).)*?^disallow:\s*/\s*$", robots or ""))
add(rs == 200 and not block_all, "robots.txt", f"HTTP {rs}" + (" — BLOCKS ALL CRAWLING (Disallow: /)" if block_all else ""))
maps = re.findall(r"(?im)^sitemap:\s*(\S+)", robots or "") or [base + "/sitemap.xml"]
for sm in maps[:3]:
ss, _, _, xml = hop(sm, limit=5_000_000)
n_url, n_map = xml.count("<loc>"), xml.count("<sitemap>")
add(ss == 200 and (n_url > 0), f"Sitemap `{sm}`", f"HTTP {ss}; " + (f"index of {n_map} sitemaps" if n_map else f"{n_url} URLs"))
days = tls_days(urlsplit(final).hostname)
add(isinstance(days, int) and days >= 21, "TLS certificate", f"{days} days left" if isinstance(days, int) else days)
gen = re.search(r'<meta[^>]+name=["\']generator["\'][^>]*content=["\']([^"\']+)', html, re.I)
cms = gen.group(1) if gen else ("WordPress" if "/wp-content/" in html else "Shopify" if "cdn.shopify.com" in html
else "Webflow" if "webflow" in html.lower() else "unknown")
add(None, "CMS / platform hint", cms)
ids = sorted(set(re.findall(r"\b(GTM-[A-Z0-9]{4,9}|G-[A-Z0-9]{6,12}|UA-\d{4,10}-\d+)\b", html)))
add(None, "Tag manager / analytics IDs", ", ".join(ids) or "none found in HTML")
print(f"## Pre-flight: {domain} ({datetime.now(timezone.utc):%Y-%m-%d})\n")
print("| | Check | Result |\n| :---: | :--- | :--- |")
for mark, check, detail in rows:
print(f"| {mark} | {check} | {detail} |")
return all(m != "✗" for m, _, _ in rows)
if __name__ == "__main__":
if len(sys.argv) != 2:
sys.exit(__doc__)
sys.exit(0 if main(sys.argv[1]) else 1)Real output for a domain that fails several checks (9 October 2026):
## Pre-flight: example.com (2026-10-09)
| | Check | Result |
| :---: | :--- | :--- |
| • | `http://example.com/` | → `http://example.com/` (200) |
| • | `http://www.example.com/` | → `http://www.example.com/` (200) |
| • | `https://example.com/` | → `https://example.com/` (200) |
| • | `https://www.example.com/` | → `https://www.example.com/` (200) |
| ✗ | One canonical host over HTTPS | 4 different end points |
| ✓ | Homepage status | 200 at `https://www.example.com/` |
| ✓ | Homepage indexable | meta robots: none; X-Robots-Tag: none |
| ✗ | Homepage canonical | missing |
| ✗ | robots.txt | HTTP 404 |
| ✗ | Sitemap `https://www.example.com/sitemap.xml` | HTTP 404; 0 URLs |
| ✓ | TLS certificate | 77 days left |Four host variants serving the same page with no redirects or canonical is a duplicate-host problem. It goes into the scope before anything else, because every crawl-based finding will be split across hosts until it's fixed.
Stage 2: Scan
Configuration rules:
- Scan the canonical host from the pre-flight, not whatever the client typed into the brief.
- Crawl production. Scan staging separately only if fixes will be verified there.
- Decide page scope deliberately: a homepage-plus-templates audit for a 10,000-URL store says more than a random 200-page crawl. Name the templates in the intake file.
- Record the date, the tool and the settings in the report appendix. A re-audit is only comparable if it uses the same settings.
Never automate: the decision to proceed when the scan hit a bot challenge or login wall. A report built on a challenge page is worse than no report.
Stage 3: Triage (Where Audits Win or Lose)
A raw audit lists everything. A useful audit lists what to do first. Triage turns hundreds of findings into a short, ordered list.
Know what's "normal" before you call it critical
Clients hear every finding as "something's wrong with my site". It helps to know which issues almost every site has. These are the shares of randomly sampled homepages (Tranco ranks 1,001–50,000) that had each issue in our 2026 measurements:
| Finding (first-pass audit) | Share of sites | Sample |
|---|---|---|
| ≥1 automatically detectable WCAG A/AA failure | 87.7% | 219 homepages |
| Mobile LCP over 2.5s on a throttled load (lab) | 74.6% | 201 homepages |
| Low-contrast text | 64.4% | 219 homepages |
| Tracking cookie set before consent | 60.7% | 196 homepages |
| Security headers grade F | 60.6% | 188 homepages |
| ≥1 internal link that redirects | 56.6% | 219 homepages |
| ≥1 tap target failing WCAG 2.5.8 | 43.6% | 188 homepages |
| Duplicate title, meta description or H1 | 42.1% | 57 crawled sites |
| Server still accepts TLS 1.0/1.1 | 32.6% | 215 hosts |
Image with no alt attribute | 28.3% | 219 homepages |
| ≥1 broken internal link | 17.4% | 219 homepages |
| Pinch-zoom blocked | 15.1% | 188 homepages |
| Horizontal scroll at 360px | 12.2% | 188 homepages |
| Near-duplicate page content | 8.8% | 57 crawled sites |
Sources: accessibility statistics, security headers statistics, pre-consent tracking, broken internal links, duplicate content and the mobile checklist, each with its full method.
Two practical uses:
- Calibrate tone. "Your site has contrast failures" is true of two-thirds of the web. Present it as standard hygiene with a clear fix, not as an emergency.
- Spot the unusual. A
noindexhomepage, a site-wide canonical to the wrong host or a broken checkout template are rare. When you find one, it goes to the top of the report, whatever its count.
The triage score
Score every finding on three axes, 1–3 each, and sort by Impact × Reach ÷ Effort:
| Axis | 1 | 2 | 3 |
|---|---|---|---|
| Impact | Cosmetic, best practice | Affects rankings, UX or compliance | Blocks indexing, revenue or legal exposure |
| Reach | One page | One template or section | Site-wide |
| Effort | Config or template change | A sprint ticket | A project or vendor change |
Group findings by template before scoring. "212 pages missing meta descriptions" is usually one blog template, so it's one fix with Reach 3.
Never automate: discarding false positives and judging business impact. A tool can count missing alt attributes. Only a person can see that the 40 failing images are decorative dividers and the 3 that matter are product photos.
Stage 4: The Client Report
Report structure (keep this order):
- One-page summary: health score, the top five actions, and what you need from the client (access, approvals, budget)
- Scorecard by area (performance, SEO, accessibility, security, mobile) with grades
- Top 5 findings in business language: what's wrong, why it costs them, the fix, the effort
- Full issue list, prioritised, with copy-paste fixes for developers
- Appendix: method, scan date, scope, settings
The anti-pattern is the 80-page export nobody reads. Lead with five actions and push the rest into the appendix. Our guide to the anatomy of a great website audit report covers presentation in depth.
Never automate: the top-five selection and the narrative. Generated summaries are a good first draft. A strategist should rewrite the opening paragraph for every client.
Stage 5: Fix Tracking
Convert approved findings into tickets the client's developers can close without a meeting:
### [SEO-014] Homepage canonical points to http:// host
**Impact:** 3 · **Reach:** 3 (site-wide template) · **Effort:** 1
**Found:** 2026-10-09 audit, Security & SEO sections
**Fix:** In `layouts/base.html`, build the canonical from the HTTPS origin:
`<link rel="canonical" href="https://example.com{{ page.path }}">`
**Acceptance criteria:** Re-audit shows canonical = `https://` final URL on all crawled pages; no "canonical mismatch" findings.
**Owner:** Client dev team · **Due:** Sprint 42Every ticket needs acceptance criteria that a re-scan can verify. "Improve performance" can't be closed. "Mobile LCP ≤ 2.5s on the product template in the scheduled scan" can.
Stage 6: Re-Audit
Run the same scan with the same settings once the client marks tickets done, and compare:
- Score and grade per area, before and after
- Each closed ticket's acceptance criterion: pass or fail
- New findings. Fixes introduce regressions, especially redirects and canonicals
Report the delta in one table. It's the most persuasive page you'll ever send, because it shows the client's money turning into measurable change.
Stage 7: Monitoring
Sites decay: new plugins, new tags, new templates, expiring certificates. A scheduled monthly scan of the same scope turns the audit into a baseline and catches regressions before the client does. It's also the natural bridge from a one-off audit to a retainer.
Where BugViso Fits in This SOP
BugViso covers the mechanical half of stages 2, 4, 6 and 7:
- Scan: a multi-page crawl (homepage links, then the full sitemap tree) rendered in Chromium. It covers Core Web Vitals and throttled-network LCP, axe-core WCAG checks, a mobile emulation pass, SEO and structured data, duplicate content, broken links, security headers and TLS, and the pre-consent privacy audit. Protected-site detection stops the scan and tells you when a bot challenge or WAF blocked it, instead of scoring the challenge page.
- Report: a white-label PDF with your company name, logo, brand colour and custom header and footer. It includes an executive scorecard, A–F grades, a plain-English summary and a prioritised remediation playbook with copy-paste fixes. Your brand kit is saved once and applied to scheduled reports.
- Fix tracking inputs: CSV export of your audit records and of each crawl, plus the full audit as JSON, ready to turn into tickets.
- Re-audit: every scan is kept in your audit records, with a one-click Re-audit that re-runs the same domain so before and after sit side by side.
- Monitoring: scheduled scans (daily, weekly or monthly, at a time and timezone you choose, labelled per client) re-run the audit and produce a fresh branded report each time.
Stages 1, 3 and 5 stay human: intake conversations, judging business impact and writing tickets your client's team will actually close. You can run your first client audit with BugViso. The site crawl and white-label reports page covers crawl limits and report fields.
Common SOP Failures
- Skipping intake. You crawl
wwwwhile the site canonicalises to the apex, or you audit a staging site withnoindex. The pre-flight takes a minute. - Reporting counts instead of causes. "1,204 issues" frightens and paralyses. "Three template fixes resolve 1,100 of them" gets approved.
- No acceptance criteria. Tickets without a verifiable "done" stay half-done forever.
- Changing scan settings between audits. A different page cap or device profile makes before/after numbers meaningless.
- One-off audits. Without monitoring, the next CMS update undoes the work, and the client remembers the audit as the thing that didn't stick.
FAQ
What should an SEO audit process include?
Intake (goals, access, pre-flight checks), a full technical scan, triage by impact and effort, a client report led by a short summary, tickets with acceptance criteria, a re-audit with the same settings, and ongoing scheduled monitoring.
How long should an agency SEO audit take?
The scan is automated and takes minutes to an hour, depending on site size. Most of the time goes on triage and the report, which depend on site complexity and how much of the fix work is in scope. Template-level grouping is what keeps that time under control.
What can be automated in an SEO audit?
Crawling, rendering, metric collection, issue detection, scoring, report formatting, branding and scheduled re-scans. Business-impact judgement, false-positive review, prioritisation and the client narrative should stay human.
How do I make clients act on audit findings?
Lead with five actions in business terms, put the effort and owner on each, turn them into tickets with acceptance criteria, and show a before/after re-audit. Progress the client can see is what earns the next approval.
Conclusion
A repeatable audit is a process, not a document: pre-flight before you scan, triage by template and impact, report five actions instead of 500 issues, and re-audit with identical settings. A BugViso scan handles the crawl, the branded report and the scheduled re-checks, so your team's hours go to the judgement calls.
See where your site stands
Run a free BugViso audit for SEO, speed, accessibility and AI search readiness — with fixes you can ship today.