Automated Accessibility Testing: What It Catches and Misses
What automated accessibility testing with axe-core catches, what it flags for review, and what it can never see. Data from 219 sites plus a manual checklist.
Automated accessibility testing reliably catches failures a machine can verify from the DOM: low color contrast, images with no alt attribute, links and buttons with no accessible name, unlabeled form fields, invalid ARIA, a missing page language, disabled zoom and undersized touch targets. It can't judge whether alt text is accurate, whether keyboard focus moves in a logical order, whether captions match the audio, or whether an error message actually helps.
The line between those two groups decides how far you can trust a green dashboard. Deque Systems, the company that maintains the open-source axe-core engine, has published research estimating that automated testing finds roughly 57% of accessibility issues by volume. Measured by WCAG success criteria, the coverage is smaller still. Of the 55 Level A and AA criteria in WCAG 2.2, only a handful can be checked end to end by a machine.
If you've already run a scan, our breakdown of the most common WCAG violations shows how to fix the failures automation does find.
This guide splits accessibility testing into what automation decides, what it flags for a human, and what it never sees. It uses axe-core results from 219 real homepages and ends with a script that reports all three buckets.
The Three Buckets Every axe-core Run Returns
Most dashboards show one number. Under the hood, axe-core sorts every check into separate result lists:
| Bucket | What it means | What you do |
|---|---|---|
| violations | The rule ran and definitely failed on these elements | Fix them |
| incomplete ("needs review") | The rule found relevant elements but couldn't decide | A person must check each one |
| passes | The rule ran and passed | Nothing, but a pass doesn't prove conformance |
| inapplicable | No elements on the page for the rule to test | Nothing |
| (not in any bucket) | Requirements no rule exists for | Manual testing |
The incomplete bucket is the one teams skip. When text sits on a background image or gradient, for example, axe can't calculate one background color. Rather than guess, it puts the element in "needs review". A tool that reports only violations will show that element as passing.
⚠️ "0 violations" doesn't mean accessible. It means no rule that could run, failed. Everything in "needs review" and everything outside the rule set is still unknown.
What Automation Found on 219 Real Homepages
We drew 420 domains at random from Tranco list 94GG2 (ranks 1,001–50,000), loaded each homepage in Chromium at 1366×900 on 3 October 2026, and kept the 219 that served a real page. We ran axe-core 4.10.2 with the WCAG 2.0, 2.1 and 2.2 A/AA rule tags, the same configuration BugViso uses.
| Metric (219 homepages) | Result |
|---|---|
| Homepages with ≥1 violation | 192 (87.7%) |
| Homepages with zero violations | 27 (12.3%) |
| Median failing elements per homepage | 10 (mean 33.7; 90th percentile 80) |
| Median distinct rules failed | 3 |
| Violations rated serious / critical (by element) | 6,256 / 1,127 |
| Homepages with ≥1 needs-review element (second pass, 209 sites) | 188 (90.0%) |
The most frequently failed rules show where automation is strong:
| axe rule | Homepages failing | What it detects |
|---|---|---|
color-contrast | 64.4% | Text below 4.5:1 (3:1 large) against a resolvable background |
link-name | 42.5% | Links with no accessible text: icon links, empty anchors |
target-size | 24.7% | Pointer targets under 24×24 px without enough spacing |
image-alt | 24.7% | <img> with no alt attribute or equivalent |
button-name | 15.5% | Buttons with no accessible name |
meta-viewport | 11.9% | user-scalable=no or maximum-scale blocking zoom |
list / listitem | 11.4% / 6.8% | Broken list markup |
label | 9.6% | Form inputs with no label |
aria-hidden-focus | 7.3% | Focusable elements inside aria-hidden regions |
html-has-lang | 4.6% | Missing lang attribute on <html> |
Every one of these is a property of the DOM: a ratio, a missing attribute, a size, a role mismatch. Automation is excellent at that kind of check, and these issues are also among the most common, which is why they dominate any automated report.
The needs-review bucket is bigger than the violations bucket
We ran a second pass on the same homepages (209 completed) and recorded axe-core's incomplete results as well. The scale surprised us:
| Needs-review metric (209 homepages) | Result |
|---|---|
| Homepages with ≥1 element needing manual review | 188 (90.0%) |
| Median needs-review elements per homepage | 14 |
| Total needs-review elements vs total violating elements | 8,016 vs 6,349 |
| Zero-violation homepages that still had needs-review items | 25 of 27 |
| Most common needs-review rule | Homepages | Why axe can't decide |
|---|---|---|
color-contrast | 86.1% | Text over images, gradients, pseudo-elements or transparent layers |
aria-valid-attr-value | 18.2% | An ARIA reference points to an element that may be added later |
link-in-text-block | 14.4% | Whether a link is distinguishable from text without color |
aria-prohibited-attr | 12.4% | Whether a label on a generic element is actually announced |
video-caption | 10.0% | A video exists, but whether captions are present and accurate can't be confirmed |
target-size | 6.2% | Overlapping or partly hidden targets |
Contrast dominates: it accounted for 92% of all needs-review elements. Text over images and gradients is common in hero sections and banners, exactly where a low-contrast headline costs the most. Automation flags those elements and steps back. Someone has to look.
The Capability Matrix: Automated vs Manual
| Requirement | Automation can… | Only a human can… |
|---|---|---|
| Images (1.1.1) | Find images with no alt; flag file-name alts | Judge whether alt text is accurate and useful in context |
| Contrast (1.4.3) | Measure text on solid backgrounds | Check text on images, gradients, hover and focus states |
| Keyboard (2.1.1, 2.1.2) | Find scrollable regions that can't be focused | Confirm every control works by keyboard, with no traps |
| Focus order (2.4.3) | — | Decide whether the order makes sense |
| Focus visible (2.4.7) | — (outline removal is hard to judge in CSS) | See whether focus is visible on every element |
| Link purpose (2.4.4) | Find empty links | Judge whether "Read more" × 12 is understandable in context |
| Headings and labels (2.4.6) | Find empty headings or missing labels | Judge whether headings describe the content |
| Forms and errors (3.3.1, 3.3.3) | Find inputs with no label | Submit the form wrong and judge the error messages |
| Captions and audio description (1.2.x) | Find <video> with no <track> | Check captions are accurate and synchronised |
| Reflow (1.4.10) | — | Zoom to 400% / 320px wide and check nothing breaks |
| Status messages (4.1.3) | — | Use a screen reader to confirm updates are announced |
| Authentication (3.3.8) | Detect onpaste blockers in scripts (sometimes) | Walk through login, reset and 2FA flows |
The pattern is consistent. Automation answers "is something there?" A person answers "does it work for someone using it?"
Why Some Failures Are Invisible to Automation
Meaning can't be computed. alt="Image" and alt="Chart showing sales doubled in Q3" both satisfy "has alt text". Only one of them is useful, and a rule engine can't tell which. Our alt text writing guide is about exactly that gap.
Behaviour needs interaction. A keyboard trap only shows up when you tab into a widget and can't tab out. A modal that doesn't return focus to its trigger only fails after you close it. A single-page scan never presses those keys.
Context changes the rule. The same red text passes as a decorative flourish and fails as an error message. "Click here" is a fine link inside a sentence that explains it, and a failure in a list of twelve identical links.
Some content isn't in the DOM at test time. Content behind tabs, accordions, infinite scroll, login walls and multi-step forms never reaches a homepage scan. A clean scan of the homepage tells you nothing about your checkout.
A Manual Testing Checklist for What Automation Misses
Run these on your key templates (homepage, a content page, a form, checkout or signup) after every major release:
- Keyboard only. Unplug the mouse. Tab through the whole page, Shift+Tab back. Can you reach, operate and leave every control? Is focus always visible, and never hidden behind a sticky header?
- Zoom to 200% and 400%. Does text reflow without horizontal scrolling at 320 CSS px wide? Do overlays still fit?
- Screen reader pass. With NVDA (Windows) or VoiceOver (macOS/iOS), navigate by headings, then by landmarks, then by links. Does the outline make sense? Are images described usefully?
- Forms. Submit empty and invalid. Is each error tied to its field, announced, and specific enough to fix?
- Dynamic updates. Add to cart, filter results, load more. Is the change announced (live region) or at least discoverable?
- Media. Play videos with sound off. Are captions accurate? Is there a transcript for audio?
- Motion and timing. Do carousels pause? Can session time-outs be extended?
This list takes about an hour per template. Paired with a clean automated scan, it covers most of what an external auditor will test first.
Run axe-core and See All Three Buckets
This script runs axe-core in Chromium and prints violations and needs-review results separately, plus how many rules passed or didn't apply.
#!/usr/bin/env python3
"""Run axe-core against a live page and separate what automation decided from what it couldn't.
axe-core returns three buckets that most dashboards collapse into one number:
violations – rules that definitely failed
incomplete – "needs review": axe found the element but cannot decide (a human must)
passes – rules that ran and passed
Everything axe has no rule for (keyboard traps, focus order, alt-text *quality*,
caption accuracy…) never appears in any bucket — see the manual checklist in the post.
Usage: pip install playwright && playwright install chromium
python3 axe_audit.py https://example.com
python3 axe_audit.py https://example.com --json > report.json
"""
import argparse, asyncio, json, sys
from playwright.async_api import async_playwright
AXE_CDN = "https://cdnjs.cloudflare.com/ajax/libs/axe-core/4.10.2/axe.min.js"
TAGS = ["wcag2a", "wcag2aa", "wcag21a", "wcag21aa", "wcag22aa"]
async def main(url, as_json):
async with async_playwright() as pw:
browser = await pw.chromium.launch()
# bypass_csp: sites with a strict Content-Security-Policy would block the axe <script> tag
page = await browser.new_page(viewport={"width": 1366, "height": 900}, bypass_csp=True)
await page.goto(url, wait_until="load")
await page.wait_for_timeout(2000)
await page.add_script_tag(url=AXE_CDN)
res = await page.evaluate("""async (tags) => {
const r = await axe.run(document, {runOnly: {type: 'tag', values: tags}});
const slim = (list) => list.map(v => ({id: v.id, impact: v.impact, help: v.help,
nodes: v.nodes.length, sample: (v.nodes[0] && v.nodes[0].target.join(' ')) || ''}));
return {violations: slim(r.violations), incomplete: slim(r.incomplete),
passes: r.passes.length, inapplicable: r.inapplicable.length};
}""", TAGS)
await browser.close()
if as_json:
json.dump(res, sys.stdout, indent=2)
return
print(f"axe-core 4.10.2 on {url}\n")
for bucket, label in (("violations", "VIOLATIONS (fix these)"), ("incomplete", "NEEDS REVIEW (a human must decide)")):
rows = sorted(res[bucket], key=lambda v: -v["nodes"])
print(f"== {label}: {len(rows)} rules, {sum(v['nodes'] for v in rows)} elements")
for v in rows:
print(f" [{v['impact'] or '-':8}] {v['id']:28} x{v['nodes']:<4} {v['help']}")
print(f" e.g. {v['sample'][:90]}")
print()
print(f"Rules passed: {res['passes']} | rules not applicable to this page: {res['inapplicable']}")
if __name__ == "__main__":
ap = argparse.ArgumentParser()
ap.add_argument("url")
ap.add_argument("--json", action="store_true")
a = ap.parse_args()
asyncio.run(main(a.url, a.json))Here's real output from a large open-source CMS homepage that scores zero violations:
== VIOLATIONS (fix these): 0 rules, 0 elements
== NEEDS REVIEW (a human must decide): 2 rules, 5 elements
[serious ] link-in-text-block x3 Links must be distinguishable without relying on color
[serious ] color-contrast x2 Elements must meet minimum color contrast ratio thresholds
Rules passed: 27 | rules not applicable to this page: 35A violations-only dashboard would show a perfect result. The honest reading is: 27 automated checks passed, and 5 elements still need a person to look at them.
How BugViso Uses Automated Testing (and Where It Stops)
BugViso runs a self-hosted axe-core engine with the WCAG 2.0, 2.1 and 2.2 A/AA tags on every audited page, after the page has fully rendered in Chromium. Each violation is reported with its rule, WCAG criterion, impact and failing elements. For contrast failures it shows the measured foreground/background colors and ratio. A separate mobile device-emulation pass adds tap-target sizing against WCAG 2.5.8 and checks for horizontal overflow and viewport problems.
On a multi-page crawl, BugViso counts a rule that fails in a shared header once per site rather than once per page, so one template bug doesn't look like 200 separate problems. Findings appear in the web report and the PDF report with remediation guidance.
BugViso doesn't claim that an automated pass makes a site compliant. It clears the machine-verifiable failures quickly and repeatedly, so your manual testing time goes to the judgement calls listed above. Start with a free BugViso accessibility scan to see which bucket your issues fall into.
Common Misconceptions
- "Our score is 100, so we're compliant." A perfect automated score covers only the criteria a machine can test. Legal standards are written against all of WCAG.
- "Overlays fix what scanners find." In our own tests, overlay widgets rarely changed axe-core results at all. See our analysis of whether accessibility overlays work.
- "Manual testing is only for big companies." The checklist above takes about an hour per template, and it finds the problems users actually complain about.
- "Automated tools produce lots of false positives." Modern engines like axe-core are designed to avoid them, which is exactly why they move uncertain cases to "needs review" instead of failing them.
FAQ
What percentage of accessibility issues can automated testing find?
Estimates vary with how you count. Deque's published research puts it at roughly 57% of issues by volume, because common failures like contrast and missing names are highly automatable. Measured by WCAG success criteria, the share that can be fully automated is much smaller, since most criteria need human judgement.
Is axe-core better than other accessibility scanners?
axe-core is the open-source engine behind many commercial and browser tools, and it's built to keep false positives low. Every engine has the same fundamental limit, though: none can judge meaning, behaviour over time or context.
How often should I run automated accessibility tests?
On every deploy, or at least every release, because regressions are cheap to catch early. Run manual testing on key templates each quarter and whenever a template changes substantially.
Do I still need an accessibility audit if my scans are clean?
Yes, if you need to claim conformance. A clean automated scan removes the most common failures, but an audit by someone using keyboard and screen readers is the only way to evaluate the criteria automation can't reach.
Why do two tools give different results for the same page?
They use different rule sets, different WCAG versions, different timing (before or after scripts load) and different rules for what goes to "needs review". Compare tools by the rules they report, not by their headline score.
Conclusion
Automated testing is the fastest way to find the failures a machine can verify, and the least reliable way to prove a site is accessible, so treat its violations as a to-do list and its "needs review" items as a human's job. A BugViso scan handles the first part on every page it crawls.
See where your site stands
Run a free BugViso audit for SEO, speed, accessibility and AI search readiness — with fixes you can ship today.