How much of WCAG can be automated? Criterion by criterion
In our assessment of the 86 WCAG 2.2 success criteria, software can decide 3 for most content, find some failures but never confirm a pass for 32, and cannot judge 51. Karl Groves puts the share a tool can definitively test at approximately 25-29% of WCAG 2.0 best practices. Either way, most of WCAG needs a person.
Published . Last reviewed .
WCAG 2.2 has 86 success criteria, and no tool tests all of them. This page rates every criterion by how much software can test it, lists the 55 at Levels A and AA with the reason for each rating, and explains why a page can pass every automated rule and still fail the people who use it.
The ratings are GotAlt's own assessment, made from what each criterion asks for and what rule-based and AI checks can see. They are not part of WCAG, and other testers may place some criteria differently. The same ratings, for all 86 criteria and with links to W3C's guidance, are in the testability column of our WCAG 2.2 criteria list.
The counts: 3 automated, 32 partly automated, 51 manual
We recounted the ratings from the criteria list on our WCAG 2.2 page. By level, they come out like this.
| Level | Criteria | Automated | Partly automated | Manual |
|---|---|---|---|---|
| A | 31 | 1 | 16 | 14 |
| AA | 24 | 1 | 9 | 14 |
| A and AA together | 55 | 2 | 25 | 28 |
| AAA | 31 | 1 | 7 | 23 |
| All levels | 86 | 3 | 32 | 51 |
The three ratings mean this:
- Automated: the requirement is a measurable property of the code or the rendered page, so software can decide pass or fail for most content. A person still reviews the edge cases a tool reports.
- Partly automated: software can find some real failures, such as a missing label or a missing title, but cannot confirm a pass, because meeting the criterion depends on meaning, context or behavior.
- Manual: a person has to judge the content or operate the page. Tools can at most point to where to look.
The counts are of criteria, not of failures or hours of work. Of the 86, 35 (3 plus 32) have an automated part, but for 32 of those a clean result from a tool is not a pass. 4.1.1 Parsing, which WCAG 2.2 removed, is not counted.
As shares of the 86, 3 (3.5%) can be decided automatically and 35 (40.7%) have an automated part. At Levels A and AA, 2 of the 55 can be decided automatically and 27 of the 55 have an automated part. These percentages are plain arithmetic on the counts in the table. Neither is a measure of how many failures a tool finds on a real page, which the research below addresses.
What published research says about automated coverage
Four published sources put numbers on it. They measure different things, so their figures do not combine into one, but they point the same way.
- Karl Groves, WCAG 2.0 best practices. In an article published on May 25, 2018, Groves wrote that an automated tool "can definitively test for approximately 25-29% of best practices for WCAG 2.0" and "cannot test for approximately 40%" (Automated lies, with one line of code). It is one practitioner's approximation, for WCAG 2.0 best practices rather than for WCAG 2.2 success criteria.
- UK Government Digital Service, one test page. The GDS accessibility team put 142 deliberate barriers on a page and ran 13 automated tools against it. The best tool fully detected 40% of them on a pass or fail basis and the worst 13%. Counting "potential barriers" that tools noticed but needed a person to check, such as whether alt text is accurate, a different tool led that column, at 50% (GDS accessibility tool audit, last updated April 13, 2018). The audit predates WCAG 2.1 and 2.2, so read it as a snapshot of tools from 2018.
- WebAIM Million, real home pages. WebAIM's 2026 scan of one million home pages detected WCAG 2 failures on 95.9% of them. WebAIM states that "Absence of detected errors does not indicate that a page is accessible or conformant," and that because it counted only automatically detectable failures, full A and AA conformance is certainly lower than the 4.1% of pages with none (WebAIM Million).
- Deque, issue volume and criteria (vendor research). Deque Systems, which makes the axe accessibility tools, analyzed 13,000+ pages and page states and nearly 300,000 issues. It reports that "57.38% of total issues were identified using Deque's automated tests." It also reports that "we found automated issues for 16 out of the 50 Success Criteria under WCAG 2.1 Level AA." The first figure counts issues and the second counts criteria. The study "focused on HTML pages only" and "concentrated on first-time audits," and Deque excludes false positives from its reports. The report is titled "The Automated Accessibility Coverage Report" and is on deque.com at /automated-accessibility-testing-coverage/. It is research by a vendor of automated testing tools, it shows no publication date, and we have not checked its data.
These sources count best practices, deliberate barriers on one page, failures detected on real pages, issues and, here, criteria, so the numbers should not be added or compared. Sources disagree on the exact share of WCAG that software can detect. Deque's criteria-based count, 16 of the 50 WCAG 2.1 AA criteria, is 32%, while the same study reports 57.38% by issue volume, because one counts criteria and the other counts issues. We attribute every figure and treat none as settled. What the sources share is direction: in that 2018 test no tool found more than 50%, WebAIM says a clean automated result proves nothing about accessibility, and W3C says evaluation tools "can not determine accessibility, they can only assist in doing so" (W3C WAI, selecting web accessibility evaluation tools).
The 55 Level A and AA criteria, by how much software can test
Laws in our table that name a WCAG level name AA (Section 508 names A and AA, which is the same set), which means all 55 criteria at Levels A and AA (see WCAG A vs AA vs AAA). Each criterion below links to its row on our WCAG 2.2 page, which links on to W3C's Understanding document. The reason column is the reasoning from that page's testability column. The last column, "Checked by GotAlt," is specific to this page: it shows which of our free scan, deep audit and bookmarklet examine some aspect of the criterion, from what we check. A mapped criterion is examined in part, never tested in full, and a dash means none of the three examines it.
The 31 Level AAA criteria are rated on the full list: 1 automated (1.4.6 Contrast (Enhanced)), 7 partly automated and 23 manual.
Automated: 2 of the 55
Both are properties a program can read: a language attribute in the code, or a ratio calculated from two colors. Even here, a person confirms that the declared language matches the text and judges text over photos, gradients or video.
| Success criterion | Level | Why we rate it this way | Checked by GotAlt |
|---|---|---|---|
| 1.4.3 Contrast (Minimum) | AA | The ratio is a calculation; text over images, gradients or video still needs a person to judge. | Bookmarklet |
| 3.1.1 Language of Page | A | A missing or invalid language attribute is detectable; a quick look confirms it matches the text. | Scan, Bookmarklet |
Partly automated: 25 of the 55
Software finds the clear failures in this group, such as a missing label, a missing title or a disabled zoom, but cannot confirm that the criterion is met. A clean automated result here means only that those checks found nothing.
| Success criterion | Level | Why we rate it this way | Checked by GotAlt |
|---|---|---|---|
| 1.1.1 Non-text Content | A | Software finds missing alt text; only a person, or a model that looks at the image, can say whether the words match the picture. | Scan, Deep audit, Bookmarklet |
| 1.2.2 Captions (Prerecorded) | A | Software can flag a video element with no caption track; a person must check that captions exist in the player and are accurate. | None |
| 1.3.1 Info and Relationships | A | Software catches missing labels, broken table markup and heading problems such as skipped levels (a best practice, not a 1.3.1 failure on its own); whether all visual structure is coded needs a person. | Scan, Bookmarklet |
| 1.3.4 Orientation | AA | Software can find code that locks the orientation; whether the lock is essential is a judgment. | None |
| 1.3.5 Identify Input Purpose | AA | Software can check that autocomplete values are valid; deciding which fields collect personal data needs a person. | None |
| 1.4.1 Use of Color | A | Software can flag links told apart from text by color alone; other uses of color need a person to interpret. | None |
| 1.4.2 Audio Control | A | Software can find media set to autoplay; whether a working control exists needs a check. | Scan |
| 1.4.4 Resize Text | AA | Software can flag zoom disabled in the viewport tag; spotting clipped or overlapping text at 200 percent needs a person. | Scan |
| 1.4.10 Reflow | AA | Software can detect sideways scrolling at that width; whether content or function is lost needs a person. | None |
| 1.4.11 Non-text Contrast | AA | The ratio is measurable, but deciding which pixels identify a control or carry meaning needs judgment. | None |
| 1.4.12 Text Spacing | AA | A tool can apply the spacing; spotting lost or overlapping content still needs a person to look. | None |
| 2.1.1 Keyboard | A | Software can flag some controls that cannot take keyboard focus; confirming everything works needs a person with a keyboard. | None |
| 2.2.1 Timing Adjustable | A | Software can flag a timed page refresh; other time limits live in scripts and need testing. | Scan |
| 2.2.2 Pause, Stop, Hide | A | Software can flag a few known patterns; most motion comes from scripts and needs a person to watch. | None |
| 2.3.1 Three Flashes or Below Threshold | A | Video analysis software can measure flashes; someone still has to find and run every animation and video. | None |
| 2.4.1 Bypass Blocks | A | Software can find skip links, landmarks and headings; whether they actually work needs a check. | None |
| 2.4.2 Page Titled | A | A missing title is detectable; whether it describes the page is a judgment. | Scan, Bookmarklet |
| 2.4.3 Focus Order | A | Software can flag positive tabindex values; the order itself needs a person to tab through. | Scan |
| 2.4.4 Link Purpose (In Context) | A | Software finds empty links and generic text such as "click here"; judging whether the context makes the purpose clear needs a reader. | Scan, Deep audit, Bookmarklet |
| 2.4.7 Focus Visible | AA | Software can flag CSS that removes the focus outline; whether a visible indicator remains needs a person to tab through. | None |
| 2.5.3 Label in Name | A | Software can compare visible text with the accessible name; labels drawn as images and edge cases need a person. | None |
| 2.5.8 Target Size (Minimum) | AA | Sizes and spacing are measurable in a rendered page; the exceptions need judgment. | None |
| 3.1.2 Language of Parts | AA | Software can validate language codes; finding unmarked foreign passages needs a reader. | None |
| 3.3.2 Labels or Instructions | A | Software finds fields with no label; whether the label or instructions are enough needs a person. | Scan, Bookmarklet |
| 4.1.2 Name, Role, Value | A | Software finds controls with no accessible name and invalid ARIA; whether custom widgets behave correctly needs testing with assistive technology. | Scan |
Manual: 28 of the 55
These need someone to operate the page or judge its content: to tab through it, submit a form with mistakes, watch a video with its captions, or compare pages across the site. A tool can at most point to where to look.
| Success criterion | Level | Why we rate it this way | Checked by GotAlt |
|---|---|---|---|
| 1.2.1 Audio-only and Video-only (Prerecorded) | A | A tool cannot tell what a recording contains or whether an alternative matches it. | None |
| 1.2.3 Audio Description or Media Alternative (Prerecorded) | A | Deciding what visual information needs describing is a judgment. | None |
| 1.2.4 Captions (Live) | AA | Live captions can only be checked while the stream is running. | None |
| 1.2.5 Audio Description (Prerecorded) | AA | Deciding what visual information needs describing is a judgment. | None |
| 1.3.2 Meaningful Sequence | A | Comparing the visual order with the code order needs a judgment about meaning. | None |
| 1.3.3 Sensory Characteristics | A | A person has to read instructions in context. | None |
| 1.4.5 Images of Text | AA | Finding text inside images and judging whether it could be real text needs a person. | None |
| 1.4.13 Content on Hover or Focus | AA | Needs a person to trigger each pop-up with a mouse and a keyboard. | None |
| 2.1.2 No Keyboard Trap | A | A trap only shows when someone tabs through the page. | None |
| 2.1.4 Character Key Shortcuts | A | Shortcuts are found by testing the page or reading its scripts. | None |
| 2.4.5 Multiple Ways | AA | Requires looking across the whole site. | None |
| 2.4.6 Headings and Labels | AA | Whether the words describe the content is a judgment. | Deep audit |
| 2.4.11 Focus Not Obscured (Minimum) | AA | Depends on layout, scroll position and overlays at the moment of focus. | None |
| 2.5.1 Pointer Gestures | A | Needs a person to find gestures and try the alternative. | None |
| 2.5.2 Pointer Cancellation | A | Needs a person to test how each control responds to press and release. | None |
| 2.5.4 Motion Actuation | A | Needs a person with a device to find and test motion features. | None |
| 2.5.7 Dragging Movements | AA | Needs a person to find drag interactions and try the alternative. | None |
| 3.2.1 On Focus | A | Needs a person to move focus through the page. | None |
| 3.2.2 On Input | A | Needs a person to operate each control. | None |
| 3.2.3 Consistent Navigation | AA | Requires comparing pages across the site. | None |
| 3.2.4 Consistent Identification | AA | Requires comparing components across the site. | None |
| 3.2.6 Consistent Help | A | Requires comparing pages across the site. | None |
| 3.3.1 Error Identification | A | Needs a person to submit forms with mistakes and read the messages. | None |
| 3.3.3 Error Suggestion | AA | Needs a person to submit forms with mistakes and judge the suggestions. | None |
| 3.3.4 Error Prevention (Legal, Financial, Data) | AA | Needs a person to walk through each such process. | None |
| 3.3.7 Redundant Entry | A | Needs a person to walk through each multi-step process. | None |
| 3.3.8 Accessible Authentication (Minimum) | AA | Needs a person to walk through the sign-in process. | None |
| 4.1.3 Status Messages | AA | Needs a person to trigger each message and listen with a screen reader. | None |
How a page can pass every automated rule and still fail users
An automated rule checks a property of the code or the rendered page: an attribute exists, a ratio is above a threshold, a heading level was not skipped. Whether the page works for a person is a different question, and the gap between the two is where the partly automated and manual ratings come from. Four common cases follow. Why rule scanners miss issues works through more examples.
Text that exists but is wrong or empty of meaning (1.1.1, 2.4.2, 2.4.4, 2.4.6)
A rule can confirm that an alt attribute is present and is not a file name. It cannot tell whether the words are true of the picture. Suppose the photo shows a tan suede boot, front view. This still passes every presence rule:
<img src="boot-tan.jpg"
alt="Black leather ankle boot, side profile">
WebAIM's 2026 scan found questionable or repetitive alternative text on 10.8% of the images that had alternative text: present is not the same as useful (WebAIM Million). Our alt text checker and deep audit open the image and compare it with the text. See WCAG 1.1.1 explained and alt text examples.
The same gap appears wherever the rule is about presence and the criterion is about meaning. A link named "Read more" has a name, and WCAG 2.4.4 lets the surrounding sentence supply its purpose, so a tool can advise but not fail it (the link text checker lists links that need a second look). A page title of "Page" or headings such as "Overview" pass any check for presence and skipped levels, but WCAG asks for titles and headings that describe their topic or purpose (the heading checker shows the outline, and the guide to heading structure explains what to look for).
Behavior that only shows when someone uses the page (2.1.2, 2.4.3, 2.4.7)
Markup can be valid and every control focusable while a keyboard trap still locks a user inside a widget, focus jumps around the screen or the focus ring is invisible. Finding these means pressing Tab and watching, which a program that only reads markup never does. Our methodology page lists keyboard traps, focus order and visibility, screen reader output and reading order among the things the GotAlt scan cannot detect. The keyboard accessibility test shows how to check them.
Contrast the code cannot see (1.4.3, 1.4.11)
A color pair can pass in the stylesheet while the text sits on a photo, a gradient or a video, where there is no single background color to measure. The GotAlt scan reads only the HTML a server sends and so never computes rendered colors, and the bookmarklet lists text over images for a manual check. The image contrast checker is for that case, and text over images explains the method. For icons, borders and focus indicators, see WCAG 1.4.11 Non-text Contrast.
What GotAlt's scan, deep audit and bookmarklet each cover
This table is drawn from our methodology page and the tool pages. A check that maps to a criterion examines one aspect of it. It is not a test of the whole criterion.
| Tool | What it reads | Criteria its checks map to | What it cannot tell you |
|---|---|---|---|
| Free scan | The HTML your server sends: 16 rule-based checks on up to three pages. | 1.1.1, 1.3.1, 1.4.2, 1.4.4, 2.2.1, 2.4.2, 2.4.3, 2.4.4, 3.1.1, 3.3.2 and 4.1.2, plus 4.1.1 for WCAG 2.0 and 2.1. | Rendered colors, layout and focus, keyboard behavior, anything a script adds after the page loads, and whether any text is true or sensible. |
| Deep audit | The same HTML plus the image files. A vision model compares up to six images per page with their alt text, and checks whether link text and headings make sense to someone who cannot see the page. | 1.1.1, 2.4.4 and 2.4.6. | Images beyond the six it opens (the rule checks still cover every image), and anything that needs someone to use the page. A model's verdict is a flag for review, not a pass. |
| Bookmarklet | The page as your browser has rendered it, in your own tab, including a staging site or a page behind a login. | 1.1.1, 1.3.1, 1.4.3, 2.4.2, 2.4.4, 3.1.1 and 3.3.2. | Whether alt text is true, whether headings make sense, whether the page works with a keyboard or a screen reader. Text over images is listed for a manual check. |
Between them, the three map to 13 of the 55 Level A and AA criteria. By our ratings, 2 of the 13 are automated, 10 are partly automated and 1 (2.4.6 Headings and Labels) is manual, so even the mapped criteria still need a person. A model's verdict in the deep audit does not change a rating: it can flag that alt text does not match an image, but a person decides what the alt text should say. That leaves 42 of the 55 outside what these three examine. Our single-purpose tools for contrast, headings, link text and PDFs (see also accessible PDFs) go further in their own areas, and the sample report shows what a run produces.
What to do with this
-
Run the automated pass first
It is cheap, and it covers the failures that occur most. WebAIM found that 96% of the errors its 2026 scan detected fall into six categories: low contrast text, missing alternative text, missing form input labels, empty links, empty buttons and a missing document language (WebAIM Million). Our guide to the most common accessibility errors maps each to a fix.
-
Test the rest with a person
Grouping the remaining criteria by test method avoids working through them one at a time. A keyboard pass covers 2.1.1, 2.1.2, 2.4.3, 2.4.7, 2.4.11, 1.4.13, 3.2.1 and 3.2.2 (see the keyboard accessibility test). A zoom and reflow pass covers 1.4.4, 1.4.10, 1.4.12 and 1.3.4. A forms pass covers 1.3.5, 3.3.1, 3.3.3, 3.3.4, 3.3.7 and 3.3.8, with a screen reader for 4.1.3. A media pass covers 1.2.1 to 1.2.5, and a comparison across pages covers 2.4.5, 3.2.3, 3.2.4 and 3.2.6. This grouping is our suggestion, not a W3C method. How to test website accessibility is the place to start.
-
Record what was tested and what was not
An audit is only as useful as its scope. What a website accessibility audit contains covers scope and method, and VPAT and ACR covers the reporting format that procurement teams ask for.
-
Run the automated pass again when pages change
Templates, plugins and content edits reintroduce the failures a scan can find. The bookmarklet works on any page you can open, and monitoring plans check your chosen pages every week and rerun the deep audit when a page's images or alt text change.
What we don't claim
The ratings above are our judgment, not a W3C classification, and the three categories are broad: some criteria have parts that fit more than one. A rating of "automated" still leaves edge cases for a person. Nothing on this page, and no scan, audit or checklist from GotAlt or anyone else, makes a site meet WCAG. Our editorial policy explains how pages like this are sourced and reviewed.
Questions about automated WCAG testing
How much of WCAG can be automated?
In our assessment of the 86 WCAG 2.2 success criteria, software can decide 3 for most content, can find some failures but not confirm a pass for 32, and cannot judge 51. Measured as a share of WCAG 2.0 best practices, Karl Groves estimated that an automated tool can definitively test approximately 25-29%. Deque, a vendor of testing tools, reports that its automated tests found issues for 16 of the 50 WCAG 2.1 AA criteria (32%) and identified 57.38% of issues by volume. Sources disagree because they count different things.
Which WCAG criteria can software test fully?
By our ratings, at Levels A and AA only 3.1.1 Language of Page and 1.4.3 Contrast (Minimum), and at Level AAA only 1.4.6 Contrast (Enhanced). Even for these, a person confirms that the declared language matches the text and judges text over photos, gradients and video.
If an automated scan finds no errors, is my site accessible?
No. WebAIM, which publishes a yearly automated scan of one million home pages, states that an absence of detected errors does not indicate that a page is accessible or conformant. A scan reports the failures it can detect and says nothing about the ones it cannot.
Why do sources give different percentages?
They count different things. Karl Groves counts WCAG 2.0 best practices, the GDS audit counts deliberate barriers on one test page, WebAIM counts detected failures on real home pages, and we count WCAG 2.2 success criteria. Deque's study shows the problem within one source: 57.38% of issues by volume but 16 of 50 WCAG 2.1 AA criteria, about 32%, which shows how much the choice of what to count changes the number. Its data is HTML pages and first-time audits only. None of the figures converts into another.
Can AI check the criteria that rules cannot?
In part. A vision model can judge things a rule cannot, such as whether alt text matches an image, which our deep audit does for up to six images per page. A model's verdict is still a flag for a person to review, so it does not move a criterion from manual to automated in our ratings. Our test of 124 alt texts shows how accurate that check is and where it errs.
Is an automated tool enough for an accessibility audit?
No. W3C says evaluation tools cannot check all accessibility aspects automatically and that human judgment is required. A useful audit combines automated checks, manual testing against the criteria and a report that states its scope. See what a website accessibility audit contains.
See what a scan finds, and what it leaves for you
The free scan runs 16 rule-based checks on up to three pages, maps each finding to a WCAG criterion and includes a section on what a scan leaves to a person. No signup.
Scan my site for free See a real report first