Blog · 12 min read

Three Part WCAG 2.2 Testing Workflow for Developers to Run in CI/CD

Developer and tester running accessibility checks

Test WCAG 2.2 by running a three-part workflow: automated scans, targeted manual and code-level checks, and usability testing with people who use assistive technology. Scope your evaluation using WCAG-EM sampling rules before making any conformance claim, and keep dated records of every scan and fix. Automation speeds up discovery, but only human judgment and real user testing confirm that a site is genuinely accessible.


TL;DR:

  • Automated scans should cover images, contrast, ARIA attributes, and document structure, with full rule sets activated for comprehensive detection.
  • Manual code checks are necessary to verify heading order, ARIA roles, keyboard navigation, and custom widget accessibility, as scanners cannot confirm correctness.
  • User testing with assistive technologies should involve realistic tasks to uncover real-world navigation and interaction issues, not just rule compliance.
  • Use structured sampling, including features like checkout and error states, to ensure the evaluation is representative and supports a reliable conformance claim.
  • Continuous monitoring and automated re-scans after fixes are essential to prevent regressions and maintain ongoing accessibility compliance over time.

AccessWiser
accesswiser.com
Make Accessibility Checks Part of CI
AccessWiser identifies WCAG 2.2 AA issues in code, explains how to fix them, and schedules re-checks to catch regressions.
Explore AccessWiser

Table of Contents

Run the tests in this order: automated, manual code checks, then assistive-technology user testing

A reliable testing sequence moves from broad and fast to narrow and precise. Each stage catches what the previous stages cannot.

  1. Automated scans sweep the whole site or app for rule-based violations: missing alt text, insufficient color contrast, empty form labels, and duplicate IDs.
  2. Manual code checks validate what the scanner flagged and look for what it cannot see, including logical heading structure, keyboard operability, and correct ARIA usage.
  3. Assistive-technology user testing confirms that real people, using screen readers, switch devices, or voice control, can actually complete key tasks.

For automated checks, enable rule sets that cover images, forms, color contrast, document structure, and ARIA attribute validity. Most scanning engines group these into categories, so turning on the full rule set rather than a narrow default catches more of the volume-level issues early.

Manual code checks should confirm that headings follow a logical order, that custom widgets expose the correct role and state through ARIA, and that every interactive element can be reached and operated using only a keyboard. This step matters because a scanner can confirm that an ARIA attribute exists without confirming that it’s the right one for the component’s actual behavior.

Assistive-technology testing works best as short, scripted tasks rather than open-ended exploration. Write three or four realistic tasks, such as completing a checkout form or registering an account, and ask a small group of testers who actually use screen readers, magnification, or voice control day to day to attempt them. Watching where they hesitate or get stuck tells you more than any rule-based report.

Tester navigating checkout with keyboard

Once all three streams produce findings, merge them into a single prioritized backlog. Each issue needs clear reproduction steps, the affected WCAG 2.2 success criterion, a screenshot or assistive-technology log, and a verification step so whoever fixes it knows exactly when it’s resolved.

Pro Tip: Tag each backlog issue with the WCAG 2.2 success criterion number so your team can track progress toward a specific conformance target rather than a vague pile of accessibility bugs.

  • Keep automated scan output and manual findings in the same tracker so nothing gets lost between teams, following best practices for ADA website compliance healthcare.
  • Re-run the automated scan after each fix batch to confirm nothing regressed elsewhere on the page.

Define scope and select representative samples using WCAG-EM

Before you test anything, decide what you’re actually testing and what conformance level you’re aiming for, typically Level AA. Trying to test every page on a large site is rarely practical, so the WCAG Evaluation Methodology gives evaluators a structured way to sample responsibly.

WCAG-EM lays out five steps: define the evaluation scope, explore the product to understand its structure and technologies, select a representative sample, evaluate that sample against the success criteria, and report the results. Skipping the exploration step is the most common shortcut, and it’s the one that produces unreliable samples.

A defensible sample mixes structured and random selection:

  • Include every distinct page template or component type (homepage, article page, product page, form).
  • Include complete processes end to end, such as checkout, registration, or appointment booking, not just their landing screens.
  • Add a few randomly chosen pages outside the structured set to catch issues that a template-based approach would miss.
  • Capture edge states like error messages, empty search results, and expired sessions, since these are frequently overlooked and frequently broken.

A sample-based evaluation supports a conformance claim only when the sample is genuinely representative and the full process, start to finish, has been checked, including dynamic states. A quick scan of a homepage and a contact page does not. When a site is large or changes frequently, WCAG-EM treats the sampling and reporting steps as something to repeat on a schedule, not a one-time exercise.

Prioritize these WCAG 2.2 criteria and common failures when testing

Not every success criterion carries equal risk. A handful of WCAG 2.2 criteria account for a disproportionate share of real-world failures, and testing time is usually limited, so start here.

  • 2.2.2 Pause, Stop, Hide: any moving, blinking, or auto-updating content that lasts more than five seconds needs a way to pause or stop it. Carousels and auto-advancing banners fail this constantly.
  • Target size (minimum): pointer targets on touch interfaces need enough space to tap reliably, a criterion new in 2.2 that specifically addresses mobile accessibility testing.
  • Focus visibility and order: every interactive element needs a visible focus indicator, and tabbing through the page should follow a logical, predictable order.
  • Consistent help and accessible authentication: new 2.2 criteria require that help mechanisms appear in a consistent place and that login flows don’t rely solely on cognitive tests like remembering a password with no alternative.

Section508 guidance cautions that automated tools provide partial coverage of these criteria at best, which means focus order, pause controls, and authentication flows usually need a human tester working through the page in sequence rather than a scanner report alone.

The W3C’s WCAG 2.2 techniques document lists specific failure patterns auditors should check for directly, such as a carousel with no pause button, a focus outline removed by CSS with nothing to replace it, or form fields that rely on placeholder text instead of a real label. Checking against these documented failure patterns, rather than general impressions, keeps findings consistent across testers and across audits.

Which tools to use, what they detect, and how to handle false positives

Automated scanners are genuinely useful for catching volume-level issues fast: missing alt attributes, low-contrast text, unlabeled form controls, and duplicate landmark regions. They’re far less reliable at judging whether an ARIA role matches what a component actually does, whether a heading structure reflects the page’s real meaning, or whether a keyboard trap exists in a custom widget. Section508 guidance describes this gap plainly: automated tools work like spelling checkers for accessibility, good at flagging the obvious, blind to meaning.

A practical triage workflow keeps scanner noise manageable:

  • Run the scan and separate results into clear violations (missing alt text, failed contrast ratios) and items needing manual review (ARIA usage, reading order).
  • Check for common false positives, such as decorative images correctly marked with role="presentation" getting flagged anyway.
  • Confirm true positives against the actual rendered page, not just the source code, since dynamic content can change after load.
  • Close out confirmed issues with a code fix and a retest, not a cosmetic patch.

Pairing a browser-based scanner with something like axe-core wired into unit tests catches new issues before they ever reach production, and a linter configured for accessibility rules flags problems directly in the code editor. Mobile apps need their own pass: automated tools built for the web often don’t read native mobile components correctly, so testing with a phone’s built-in screen reader, such as VoiceOver or TalkBack, remains necessary even when a scan reports a clean result.

Pro Tip: Run your automated scan after the page fully renders its dynamic content, not on the initial HTML, or you will miss real issues and flag ones that no longer exist.

Embed accessibility checks into design, development, and CI to stop regressions

Catching issues before launch is far cheaper than fixing them after. Build accessibility checks into the stages where a mistake is easiest to prevent.

  1. Design stage: give your component library accessibility acceptance criteria, such as minimum contrast ratios and required focus states, so designers catch problems before a single line of code exists.
  2. Development stage: add accessibility linters to the editor and unit tests built on tools like axe-core at the component level, so a developer sees a failure the moment they introduce it.
  3. Pre-merge CI checks: run an automated scan against the build as a blocking gate for serious violations and an advisory warning for lower-severity ones, so releases don’t stall over minor contrast tweaks.
  4. Post-deploy monitoring: schedule recurring site scans that alert the team when a new issue appears, since content updates and third-party scripts introduce regressions outside the normal release cycle.

Assign clear ownership at each stage: designers own component patterns, developers own implementation and unit tests, and a QA or accessibility lead owns final sign-off and the scheduled monitoring. Our solutions page walks through how scanning and monitoring fit into an existing development workflow without adding a separate manual process for every release.

Document test results and manage fixes so you can prove progress and compliance

A finding that isn’t documented is a finding that didn’t happen, as far as any compliance review is concerned. Every test report needs the same core elements: the scope and sample tested, the method used, step-by-step reproduction notes, screenshots or assistive-technology logs, and a date.

  • State the exact success criterion each defect violates, not just a general description.
  • Include a code pointer (file, component, or selector) so the developer fixing it doesn’t have to hunt for the problem.
  • Define a verification step so whoever retests knows exactly what “fixed” looks like.
  • Route every fix through a ticketing flow that closes with a retest, not just a code merge.
Report element Purpose
Scope and sample Shows what was tested and what the findings represent
Method Documents whether the check was automated, manual, or user testing
Reproduction steps Lets a developer confirm and fix the issue without guesswork
Verification criteria Defines what “resolved” means before retesting
Date Creates a timeline of progress for compliance records

Keeping dated scan and fix records over time is what makes an accessibility statement credible rather than aspirational. A formal conformance claim is appropriate once a representative sample, evaluated under WCAG-EM, shows no outstanding failures for the target level, and that claim should always reference the date and scope it covers, since conformance can change the moment new content is added.

AccessWiser example: scan, fix guidance, and scheduled re-checks

A platform builds the three-part workflow directly into its system rather than leaving teams to assemble it from separate tools. Each scan checks a site against WCAG 2.2 AA success criteria, with mappings to Section 508 and EN 301 549, so a single result speaks to multiple compliance frameworks at once.

  • Automated scans identify code-level issues and tie each one to the specific element and criterion it affects.
  • Every finding includes plain-language guidance for fixing the issue in the site’s own code, so the fix is permanent rather than a runtime workaround.
  • Scheduled re-checks catch regressions after a scan, which matters most during active development or frequent content updates.
  • Maintaining dated records of every scan and fix helps build the evidence base for an accessibility statement and a broader compliance program.

These records shorten remediation cycles because developers get specific fix recommendations instead of generic violation codes, and they give compliance officers something concrete to point to when documenting progress. Automated checks still catch a substantial subset of issues rather than all of them, which is why the manual and user-testing stages described earlier remain part of a complete workflow.

Why testing must be ongoing product work, not a one-time audit

Accessibility holds up only when it’s treated as continuous product quality work with real ownership and measurable goals, not a report filed once a year. Automation is necessary for catching regressions at scale, but it’s not sufficient on its own. The strongest programs pair routine automated checks with periodic human review and genuine user testing, scaled to risk rather than run uniformly everywhere.

— The AccessWiser Team

Getting started with AccessWiser

Running this workflow by hand across every page, every release, is a lot to ask of a small team. AccessWiser turns the testing playbook above into something you can run on a schedule: automated scans mapped to WCAG 2.2 AA, plain-language fix guidance tied to your actual code, scheduled re-checks that catch regressions, and dated records you can hand to a compliance officer without extra work.

Plans start with the Starter plan at a low monthly rate and scale up to higher-priced options for larger sites, with custom setups available for teams. Extra specific page scans and full site scans can be added on as one-off purchases when projects need additional coverage. Full details are on the pricing page.

If you’re ready to see how a scan maps to your own code, start with a trial run and review the findings against the workflow in this guide.

FAQ

What is accessibility testing?

Accessibility testing is the process of checking whether digital content and applications can be used by people with disabilities, including those using screen readers, keyboard navigation, or switch devices. It typically combines automated scans, manual code review, and testing with assistive technology to confirm both technical conformance and real-world usability.

What is digital accessibility?

Digital accessibility means designing and building websites, apps, and digital documents so people with disabilities can perceive, navigate, and interact with them. It covers visual, auditory, motor, and cognitive needs, and it’s typically measured against a published standard like WCAG.

What is WCAG 2.2 AA compliant?

A site is WCAG 2.2 AA compliant when it meets all Level A and Level AA success criteria defined in the WCAG 2.2 conformance guidance, evaluated across a representative sample of pages and processes. W3C and government guidance both recommend pairing that technical check with usability testing involving people with disabilities, since conformance alone doesn’t guarantee a page is easy to use.

What are the key guidelines for web content accessibility?

The core guidelines come from the W3C’s WCAG, organized around four principles: content must be perceivable, operable, understandable, and robust. WCAG 2.2 adds criteria addressing mobile pointer targets, focus visibility, and accessible authentication on top of the established 2.1 requirements.

Sources

Articles on this blog are general information about web accessibility, not legal advice. Laws change and their application depends on your specific situation — for decisions with legal consequences, consult a qualified legal professional.

Articles on this blog are general information about web accessibility, not legal advice. Laws change and their application depends on your specific situation — for decisions with legal consequences, consult a qualified legal professional.