Accessibility testing checks whether web applications meet WCAG, an international catalog of criteria with more than 80 checkpoints. Around 20 to 25 of them can be automated with conventional tools. AI language models can take on language-related criteria as well, such as whether headings match the page content. Clustering algorithms help pick meaningful representative pages from a large set.
Key Takeaways
- Accessibility testing today runs on an unproductive mix of Word, Excel, screen reader and browser, which should give way to an integrated test environment without constant switching between tools.
- Of more than 80 WCAG criteria, only around 20 to 25 can be tested automatically with existing standard tools. The rest still takes manual work.
- GPT-4 can reliably assess language-related WCAG criteria such as whether a heading fits its text, while visual checks like spotting broken layouts don’t yet work well enough with current models.
- A crawler combined with clustering algorithms cuts the manual effort of an accessibility audit by automatically grouping pages by relevant attributes, such as embedded PDFs or videos, and suggesting representatives.
Accessibility Testing Has Been Mandatory for Many Companies Since 2025
Since June 2025, the European Accessibility Act has required private companies as well to make their digital offerings accessible. AI accessibility testing can take over part of that work, mainly where language is involved. In the public sector, the obligation has applied for some time. Banks and insurers have been preparing for years, because breaking the rules has consequences.
The effort is considerable. Procurement portals list testing contracts worth many person-days. If you have to test a large web application, you’re not dealing with a few random samples but with a structured process across many pages.
The technical basis is WCAG, an internationally recognized catalog of criteria. Its more than 80 criteria describe what a website has to meet to count as accessible. National regulations and EU law build on this standard as well.
Which Accessibility Criteria Can Be Tested Automatically?
Of the more than 80 WCAG criteria, around 20 to 25 can currently be automated with standard tools. Open source tools cover these checks, and many teams already have them in their setup.
These tools mostly work with data that can be pulled from a page’s rendering. Contrast between text and background is a typical example: it follows directly from the rendered layout and can be checked objectively.
The rest stays manual. Realistically, no method will check all 80 criteria automatically anytime soon. Some of them need human judgment, for example where interpretation and context come in.
How an Accessibility Audit Works
An audit follows a defined process model that exists as a supplement to WCAG and describes the steps in broad strokes. The process is easy to follow and there’s nothing magical about it.
The steps:
- Capture the website. Get an overview of how the application is structured.
- Identify technologies. Check which technologies the site uses.
- Determine page types. Find out what kinds of pages exist, around 20 different types instead of 10,000 individual pages.
- Select representatives. Draw a sample for each page type and document every decision so the result stays traceable later.
- Check criteria. Go through the criteria catalog for each representative.
- Document. Write down the result.
Documentation runs through the entire process. Which page types were selected, and why, has to be recorded, or the audit can’t be verified.
The Biggest Pain Is the Tool Clutter, Not the Audit Itself
Anyone testing accessibility today is juggling a pile of tools. Findings go into Word, page lists into Excel, and the check itself needs a screen reader, an open source testing tool and, of course, the browser. Screenshots are taken, pasted, copied.
All that switching costs time and nerves. When many pages need checking, the back and forth between applications becomes the real obstacle, not the technical check.
A comparison with software development makes the problem clear. Developers start from a similar place: code, compiler, test frameworks, runtime environment, database tools. Modern IDEs bundle all of it into one interface through plug-in architectures.
That’s exactly the principle missing in accessibility testing. An integrated test environment where testers work without switching tools, link tools together and add new functions without leaving their working context would have the biggest impact. And by the way, most users simply expect a dark mode.
AI Accessibility Testing Helps Where Language Understanding Is Needed
AI shows its strength in accessibility testing mainly on language tasks. Current language models are already good in the language domain, and that’s where using them really pays off.
A concrete example is the criterion that a heading must match the text that belongs to it. Even for humans, that check isn’t trivial or clear-cut. There’s room for interpretation.
A workable approach looks like this: a tool extracts the headings on a page together with their text blocks and sends a matching prompt to a language model. The model rates how well heading and text fit together. The results of evaluations like this are consistently good enough to accept in production.
For all the enthusiasm, some restraint pays off. AI isn’t a universal fix for every testing problem. Anyone who slaps “AI” on everything today tends to look suspicious, because the topic comes with sky-high expectations. The better question is: what specific problem should the tool solve at this point?
Why Visual Defect Detection with AI Doesn’t Work Yet
Getting an AI model to reliably recognize whether a page looks broken didn’t work on the first attempt. The idea was appealing: feed in a screenshot, and the model reports a broken layout or other visual defects.
Models like that have to be trained first, and training needs labeled data. Over several weeks, artificially broken pages were shown to customers in a small web application, with the question of whether the page looked broken or not. The labeling ran as a contest with a prize draw.
Even with this training data, detection accuracy wasn’t good enough in the first round. The goal behind it is still attractive: a sensor that runs in the background during every test execution anyway and notices when the application under test is visually broken. Most automated tests don’t check non-functional aspects like this at scale.
How a Crawler with Clustering Speeds Up Page Selection
Getting an overview of every page in an application is one of the most time-consuming steps in an audit, and it’s where automation helps most directly. Building a map of all pages by hand is tedious.
A crawler that works through the application from one or more start URLs takes over the mapping. Crawlers aren’t new, and neither are their problems. How do you recognize that you’ve landed on the same page again when only a date, a time or an ad has changed? That state abstraction is the core difficulty, along with handling forms and form data.
Instead of making the crawler ever smarter, combining it with clustering helps. The crawler runs first. Then a clustering algorithm groups the pages it found by attributes that matter for accessibility, such as whether a page contains a PDF or a video.
That way you find meaningful representatives without testing ten similar pages twice. It saves work and hands the tester prepared groups: here all the pages with PDFs, there all the ones with video. Collecting and presorting at this stage takes a lot off the human tester’s plate, precisely because the manual check still comes afterward.
Why Local Language Models Aren’t Yet an Alternative to OpenAI
At the time, in 2024, OpenAI was the gold standard for the language checks, because no other model handled these tasks as well. That came with two problems.
First, the API connection wasn’t stable enough. Whenever the status page turned red, it hit your own application directly. Second, data privacy quickly raised the question of where the data actually goes, especially when customers are restricted to the German data space.
That’s why local models are an important topic for the future. The hope is to fine-tune a general open source language model such as Llama for the specific use case. In the first attempt, that didn’t work well yet. OpenAI was clearly ahead.
Deliver Iteratively and Involve Users Early
The most sensible route is to ship a first set of working tools and feed user feedback straight back in. Instead of waiting for the big, finished system, testers first get four or five additional criteria that can now be automated.
The message to users: here’s a set of tools that works. These criteria can be tested automatically on top of what was possible before. Work with it and tell us how it feels. Everything else follows step by step.
This keeps development close to users’ real problems. Every AI feature starts with the question of what the tool actually does better, not with technology for its own sake. One question remains open: how transparent should the use of AI be? Are users simply happy with a button labeled “Analyze,” or do they want to know what happens behind it?
Frequently Asked Questions
Who Has Been Required to Make Digital Services Accessible Since 2025?
Since June 2025, the European Accessibility Act has also required private companies to make their digital services accessible. This requirement has long been in effect in the public sector. Banks and insurance companies have been preparing for this for years because violations have serious consequences. The effort involved is considerable: test contracts requiring many person-days are appearing on procurement portals.
Why can’t accessibility be tested fully automatically?
Of the more than 80 WCAG criteria, about 20 to 25 are considered automatable using standard tools. These tools work with data that can be extracted from a page’s rendering, such as contrast values between text and background. The rest require human judgment because interpretation and context come into play. It is unrealistic to expect that all 80 criteria will be testable automatically in the foreseeable future.
How do you test a website with thousands of pages for accessibility?
Not every single page is checked. An audit first assesses the application’s structure and the technologies used, then identifies the page types (about 20 different types instead of 10,000 individual pages) and selects a representative sample from each type. Only then is the criteria checklist applied. Every selection decision is documented; otherwise, the audit cannot be verified.
What makes accessibility testing so time-consuming in practice?
It’s not the technical testing itself, but the mix of tools. Findings end up in Word, page lists in Excel; testing is done using screen readers, open-source testing tools, and browsers, plus screenshots for copying and pasting. For many pages, this back-and-forth between applications takes up the most time. Software development serves as a model here, where development environments bundle all tools into a single interface via plug-in architectures.
For which aspects of accessibility testing are language models suitable?
For language-related criteria. Even for humans, it’s not always clear whether a heading matches the accompanying text, and this is precisely where language models provide useful evaluations: A tool extracts headings along with text blocks and sends them to the model as a prompt. The results were good enough for production use. AI is not suitable as a one-size-fits-all solution for every testing problem.
Can AI use screenshots to detect whether a page is visually broken?
Not on the first try. For the training data, artificially broken pages were shown to customers over several weeks in a small web application, along with the question of whether the page looked broken. The labeling was conducted as a contest with a prize drawing. Nevertheless, the detection accuracy was insufficient. The idea of a sensor that detects visual defects in the background while test runs are already in progress remains appealing.
How does clustering help in selecting which pages to test?
A crawler starts from one or more initial URLs and maps the application; afterward, a clustering algorithm groups the found pages based on attributes that matter from an accessibility perspective, such as embedded PDFs or videos. This creates meaningful samples without having to perform testing on ten similar pages. The core challenge in crawling remains state abstraction: recognizing that only the date, time, or advertisements have changed.
Can such AI tests be run without cloud services?
In 2024, OpenAI was the gold standard for language testing because no other model handled these tasks as well. Two factors worked against it: an API connection that wasn’t stable enough, whose disruptions directly affected the application itself, and data privacy concerns for customers restricted to the German data space. Retraining an open-source model like Llama didn’t work well as a first step.


