Legacy testing with AI means testing legacy software that exposes no usable element IDs by driving it through a visual model. An AI tool analyzes screenshots and finds elements from natural language descriptions, without cropping reference images. The goal is to cut manual test cycles from several weeks to a few hours.
Key Takeaways
- Manual regression testing of Deutsche Bahn’s mobile checkout app tied up two to three people for two weeks. With AskUI, 60 automated test cases now run in about three hours.
- AskUI recognizes UI elements purely visually, from a screenshot and an AI model, without touching element IDs. That makes it usable for applications that classic frameworks such as Appium or Selenium cannot automate.
- Leaving element IDs out of the code blocks test automation later on. Deutsche Bahn therefore now requires service providers to deliver automatable software for every new or purchased application.
- The biggest open problem with AI-based UI testing is execution speed: every step needs an inference call to a model, which makes the suite much slower than conventional frameworks.
Why Legacy Testing Fails With Classic Automation Tools
Legacy testing hits a wall early: classic test automation tools often can’t drive legacy applications at all, because the technical anchors are missing. Selenium, Appium, and TestComplete all rely on element IDs. Without those IDs, the whole approach goes nowhere.
That is exactly what happened with Deutsche Bahn’s mobile checkout, a point-of-sale app. Staff on long-distance trains use it to sell coffee, beer, and other items and to take card payments. The software is Android-based, runs on .NET 6, and was bought in, not built in-house. Nobody thought about automation back then.
The result: no cleanly addressable elements and no standard tool that works. Several proofs of concept with open-source tools failed, simply because the application does not expose any element IDs.
Manual Regression Testing Takes Two to Three Weeks
Manual regression testing of the mobile checkout ties up two to three people for two weeks. That is where the pain starts.
It gets worse with every bug found. If bugs turn up after two weeks and a new version ships, the regression has to run again. In practice, that cycle is barely sustainable when releases are supposed to follow each other quickly.
For newly purchased or in-house software, Deutsche Bahn now starts earlier. Service providers and teams have to secure quality at every level, from unit tests and interface tests upward, before acceptance testing kicks in. Legacy applications like the mobile checkout never had that foundation, which is the root of the problem.
How Visual AI Selectors Drive Legacy Interfaces
AskUI does away with element IDs completely and works from screenshots instead. A controller, the so-called AgentOS, connects to an inference layer that bundles several models: OCR, image recognition, an LLM, and multimodal models.
At runtime, the tool takes a screenshot of the operating system and finds the controls on it. There is no element matching and no recorded click sequence as with older tools.
The difference from earlier image-based approaches matters. Old tools cropped out images of individual elements and compared them pixel by pixel. As soon as anything on the screen changed, the whole run broke. AskUI crops nothing.
Instead, the model has learned in advance what a login button looks like. You describe the action, not the appearance. Jonas Menesklou sums it up:
“The way ChatGPT understands text, we understand images.”
(Jonas Menesklou)
From Test Case to Execution: Two Paths
There are two ways to build test cases, depending on how technical the team is. For the mobile checkout, Deutsche Bahn takes the code path, using a framework with TypeScript and Node.js.
In code, the AskUI selector is just one more selector in the library. Instead of an element ID, you write a description in plain language: click the green login button in the top left corner. At runtime, that description becomes the selector. The library plugs into PyTest, TypeScript test runners, and others.
For less technical users there is a no-code path via a CSV file. If the test cases live in a test management tool with a test case ID and a step-by-step description in natural language, an LLM reads that description and turns each step into an action. You upload the file and start the suite.
The selector behaves like a virtual tester who understands user interfaces. You describe the element the way you would explain it to someone seeing the screen for the first time.
Maintenance Through Training Instead of Code Rewrites
When the interface changes, the effort to adapt stays small, because you don’t have to touch all the code. The tool picks up the new screens automatically and handles part of the change itself. The test logic only changes for the affected case.
For recognition errors there is a separate training tool. If a screenshot shows a blurry element and the model fails to recognize, say, “Registration,” you can teach the tool exactly that mapping. That way you extend the underlying model by hand when automatic recognition falls short.
Being able to fine-tune the model yourself sets this approach apart from pure black-box tools. You don’t have to wait for the vendor to cover every edge case.
What the Numbers From Practice Show
The mobile checkout has 270 longer test cases at acceptance level, 60 of which are automated so far. Those 60 run automatically in about three hours. Manually, one tester needs at least eight hours for around 50 cases.
The goal is far more ambitious. Once 210 to 220 of the 270 test cases are automated and run in parallel on several devices, the whole suite should finish in one to one and a half hours. That would free up about half of the current resources.
The mobile checkout can’t be fully automated. It uses peripherals: a printer for receipts and a device for reading credit cards. Those peripheral tests are not covered yet, but they are the smaller share.
Railway-specific devices with their own certificates add another hurdle. The app won’t run on just any hardware, and testing against emulators fails because of those certificates too.
Speed Is the Biggest Open Issue
The biggest weakness of the AI-based approach is performance. Because every step processes a screenshot and finds a selector visually, execution is noticeably slower than with a classic tool. Umar Usman Khan is clear about it: if the application could have been automated with Appium, he would have used Appium, even though AskUI needs less code.
Work on this is ongoing. The models run on Deutsche Bahn’s own inference infrastructure rather than in the cloud. That keeps control over the data, but costs communication time between server and tool. A newly built caching system keeps commands local instead of sending each one to the server.
Other levers are compression, smaller models, and as few inference calls per run as possible. The direction is clear: process as much as possible locally and cut down request-response traffic.
Why Partnering With a Startup Works Here
Working closely with a young vendor pays off for both sides. Missing features and problems from day-to-day operation go straight back to the vendor, and fixes and new functions arrive faster than an established tool would usually deliver them.
For that to work, you need backing inside the company. Proofs of concept cost money, and someone internally has to stand behind that investment. That was the case here, with support from team leads and department heads.
For you as a tester, this means a new, still unfinished tool can be the right choice when established tooling simply fails on the application. Going from two to three weeks of manual effort to a few hours is reason enough to help shape a growing product instead of waiting for the perfect off-the-shelf solution.
Frequently Asked Questions
Is it enough to plan test automation only after an application has been deployed?
No. If you don’t embed element IDs in the code, you’ll make it impossible to automate later on, because traditional tools rely precisely on those IDs. That’s why Deutsche Bahn already has a requirement that service providers ensure applications are automatable when they’re newly developed or purchased. Legacy applications like the mobile point-of-sale system lack this foundation, and that’s exactly where the problem arises.
Why does two-week manual regression testing become a problem with frequent releases?
Because the effort is repeated with every new issue found. The manual regression testing for the mobile point-of-sale system ties up two to three people for two weeks. If bugs surface afterward and a new version is released, the entire process must start over. With releases coming in rapid succession, this cycle is hardly sustainable.
How do AI-powered visual selectors differ from older image-based testing tools?
Older tools cropped individual element images and compared them pixel-by-pixel. If anything on the screen changed, the entire run would break. The AI approach does not crop reference images: the model has learned in advance what, for example, a login button looks like. The action is described, not the element’s appearance.
How do you describe test steps if an application doesn’t provide element IDs?
Using natural language. Instead of an element ID, the code contains a description such as “click the green login button in the upper-left corner,” which becomes the selector at runtime. There’s also a no-code approach: A language model reads step-by-step descriptions from the test management system and converts each step into an action.
How much maintenance is required if the user interface changes?
It remains minimal because the entire code does not need to be modified. The tool automatically captures the new screenshots and handles part of the change itself; only the test logic of the affected case needs to be adjusted. If the model fails to recognize a blurry element, the mapping can be manually refined using a dedicated training tool.
Are AI-powered UI tests slower than traditional frameworks?
Yes, significantly. Each step processes a screenshot and requires an inference call against a model. If inference runs on your own infrastructure rather than in the cloud, this provides data control but incurs communication time between the server and the tool. This is mitigated through local caching of commands, compression, smaller models, and as few inference calls per run as possible.
Can an application with connected peripherals be tested fully automatically?
No. The mobile POS system uses a printer for the receipt and a device to scan credit cards; these peripheral tests are not covered, but they account for only a small portion of the process. In addition, there are rail-specific devices with their own certificates: The app does not run on just any hardware, and testing against emulators fails due to these same certificates.
When is it worth relying on a tool from a young provider that is not yet fully developed?
When established tooling simply fails to handle the application. The shift from two to three weeks of manual effort to just a few hours justifies actively helping to shape a growing product. This requires support within the company, because proof-of-concept projects require an investment that someone internally must help cover (in this case, team leaders and department heads).


