AI quality assurance means assessing quality statistically instead of with a clear yes or no. Classic quality characteristics such as functionality, performance and portability still apply, but AI adds its own criteria: functional performance metrics, autonomy, transparency and ethics. Because there is no clear test oracle, techniques such as metamorphic testing and pairwise testing take on new importance.
Key Takeaways
- With AI, classic quality characteristics such as functionality are no longer a yes-or-no statement. They become statistical measures, tracked through metrics such as accuracy, precision and sensitivity.
- Metamorphic testing addresses the test oracle problem: if you don’t know the expected output, you vary the input in a way that must leave the output unchanged, and gain verifiable confidence in the system.
- AI systems learn from the data, not from the intention behind it. An image classifier trained on photos with timestamps learns the timestamp instead of the actual image content.
- Reproducibility is structurally hard in AI systems, because training parameters are often set by random number generators and not every state can be captured exactly.
- Established techniques such as pairwise testing and A/B testing gain new relevance with AI, because huge numbers of parameter combinations and missing test oracles call for exactly their strengths.
AI Quality Assurance: Why Quality Is No Longer a Yes-or-No Question
In classic software, a test case gives you a clear result: pass or fail. AI quality assurance loses that clarity. Functionality becomes a statistical measure, expressed through values such as accuracy, precision and sensitivity.
Testers are trained to rate every test step as “right” or “wrong.” With AI, that step-by-step verdict is often missing entirely. At the end of a run across many test steps, you may be able to say “that was good,” but there is no hard statement about the individual steps in between.
The core of the problem has a name: the missing test oracle. In classic software, a specification prescribes how the system must behave. With AI, that clear yardstick no longer exists. There is only how the system should normally behave. Sometimes it doesn’t, and that can still be fine.
An image classifier usually recognizes an apple as an apple. It is allowed to miss one now and then. That is exactly why a test case can’t read: “Feed in this one image, ‘apple’ must come out, checked once, done.” You need lots of images instead, and the question of how many are enough to build trust.
AI Is Still Software, with Extra Quality Characteristics
The classic quality characteristics still apply. Functionality, performance, portability: these staples of quality management stay relevant, because an AI is embedded in a larger overall system. Performance testing, usability testing, regression testing and requirements analysis don’t go away.
What’s new are characteristics that classic software didn’t have in this form. Autonomy describes how independently a system can act, how long it does so under which conditions and when it hands back control. An automated car doesn’t drive itself forever. At some point the lane assist chimes in and asks you to put your hands back on the wheel.
Ethics becomes something to test. In an accident, a decision has to be made, and testers will increasingly have to question the requirement behind it. What should the AI really do at that moment? Should it decide on its own or hand control back to the human?
Transparency matters especially to testers. Without insight into why an AI decides the way it does, quality is hard to measure. The research field behind this is called explainable AI, or XAI.
How an AI Learns the Wrong Thing
An AI learns from the data it gets, and sometimes it learns the wrong thing. Garbage in, garbage out: feed a network misleading data, and you end up with a problem that is hard to spot.
A real case shows how this happens. An AI was supposed to use image recognition to read the settings on a smart heating control. Accuracy stubbornly stayed at 40 to 45 percent, 68 percent in one case, which was too low for use.
A heat map revealed the cause. Every photo carried a timestamp from the camera. The AI had learned the timestamp and nothing else. Hidden correlations like this are exactly why testers need to dig into the input data before they interpret any results.
There is a language trap here, too. In the AI workflow, “test data” means the data used to train a network. For classic testers, test data is what you test with. Two different things share one word, and keeping the terms apart is part of getting up to speed.
Metamorphic Testing: Confidence Without a Known Result
Metamorphic testing addresses the oracle problem by deliberately changing inputs in ways that must not change the expected output. You don’t know exactly what should come out, but you know the output has to stay the same.
A triangle shows the principle. Extend all sides by five centimeters, and it is still a triangle. This metamorphic relation describes a change that must not affect the result. If the AI handles it correctly, you gain some confidence, even without knowing the absolute expected result.
The technique is especially useful for uncovering hidden misprioritization. If an image classifier learns shapes that each appear in their own color, say a red circle, a blue rectangle and a yellow triangle, it may end up learning the color instead of the shape.
So in the test, you vary the color. If the circles were red, you show the AI a red rectangle. If it classifies that as a circle, it is using color as a feature, even though color shouldn’t matter. The metamorphic relation varies irrelevant properties on purpose to flush out exactly these errors.
The more metamorphic relations you apply, the more confidence you gain. The technique isn’t new, but it becomes more important wherever the test oracle is missing.
Proven Test Techniques for AI Systems
Not everything has to be reinvented. Several established test techniques apply directly to AI systems, because their strengths fit these old problems particularly well.
- Pairwise testing: AI systems have a great many parameters and therefore countless parameter combinations. Testing them all is impossible, so pairwise combination keeps the number of test cases manageable.
- A/B testing: where the expected result is unclear, you can develop two AIs with the same goal and have two user groups evaluate them. That compensates for part of the missing oracle.
- Metamorphic testing: change inputs without changing the expected output, and check the behavior.
These techniques were known years ago. In the AI context, they gain new weight, because the missing oracle and the high number of parameters are exactly the gaps they fill.
Why Test Environments for AI Take So Much More Effort
In classic software, inputs and outputs are usually structured: databases, tables, sensor data from physical systems. For AI applications, the range of inputs grows enormously, and a small set of test data is no longer enough.
Autonomous driving requires simulating an entire environment: cities, streets, buildings, pedestrians, other vehicles. That simulation serves not only classic components but above all the AI components, with their multimodal input from cameras, lidar and radar sensors.
Real-world testing is barely affordable. Running through scenarios with extras and elaborate setups in reality isn’t feasible, which is why there is a strong focus on virtualization. Computing power, data provisioning, reusability and anonymization of the data add to the effort.
Reproducibility Is Not a Given with AI
Reproducing results exactly is extremely hard with AI, because a lot of training relies on random number generators. The parameters of neural networks are often initialized at random, and that randomness has to be repeatable.
Frameworks can save random seeds, which helps, but not always. Not every state is captured, and some results simply can’t be reproduced.
Scenario-based testing shifts the goal. Instead of repeating every run bit for bit, you describe the scenario as precisely as possible so that the underlying principle is reproducible. Since statistics determine the result, not every single run is reproducible, but the statistics as a whole are.
A driving scenario consists of many parameters at once: the car is approaching a red light, a truck is coming from the right, a bus from the front, a woman is standing on the left, a child is playing on the right, plus weather and sunlight. What used to be a short list of constraints becomes a long list of real parameters that describe the scenario.
Today’s Tools Test with AI More Than They Test AI
Right now, the tool market focuses more on building AI into testing tools than on supporting the testing of AI. The urge to play with new technology is strong and very human, so testing with AI actually came before testing AI.
For testing AI itself, generative models help when there isn’t enough test data. They produce text for chatbots, images for person recognition or synthetic data for finance, for example. Whether they were built for this purpose or not, you can use them.
A practical way to start is to have a language model write a test for you. Good testers who get disappointed quickly shouldn’t give up but should play with their prompts.
“AI won’t replace us. It will make our work easier. And the work won’t get any less.”
(Gerhard Runze)
Standards for Testing AI Are Still Taking Shape
Standards for testing AI are only now emerging. It started with an A4Q syllabus, which was replaced by the ISTQB CT-AI syllabus, published a few years ago and later followed by a German version.
In parallel, the German standardization roadmap on AI is working on the topic. The federal government initiated it, and it was presented at the Digital Summit. Now in its second edition and written by several hundred authors, it shows where standardization is needed. It covers domains such as finance, medicine and transport as well as cross-cutting topics, such as the impact of the EU AI Act.
Reproducibility remains an open flank. How to standardize something that can’t be reliably reproduced hasn’t been fully thought through yet. For large language models, a group has already formed to structure the needs and get standardization started.
What Testers Need to Change Now
For AI, testers need a new mindset, not an entirely new craft. Statistical thinking replaces the old yes-or-no logic: probabilities, inaccuracies and confidence levels take the place of one clear expected result.
The focus shifts to the data. Testers need to understand what data a network was trained on, which hidden features it might have picked up and whether the training data actually reflects what the AI is supposed to learn.
At the same time, the classic foundation still carries the weight. If you know CTFL, you have the basis. The AI-specific techniques build on it, from metamorphic testing to pairwise and scenario-based approaches. What is really new isn’t the tool. It is the willingness to think about quality without a fixed oracle.
Frequently Asked Questions
Is a single test case enough to validate an image classifier?
No. A classifier usually recognizes an apple as an apple, but it may occasionally misclassify it without that being considered an error. That is why a statistical evaluation based on a large number of images replaces the simple pass/fail assessment. Metrics such as accuracy, precision, and sensitivity are used for evaluation. The real question is how many cases are needed to establish robust confidence.
What quality characteristics, in addition to the traditional ones, are relevant for AI systems?
Autonomy, transparency, ethical considerations, and functional performance metrics. Autonomy describes how long a system operates independently under which conditions and when it relinquishes control, for example when the lane-keeping assist system prompts the driver to place their hands on the steering wheel. Transparency is part of the research field of Explainable AI. Functionality, performance, and portability remain relevant because AI is embedded in a larger overall system.
Does the term “test data” mean the same thing in the AI context as it does in traditional testing?
No. In the AI workflow, test data refers to the data used to train a neural network. For traditional testers, test data is what is used to perform testing. Two different things share the same term. Clearly distinguishing between these concepts is part of the onboarding process when testing professionals transition to AI projects.
How do you test a system when no one knows exactly what the correct result should be?
Through metamorphic relations. This involves deliberately modifying the input in such a way that the expected result must not change: If you extend all sides of a triangle by five centimeters, it is still a triangle. If the system responds correctly, trust is built even without an absolute target result. The more relations are tested, the more robust the picture becomes.
Does testing AI require completely new test techniques?
No. Pairwise testing keeps the number of test cases manageable when a system has a very large number of parameter combinations. A/B testing helps when the desired outcome is unclear by developing two AIs with the same goal and having them evaluated by two user groups. Metamorphic testing complements both. These methods have been known for years and, in the context of AI, fill precisely these new gaps.
How can you tell if a neural network has learned the wrong feature?
Through heatmaps and purposefully varied inputs. In an image recognition project for a smart heating control system, accuracy remained between 40 and 45 percent; the heatmap showed that the AI had learned the camera’s timestamp. If a classifier learns color instead of shape, a rectangle in the circle’s color that gets recognized as a circle reveals the misprioritization.
Can AI functions for autonomous driving be realistically tested on the road?
Hardly. It’s not affordable to actually run through scenarios with extras and elaborate setups, so testing is heavily virtualized. Entire cities are simulated, complete with streets, buildings, pedestrians, and vehicles, because the AI components process multimodal inputs from cameras, lidar, and radar sensors. Computing power, data provision, reusability, and data anonymization further drive up the costs.
Can the training results of an AI be reproduced exactly?
Often not. The parameters of neural networks are frequently set using a random number generator. Frameworks can save random seeds, but not every state is recorded. In scenario-based testing, therefore, the goal shifts: the scenario is described in such detail that the underlying principle remains repeatable. What is reproducible, then, is the statistics as a whole, not every individual run.


