An AI audit needs a clear picture of what can be tested in an AI system at all. The AI Assessment Matrix provides that picture: it is a framework that structures the testing and certification of AI systems. Test dimensions, from technical performance through robustness and fairness to environmental impact, run along the X-axis and are mapped against the data and model lifecycle on the Y-axis. The goal is a complete overview from which targeted testing decisions can be made.
Key Takeaways
- The AI Assessment Matrix from TÜV AI Lab arranges AI test criteria along two axes: test dimensions, from technical performance to global environmental impact, against the phases of the data and model lifecycle.
- AI testing takes three forms: direct testing of the product, evaluation of the provider’s documentation, and assessment of processes and people. All three require a solid understanding of testing.
- The EU AI Act doesn’t regulate every test dimension to the same degree. Explainability, for example, is only hinted at, because reliable test methods for it don’t fully exist yet.
- Fairness and non-discrimination are different requirements, legally and conceptually. They can contradict each other, so they have to be defined and tested separately.
- The energy consumption of hardware and software and the working conditions in data labeling belong in the matrix as environmental and social criteria, alongside the technical ones.
Why an AI Audit Needs Its Own Testing System
AI is a powerful technology, and power cuts both ways. Whatever can do a lot of good can also do harm. That is where every AI audit starts: how do you bring innovation and safety together without sacrificing one for the other?
Christoph Poetsch of TÜV AI Lab describes this as a mission for Europe. Trustworthy AI needs public support, and that support only comes when it’s clear what gets tested in an AI system, what it is supposed to be like and what it must not do.
The comparison with TÜV’s traditional role goes further than you might think. TÜV once dealt with steam boilers. Today the AI system is the steam boiler of the 21st century: something whose effects you want to keep under control without choking off its benefits.
For testers, this breaks with familiar logic. Traditional testing relies on clear steps and an expected result. An AI system produces outputs that aren’t fixed in advance. That gap between the clarity testers expect and the behavior they actually get is why AI testing and certification need their own structure.
AI Is Not Just a Technical System but a Quasi-Actor
The key conceptual step is to treat AI as a quasi-actor. As long as a system only performs functional tasks, functional safety is the right frame, the classic case for testing. Once a system takes on tasks that would otherwise need human judgment, the testing requirements shift.
An AI system in an HR process decides on job applications. It does something a person used to do. At that point, a purely technical view is no longer enough, because decisions with social consequences are now involved.
Even so, testing stays tied to technical reality. Nobody can have a person talk to an AI system for two years and then give a gut verdict. In the end, the assessment has to be measurable, reproducible and technically feasible.
What the AI Assessment Matrix Is
To make that possible, TÜV AI Lab developed the AI Assessment Matrix, a framework that brings test methods, metrics and benchmark data into one ordering structure. The matrix is built as a two- to three-dimensional system.
The X-axis holds the test dimensions, the properties measured on the AI system. Poetsch describes them as the sensors you hold up to the system, each one sensitive to a different aspect such as performance, robustness or fairness.
The Y-axis holds the areas of testing along the software lifecycle, from inception to retirement. This lifecycle is doubled, because with AI the data lifecycle runs alongside the model lifecycle. The role of data in development is a clear difference from traditional software engineering.
One detail is easy to misread: the Y-axis doesn’t describe when you test, it describes what you focus on. If you want to test robustness based on design decisions, you need documentation from the design phase, but you do the testing later. If you want to assess a training data set, it has to still be available.
Combining the two axes produces a maximum grid. It is explicitly not meant to be filled in completely, spreading test resources evenly across every cell. The point is a complete overview from which you consciously choose the fields that matter for your AI assessment.
“The idea is not to fill this maximum grid evenly with testing resources, but to know deliberately: can we get something like a complete overview, from which we then say, now we’ll focus on certain aspects.”
(Christoph Poetsch)
How Zooming Out Orders the Test Dimensions
The real innovation lies on the X-axis. Discussions about trustworthy AI often end in a bouquet of demands: robust, fair, high-performing, sustainable. What’s usually missing is the question of how these criteria relate to each other and whether the list is complete.
The matrix orders these criteria with the image of zooming out. The starting point is the individual AI system. From there, the view widens step by step until it reaches a global scale, and each zoom level brings its own test dimensions into focus.
At the very center are questions that are deliberately left out, such as autonomy or a conscious inner life. One step further, where the system acts on the outside world, performance and safety come into focus. Where something from outside acts on the system, the concern is robustness against random influences such as bad weather or dirty road signs, and cybersecurity against targeted attacks.
Add an individual human and the epistemic area opens up: explainability and transparency, split by what laypeople and what experts can understand. In the opposite direction, where the system acts on the person, privacy and nudging come into view.
With several individuals, the ethical questions arise: fairness, non-discrimination, bias. Here the AI system shows up as something that distinguishes between two people and decides who gets a job. At the level of society, legal questions of accountability follow. At the global level come supply chain responsibility, working conditions in data labeling and the energy and resource consumption of the hardware.
This overview sums up the logic of the zoom levels:
| Zoom Level | Direction | Example Test Dimensions |
|---|---|---|
| AI system acting outward | Effect of the system | Performance, safety |
| Influence on the system | From outside onto the AI | Robustness, cybersecurity |
| System and one individual | Human understands AI | Explainability, transparency |
| AI acting on an individual | AI influences the human | Privacy, nudging |
| Several individuals | AI distinguishes between people | Fairness, non-discrimination, bias |
| Society | Responsibility | Accountability |
| Global scale | Humans and ecosystem | Supply chain, resource consumption |
The EU AI Act Doesn’t Regulate Everything, and That’s Deliberate
One finding from working with the matrix contradicts a common assumption: the EU AI Act does not regulate every aspect. If you map the requirements that address the AI system directly onto the matrix, some fields stay empty.
Explainability in particular gets only hints, even though much more could be demanded technically and in substance. Behind this is deliberate restraint. The law doesn’t require anything when it’s unclear whether and how it can be done technically.
That honesty isn’t a flaw. AI is evolving at a speed that would quickly overtake any regulation. Writing requirements into law today that nobody can meet would harm both safety and innovation.
Three Forms of AI Audit: Product, Documentation and Process
The third dimension of the matrix distinguishes how testing is done. The first form is direct product testing. The AI system goes on the test bench like a car, and the questions are how to apply the measuring tool and which thresholds apply.
The second form is testing based on documentation. In many places, the EU AI Act provides for an assessment of the technical documentation. The provider runs its own accuracy tests, and the auditor checks whether the results are plausible.
This second form still needs a full understanding of testing. You have to be able to judge whether the right test was used, whether the values are plausible and whether the interpretation holds. Without that understanding of the content, documentation can’t be assessed seriously.
The third form covers processes and people. Risk management and quality management play a big role anyway. On top of that comes human competence, anchored in AI literacy under Article 4 and in human oversight. Article 26 requires deployers to ensure the expertise of the people providing oversight, which raises the question of what criteria you use to test human competence.
Fairness Is Not the Same as Non-Discrimination
A precise set of definitions is the basis of any reliable testing. Buzzwords aren’t enough. Every test dimension needs a definition, aligned with international standards and the EU AI Act, so the whole set stays consistent.
The difference between fairness and non-discrimination shows why this matters. In its articles, the EU AI Act only speaks of non-discrimination, once, in Article 10. Fairness appears only in the recitals.
Non-discrimination means what the law requires, for example under Germany’s General Equal Treatment Act (AGG) or the EU Charter. Fairness, depending on the reading, covers different concepts for individuals and groups that go beyond the law and can even be at odds with it.
As soon as two notions of fairness contradict each other, no system can satisfy both at once. So before every test, it has to be settled which concept is meant. You can define what you mean by fairness. Which concept of fairness is the right one remains a different, very old question.
Why Philosophy Helps with AI Testing
AI works like a magnifying glass for questions humanity has been asking for thousands of years. Because an AI system develops something like cognitive capacities, old questions about justice, understanding and responsibility can be looked at again in sharper form.
Both sides gain from this double perspective. The technical view is sometimes too quick to use the black box label. Strictly speaking, a neural network isn’t a black box, because all the information about the network is available. What’s unknown is why that information produces exactly this behavior, not the information itself.
The philosophical view, in turn, brings centuries of research on concepts such as justice and fairness, which the AI debate needs. Bringing both disciplines together lets you order a testing field with the necessary depth instead of covering it with buzzwords.
The next step for the matrix is to work from the top down. The top-down design needs to become concrete, because you don’t test robustness the same way for every system. The challenge is finding the right altitude: as general as possible, but as concrete as necessary, so that in the end real testing can happen.
Frequently Asked Questions
Why Can’t an AI System Be Tested Using Traditional Test Techniques?
Traditional testing relies on defined steps and a predetermined expected outcome. An AI system, on the other hand, produces outputs that are not predetermined. It is precisely this gap between expected clarity and actual behavior that necessitates a dedicated testing methodology. Nevertheless, the evaluation remains tied to technical reality: it must be measurable, reproducible, and feasible, not a gut feeling based on prolonged use.
When is functional safety no longer sufficient as a testing framework for an AI system?
As long as a system performs only functional tasks, the framework of functional safety is sufficient. As soon as it takes on tasks that would otherwise require human judgment, the testing requirements shift. An AI system in an HR process that decides on job applications thus becomes a quasi-actor. Decisions with societal implications can no longer be adequately evaluated from a purely technical perspective.
Why is it necessary to consider two lifecycles in AI rather than just the software lifecycle?
In AI, the data lifecycle stands alongside the model lifecycle, from the inception phase through retirement. The role of data in the development process is the clear distinction from traditional software development. It is important to note that the lifecycle axis does not specify the timing of testing, but rather the focus. To evaluate a training dataset, its availability must still be ensured; design decisions are reviewed later based on documentation.
Does an AI audit have to cover all combinations of audit dimensions and lifecycle phases?
No. The combination of both axes results in a maximum set that should explicitly not be filled with audit resources using a scattergun approach. The benefit lies in a comprehensive overview: based on this, one makes a conscious decision about which aspects the evaluation should focus on. Without this structure, demands for robustness, fairness, or sustainability remain a disjointed collection of buzzwords.
Are energy consumption and working conditions part of an AI assessment?
Yes. At the global zoom level of the matrix, supply chain responsibility, working conditions in data labeling, and the energy and resource consumption of hardware and software are listed as separate assessment dimensions. They complement the inner levels, which include performance and safety, robustness and cybersecurity, explainability and transparency, privacy and nudging, and fairness and accountability.
Does the EU AI Act cover all requirements for trustworthy AI?
No. If you map the requirements directly addressing the AI system onto the evaluation matrix, some fields remain empty. Regarding explainability, there are only vague references, even though more could be required in terms of both technical specifications and content. This reflects a deliberate approach of restraint: nothing is required if it is unclear whether or how it is feasible. Unmeetable requirements harm both safety and innovation alike.
Is it sufficient to evaluate the provider’s technical documentation for an AI assessment?
Documentation review is one of three forms of testing, alongside direct product testing on the system and the evaluation of processes and personnel. However, it requires a thorough understanding of testing: You must be able to assess whether the correct test was applied, whether the values are plausible, and whether the interpretation holds up. If the provider conducts its own accuracy tests, the plausibility of the results is verified.
Can an AI system meet all fairness requirements at the same time?
No. As soon as two concepts of fairness contradict each other, it is impossible to build a system that satisfies both. Therefore, it must be clearly established before each test which concept is being referred to. Nondiscrimination refers to what is required by law, such as under the AGG or the EU Charter. Fairness encompasses broader concepts for individuals and groups that may conflict with the law.


