An AI test description is the structured documentation of test cases for AI systems, organized by test objective, test steps and acceptance criteria. It rests on two dimensions: the capabilities of an AI system (perception, processing, action, communication) and method-specific quality criteria such as correctness, robustness and protection against bias.
Key Takeaways
- An AI test description is structured like code: test objective, test steps, a threshold such as a confidence score of 0.5, and at the end the acceptance criteria that decide when the model is assumed to be correct.
- AI systems can be tested along five processing capabilities: identification, classification, extraction, selection and generation. They apply equally to image, sound and language processing.
- The EU AI Act has been in force since August 2024 and requires makers of high-risk AI systems to meet standards on ten topics, including correctness, robustness, cybersecurity and risk management.
- General purpose AI systems such as GPT-4 face stricter documentation and security obligations once their training compute exceeds the threshold of 10^25 FLOPs.
- Under European law, notified bodies for AI conformity assessment had to be designated by August 2025. The European standards meant to underpin them were already behind their April 2025 deadline at the end of 2024.
AI Test Description: How Standards Make AI Systems Testable
AI systems become testable once you translate their capabilities and quality criteria into a structured test description with a test objective, test steps and acceptance criteria. Taras Holoyad, who works in telecommunications regulation at Germany’s Federal Network Agency (Bundesnetzagentur), is helping to write the standards behind that approach. His agency regulates AI through standards, not bans.
The Federal Network Agency is the regulator for telecommunications, postal services, railways, electricity and gas. In telecommunications, it monitors the market for radio equipment such as Bluetooth radios and cell phones and works on the conformity assessment of products for the European single market.
With AI regulation, the job is shifting toward standardization. Since August 2024, a European law on artificial intelligence, the EU AI Act, has been in force, and every company entering the market has to implement it. The standards Taras is working on turn that law into concrete requirements.
The focus is on high-risk systems and general purpose AI systems. Standards are meant to define which requirements a product has to meet before it may go on the market. The European Commission has issued a standardization mandate with ten topics, and authorities are working on them together with industry, consulting firms and certification bodies.
Why Standardizing AI Is So Hard
Artificial intelligence is hard to standardize because the technology moves faster than committees can work. Many call it old wine in new bottles. NASA was already using neural networks in space programs in the 1990s.
According to Taras, there’s little that is truly groundbreaking. The Transformer method is more accurate than older approaches. Even so, AI doesn’t come close to human intelligence. It’s more of a very elaborately wired algorithmic system.
That gap is exactly what makes standardization tricky. Specify too much detail, and new kinds of systems may no longer be testable in a meaningful way. Keep the standard too abstract, and it’s no help to the tester. In Taras’s view, the decisive research breakthrough is still missing, and that shapes every decision about how specific a standard can get.
Testing AI Means Testing Its Capabilities
You test the functional range of an AI system through its capabilities, not by asking whether it’s intelligent. That idea is at the heart of a standardization approach that describes AI along two dimensions.
The first dimension is the methods implemented in algorithms: classic AI with optimization and planning methods, symbolic AI with knowledge representation, machine learning, and hybrid methods that combine rule-based and data-driven approaches.
The second dimension is the capabilities those algorithms deliver. They include perceiving images or smells, processing knowledge, acting (robotically or in software) and communicating, the way a ChatGPT-style system does.
For processing inside AI models, five basic capabilities stand out: identification, classification, extraction, selection and generation. Metrics can be defined for each of them, the same way for sound, images or natural language. If you want to test a Hugging Face model, you can do it along these five capabilities, repeatably and at scale.
This approach is part of the international standard ISO/IEC 42102, which Taras leads. It’s being developed with colleagues from France, the USA and Germany. Through the Vienna Agreement between ISO and CEN, an international standard and a European standard are being developed in parallel. In terms of content, the standard maps to the transparency topic of the standardization mandate.
Quality Criteria Make AI Testing Tangible for Testers
Quality criteria give testers a familiar lever for making abstract AI requirements measurable. A second standard describes which criteria apply to individual methods once the algorithms are implemented.
For supervised, unsupervised and reinforcement learning, five quality criteria can be defined:
- Correctness
- Robustness
- Avoidance of unwanted bias
- Protection against adversarial attacks
- Information security
Each criterion comes with method-specific metrics. For correctness in supervised learning and image recognition, for example, you can use the confidence score. The result is a kind of matrix you can lay over an AI system to check it in concrete terms.
The two standards interlock. One describes what artificial intelligence actually is, meaning methods and capabilities. The other defines which quality criteria can be verified. There’s a practical reason the level of detail is spread across several documents: for political reasons, the different interested parties don’t always allow every level of detail to go into a single document.
A Test Description Language Structures AI Tests Like Code
A test case for AI can be written as structured text, much like a function in program code. At the time of the conversation, Taras was working on this as a new proposal, while the other two standards were already well advanced. The inspiration is the Test Description Language from the ETSI committee MTS (Methods for Testing and Specification), where descriptions of this kind were developed for protocol tests in mobile communications and the automotive industry.
The idea: you see at a glance what it’s about. Instead of a function definition with def as in Python, you write the defined syntax directly. A testcase in curly braces, for example Vehicle Recognition, with the Test Objective, the Test Activities and the Test Steps below it.
An example from image recognition: a model reads every frame of a video, runs inference and classifies objects. The structured text can set a threshold, such as a confidence score of 0.5. Only once that value is reached is an object classified. At the very bottom are the acceptance criteria that decide when the model is assumed to be correct in the first place.
“If you commission me to run a test, I take my two standards and use them to put together a structured text with the test description.”
(Taras Holoyad)
After the conversation, Taras planned to launch this test description as a new document at ETSI MTS, most likely as a technical specification. The related ETSI report carries the number 103 910.
When the EU AI Act Affects You
Whether the EU AI Act applies to you depends on which of two categories your product falls into. The law has been in force since August 2024 and is being phased in over set time windows. Enforcement runs through market surveillance authorities, each responsible for a segment such as medical devices, toys or radio equipment. Around 13 segments are to be covered in total.
The first category is high-risk AI systems. These are systems that are part of a safety component. Taras draws a clear line between safety and security: safety protects people from the machine, security protects the machine itself. Anyone operating a high-risk system has to meet the standards from the mandate’s ten topics or else go through a certification body.
The ten topics of the standardization mandate include correctness, robustness, cybersecurity, quality management, conformity assessment, transparency and risk management. The mandate goes out as a standardization request to the organizations CEN, CENELEC and ETSI, in this case only to CEN and CENELEC.
The second category is general purpose AI systems, meaning systems with an especially broad range of functions, such as ChatGPT. An extra threshold applies here: if the compute used in training exceeds 10^25 FLOPs, the system counts as a general purpose AI system with systemic risk. The FLOPs figure comes from an algorithm-specific constant, the token length of the training data and the number of parameters. In Taras’s assessment, GPT-4 has crossed that line. The consequences are heavier documentation requirements and additional preventive cybersecurity measures.
Everyone Is Under Time Pressure
The tight schedule puts authorities, manufacturers and certification bodies under pressure at the same time. Where no standards exist or applying them isn’t enough, manufacturers have to go to a certification body. That body reviews the test results and issues a CE mark for market access.
But a certification body may only do that once it has been notified. In every EU member state, a responsible authority has to assess the body together with an independent expert. Only after that assessment does it become a notified body that can help with market access.
At the time of the conversation, at the end of 2024, the deadlines were ambitious. Notified bodies had to be designated by August 2025, which meant the authorities’ assessments had to be finished by then. By August 2026, these bodies were supposed to have built enough expertise to test high-risk systems.
European standardization itself was also behind schedule. The deadline had been set for April 2025, but the official CEN and CENELEC website already showed time windows in 2026 back then. Some standards were being pushed through heavily accelerated procedures to get something suitable to the Commission in time.
Frequently Asked Questions
Is artificial intelligence really as new a technology as the debate suggests?
Much of it is old wine in new bottles. As early as the 1990s, NASA was using neural networks in its space programs. Compared to older approaches, the Transformer method delivers greater accuracy, but it remains a highly complex algorithmic system and does not reach the level of human intelligence. From Taras Holoyad’s perspective, the decisive research breakthrough has yet to come.
How specific should a standard for AI systems be?
It walks a fine line between two pitfalls. If too much detail is specified, it may no longer be possible to perform effective testing on novel systems. If the standard remains too abstract, it is of no help to the tester. A practical solution is to distribute the level of detail across multiple documents, partly because different stakeholders will not politically allow every level of detail to be included in a single document.
How can an AI model be tested without assessing whether it is intelligent?
Through its capabilities. At the highest level, these are perception, processing, action, and communication. For processing within models, five capabilities can be identified: identification, classification, extraction, selection, and generation. Each of these capabilities has associated metrics that apply uniformly to sound, images, and natural language. This also makes a model from Hugging Face testable in a repeatable and scalable manner.
What quality criteria can be demonstrated for machine learning systems?
Five criteria apply to supervised, unsupervised, and reinforcement learning: correctness, robustness, avoidance of unnecessary biases, protection against adversarial attacks, and information security. The corresponding metrics depend on the method. For example, the confidence score is suitable for assessing correctness in supervised learning for image recognition. Together, these form a matrix that can be applied to a specific system.
What are the benefits of describing AI test cases in structured text rather than in prose?
A test case becomes readable at a glance, similar to a function in program code: test objective, test activities, test steps, and, at the very bottom, the acceptance criteria. Threshold values can be specified directly, such as a confidence score of 0.5, at which point an object is classified. The model for this is the Test Description Language developed by the ETSI MTS committee, which has given rise to protocol tests for the mobile communications and automotive industries.
When is an AI system considered a high-risk system?
When it is part of a safety component. The key distinction here is between safety and security: safety refers to protecting people from the machine, while security refers to protecting the machine itself. Anyone operating such a system must comply with the standards covering the ten topics of the standardization mandate (including correctness, robustness, cybersecurity, quality management, transparency, and risk management) or go through a certification body.
What does the threshold of 10^25 FLOPS in the EU AI Act mean?
It distinguishes general-purpose AI systems from those posing systemic risk. If the hardware performance during training exceeds this value, more extensive documentation requirements and additional cybersecurity safeguards apply. The FLOPS are calculated based on an algorithm-specific constant, the token length of the training data, and the number of parameters. According to Taras Holoyad’s assessment, GPT-4 has exceeded this threshold.
What should manufacturers do if there is no applicable standard for their AI product?
They must approach a certification body, which evaluates the test results and issues the CE mark to grant market access. However, the body may only do so after it has been notified: A competent authority in the respective member state evaluates it together with an independent expert. The timeline called for designation by August 2025 and the development of expertise for high-risk systems by August 2026.


