Skip to main content

Search...

AI-Generated Test Cases: Compliant Without Tool Validation

AI-generated test cases for medical devices: a RAG system retrieves in-house documents, a human reviews the results, so no tool validation is needed.

• • Updated: • 11 min read
Cover of the expert talk on 'AI-Generated Test Cases: Compliant Without Tool Validation' with Alexander Frenzel and Richard Seidl.

AI-generated test cases can work in the regulated medical device industry if the system does not interpret data and only retrieves the company’s own documentation. A RAG system supplies text chunks through similarity analysis, and an LLM turns them into structured test cases. A human reviews the result and formally approves it. That creates traceability without the need for tool validation.

Key Takeaways

  • The HyDE principle solves the tagging problem for hundreds of thousands of document pages: instead of asking questions, the system formulates statements and uses similarity analysis to check which document chunks support them.
  • AI-generated test cases in a regulated environment need no tool validation as long as a human formally reviews every output and approves it by e-signature.
  • Automatic prompt engineering, where one AI subsystem writes optimized prompts for another, makes one-click test case generation possible without training testers to write prompts.
  • Modularity beats model choice: if the AI architecture lets you swap out individual models, you don’t depend on whichever model happens to be current.

Why AI-Generated Test Cases Look Different in Medical Technology

In a regulated medical device environment, AI-generated test cases have to meet one core condition: the AI must not add anything of its own. That is the principle behind the test case generator Alexander Frenzel built at Fresenius Medical Care, an assistant that supports testers without undermining regulatory traceability.

Fresenius develops hemodialysis machines like the ones in dialysis clinics. A patient’s blood circulation depends on these devices. Mistakes are not an option, and innovation has to bow to that reality. “Good is not good enough for us” is how Alexander describes the standard that shapes every technical decision.

Testing these systems has many layers. It covers not just software but hardware, electronics, materials, and biocompatible components, all interacting with one another. The proof of concept deliberately started with software, with the plan to roll it out more widely later.

What Sets the Test Case Generator Apart from a Chatbot

The generator is not a prompting tool for end users. It is a one-click approach. The tester enters a requirement, presses a button, and a test case comes out at the end.

That choice comes down to usability. In a large test organization, you would otherwise have to teach a lot of people to write good prompts. Instead, an AI subsystem handles prompt creation itself: from the input, it produces several prompts optimized for a downstream AI system.

Behind it sits not one AI but an architecture of several systems, some running different LLMs and models. That complexity is exactly why success depends on the question “How do I set something like this up?” and not on picking a single tool.

Why Traceability Has to Be Built In from the Start

In a regulated field, there is no room for hallucinations. That is why reasoning was a mandatory component for Fresenius even before the big AI wave, at a time when common chatbots didn’t offer it yet.

Testers need to know where a statement comes from. How does the system get from a piece of text to the boundary values of a boundary value analysis? How are the equivalence partitions derived? Those lines of reasoning have to be traceable. Otherwise the result is worthless from a regulatory point of view.

Fresenius didn’t bring AI expertise to the project. It brought domain knowledge. The AI expertise came from external partners. The company’s own contribution was the business concept: defining what the system has to look like to deliver real value.

How the HyDE Principle Makes Scattered Documentation Usable

The system works with the company’s own documents through a RAG system and doesn’t interpret them. Nothing new gets made up, and testers keep their original documentation.

The underlying problem: the data sits in scattered systems and repositories, some of it as Word files from 2002 that never made it into a current ALM system. With hundreds of thousands of pages, manual tagging is out of the question.

The solution came from one of the architects: the HyDE principle. Instead of asking a question, the system makes a statement and has the data in the RAG system back it up. An AI system recognizes similarity more easily than it produces direct answers.

Using similarity analysis, the system checks what percentage of a text section matches the statement at the vector level. That is how it finds the relevant chunks without the AI changing any content. There was another benefit, too: Fresenius wouldn’t have had enough of its own test data to train an LLM directly for a single project. The RAG approach removes that dependency.

A Human Stays in the Loop at Every Step

A human in the loop was a given from day one. Testers can step in at any point, adjust things, and add their own experience without having to start over.

That makes the output a draft, not a finished result. Testers decide for themselves when to take what was generated as their starting point. After that come testing on the machine, review, and approval, as before, as the prerequisite for formal test execution.

Development ran for about nine months in 2024, with a small internal team and a larger external one, around ten people on average, who did this alongside their regular work. A dedicated team could have done it in two to three months, and probably faster today.

Modularity Beats Picking the Perfect Model

The architecture is built so that models can be swapped out. That mattered more to Fresenius than committing to a particular model, because the models keep evolving.

The team tried a lot along the way: first Llama, then Claude, then Llama again. There were also technical hurdles around running everything in a private cloud, because intellectual property can’t be pushed to public services. The system has to run separately and stay under control.

When a new model comes out, it can be swapped in for the existing one and checked for better output. That flexibility keeps the solution viable across model generations.

Where AI-Supported Testing Hits Its Limits

The generator worked very well for software. It gets hard as soon as physical dependencies come into play.

One example: a tubing system being filled. The pressure isn’t static, it is highly dynamic. Effects like that are not easy to model. You need a physical model behind it, and that model would have to be updated for every new feature, every new motor, every replaced part. That effort puts tight limits on AI support in the physical domain.

Why These AI-Generated Test Cases Need No Tool Validation

A generative, non-deterministic system can hardly be validated in the classic sense. After a long discussion, the deciding question came up: does it even need to be? The answer was no, and for good reasons.

Several conditions support that assessment:

  • The output serves as a draft, not as a final result.
  • A human reviews it and formally approves it by e-signature.
  • The data sits unlabeled in the RAG system, and the system doesn’t change it.
  • Every process step is documented, and you can jump in and adjust at any time.

“Then it’s just a tool that helps you, not one that does the work for you. And that puts you back on solid regulatory ground.”

(Alexander Frenzel)

With good logging, what comes out stays traceable. Exactly how the system reaches a result can’t be reconstructed one hundred percent, and in terms of wording, sometimes one phrasing turns out better, sometimes another. The underlying information stays the same, though, because it isn’t altered.

What Testers Found Surprisingly Useful

In the internal beta test, the first reaction was excitement about the output. But testers saw the biggest value somewhere the team hadn’t expected: knowledge transfer.

Because the system shows, for every expected test step, where that expectation comes from, testers came across documents they hadn’t known about. With documentation spread across many places, that reveals where the relevant information actually lives.

The second big point was being able to step in throughout the whole process. Nobody has to start from scratch. A draft that isn’t perfect but is solid saves time that can go into fine-tuning before review and approval.

From Draft to Pipeline: The Obvious Next Steps

The next stages follow almost naturally from the existing system. First, the generator, which has been running as a web application, is to be integrated directly into the ALM and PLM systems.

The RAG database is to be connected to various document management systems and databases through service interfaces, so the data becomes easier to reach. From a well-documented test specification with traceability, the system can generate a keyword-driven test once you hand it the company’s own keyword library.

From there, it is a short step to a test script for a software-in-the-loop system. The script could go straight into the pipeline, be tested in simulation, and be revised by the system itself if it fails, for example with an LLM-as-a-judge approach.

One thing stays the same at every stage: the human in the loop for approval and traceability. People need a solid understanding of the system and of test design to judge whether the result really tests what it is supposed to test. Syllabi such as the new GenAI syllabus from the GTB and ISTQB help make that judgment on solid ground.

Frequently Asked Questions

Why is generative AI considered a sensitive issue in testing within regulated industries?

Because a generative system must not invent anything if the results are to remain usable for regulatory purposes. Testers must be able to trace the origin of a statement: How does the system go from a text snippet to the boundary values of a boundary analysis, and how are the equivalence partitions generated? Without traceable lines of reasoning, the result is worthless in a regulatory sense. That’s why reasoning has been a mandatory component from the very beginning.

Do testers need to learn how to create prompts in order to use AI for test cases?

No. The generator described here works as a one-click solution: The tester enters a requirement, presses a button, and a test case is generated. Prompt generation is handled by an upstream AI subsystem that generates multiple prompts optimized for the downstream system from the input. In a large testing organization, this eliminates the need to train many people in prompt creation.

How can hundreds of thousands of pages of distributed documentation be made usable without manual tagging?

Through the HyDE principle. Instead of asking a question, the system formulates a claim and has it substantiated using the data in the RAG system. A similarity analysis checks the percentage to which a text segment matches the assertion based on its vectors. An AI system recognizes similarities more easily than direct answers, and the content remains unchanged.

What are the advantages of a RAG approach over training your own model?

Two reasons. First, the original documentation is preserved because the system only retrieves the chunks and does not interpret them. Second, there simply isn’t enough data for a single project: Training an LLM directly would have required far more in-house test data. The RAG approach eliminates this dependency and even works with Word files from as far back as 2002.

Should you commit to a specific model when designing an AI architecture for test case generation?

No; replaceability is more important. Because models are constantly evolving, the architecture was designed so that a new model can be swapped in for the existing one and tested for higher performance. Over the course of the project, the team switched from Llama to Claude and back again. The system ran in a private cloud because intellectual property does not belong in public services.

Does an AI tool for test case generation need to be validated in a regulated environment?

Not necessarily. A generative, non-deterministic system is difficult to validate using traditional methods, and with this setup, it wasn’t necessary either: The output is a draft; a human reviews it and formally approves it via e-signature; the data remains unlabeled in the RAG system and is not altered; and documentation is generated for every process step. The tool supports users; it does not do the work for them.

Can tests for hardware and physical effects also be generated?

Only to a limited extent. The generator worked very well for software, but it becomes difficult when dealing with physical dependencies. Take a hose system as an example: During filling, the pressure is not static but highly dynamic. This would require a physical model that would need to be updated for every new feature, every new motor, and every replaced component. This effort imposes strict limitations.

What benefits do testers report from using such systems in practice?

The testers saw the greatest added value during the internal beta testing in knowledge transfer, something the team hadn’t anticipated. Because the system shows, for every expected test step, where the expectations originate, they came across documents they hadn’t been aware of. The second benefit was the ability to intervene throughout the entire process: No one starts from scratch.

Share this page