Skip to main content

Search...

Testing AI Systems: Determinism Decides the Test Strategy

Testing AI systems often boils down to one question: does the system behave deterministically? The answer decides whether classic methods suffice.

• • Updated: • 11 min read
Cover of the expert talk on 'Testing AI Systems: Determinism Decides the Test Strategy' with Marco Achtziger, Gregor Endler and Richard Seidl.

Testing AI systems means testing an AI component as part of the overall system, not testing the model in isolation. The deciding question is determinism: if the system always responds the same way to the same input, classic test methods apply unchanged. If it doesn’t, you need statistical metrics and test datasets that testers and data scientists agree on together.

Key Takeaways

  • Whether an AI system is deterministic or not determines the entire test strategy. If it behaves deterministically, an AI component changes almost nothing about how you test.
  • Testers and data scientists overlap when they build test data, so close contact between the two roles can directly improve the quality of the training data.
  • Basic statistics is no longer optional for testers: AI systems deliver functional outputs only with a certain probability, and a plain pass/fail verdict is not enough.
  • Model updates need their own, smaller test datasets so that a retrained system doesn’t drift off in an unwanted direction.

Testing AI Systems Starts with One Question: Is It Deterministic?

Before you think about statistical methods, model internals or new test tools, settle one thing when testing AI systems: does your system behave deterministically? The answer shapes your whole test strategy, and it is the first item on any AI testing checklist.

If the system is deterministic, very little has changed for you as a tester. You have functional requirements and a subcomponent that does its job. Whether AI, machine learning or linear regression sits behind it hardly matters. The same input has to produce the same output every time, so you can define fixed input-output pairs and test against them.

Sometimes a system is only partly deterministic. Maybe the output for a given input changes only when the model is retrained. If that happens twice a year, you can work with fixed test data and simply check, well before the next model release, whether your assumptions still hold.

Things only get harder when the answer is clearly “not deterministic”. Then the question becomes how often you have to run a test case to be statistically confident that your assumptions are met. At its core, though, that means running familiar tests more often. It is not a whole new game.

Nondeterminism Can Appear Even When the Model Doesn’t Change

A fixed model does not guarantee deterministic behavior. The same model can produce different results on different hardware architectures, even though nothing in the model itself has changed.

In that case you have to look more closely. Where does the system run? Is the hardware always the same? Which algorithms does the GPU use for matrix operations? This is where testing gets uncomfortable, because race conditions between processing units can creep in.

Often you can pull the system back to a deterministic case. If it runs on-premises on fixed hardware, you may not need GPU acceleration at all and can use a single-core CPU with no race conditions. Reproducibility is back. That is why it pays to find out early where, and on which hardware, the model will actually run.

Testers and Data Scientists Work on the Same Thing from Two Directions

The biggest lever is the conversation between testers and data scientists. Their activities overlap, but each role sees different gaps.

A data scientist gets a problem and builds a model that solves exactly that problem. They evaluate it with mathematical methods and conclude that the model works. What can get lost along the way is that the model is only one part of a larger system with a specific job to do. Tunnel vision sets in quickly: model built, model evaluated, done.

This is where the tester comes in. Testers look at the whole system and check whether the assumptions made during training also fit the real application. They can point out boundary conditions in the data that nobody considered during training.

Those boundary conditions can leave a data scientist thrilled or heartbroken. Thrilled, because new ideas lead to a better training dataset. Heartbroken, when it turns out the existing training dataset was useless. Both outcomes are valuable, because both push quality forward.

Which Statistics Skills Testers Need for AI

Statistics is becoming more important. If a model delivers the expected output only a certain percentage of the time, you need to understand what that means for your test. “Stats 101” is often enough to know where to ask questions.

If a model claims 99.9 percent accuracy, the question that matters is: what exactly does that mean, and is it good enough for the intended use? You don’t have to invent terms like accuracy and precision, but you do have to be able to interpret them.

These methods help in practice:

  • Statistical and stochastic methods to evaluate probabilistic outputs
  • Combinatorial testing, which is gaining weight in AI projects
  • Exploratory testing, which is becoming more important, not less

You don’t need deep knowledge of what a data scientist does. These are two separate roles with two skill sets. If you work at a small company and cover both roles, learning both pays off. Otherwise it is enough to understand what the models output, because that output is what shapes your test design and test strategy.

The difference between machine learning, deep learning and other approaches is nice to know. It is not a must-have skill for testers, even if many want to learn it because the topic is hot right now.

How Much of the AI Do You Need to Understand?

You can treat most of it as a black box. Fully understanding what happens inside a model is an open research question. Explainable AI exists for a reason, and it is far from solved.

So you won’t always be able to understand everything. You can, however, understand it better than you did at the start. Ask yourself: which parts of my system use AI at all, and where do I need to look more closely?

How readable the model is determines how far you can get by hand. A decision tree you can still walk through, depending on how wide it is. A deep neural network whose formulas would fill hundreds of pages is hard for anyone.

A small slice of the basics is still useful. It gives you a feel for where to ask questions and helps you spot your own blind spots.

Do We Still Need Testers? Yes

Yes, we need testers. The question came up in the working group, and some tool vendors may well claim otherwise. The answer stands: we need people who approach a system with a testing focus.

That focus is not an add-on. A data scientist evaluates their model with the metrics that seem sensible to them. The tester checks whether it holds up in the overall system and in the real use case. Neither perspective replaces the other.

Test Data: Do You Need a Golden Dataset?

When it comes to test data, it is worth talking to the data scientist again. They have already produced training data, and you need to work out what you can reuse and what you still need.

Ask yourself whether you really have to test your system against the same dataset every time. Maybe that isn’t a use case at all, maybe it doesn’t matter. Maybe your tests also have to cover the model’s training. In that case, agree on what the golden dataset means for testers and what it means for the data scientist.

Data scientists already create test data as part of their work, because they need data to train the model. That produces datasets you can use for testing directly, or at least gives you a feel for the data that helps when you generate your own test data later.

Model Updates Need Their Own Tests

A point that often gets forgotten: if the system keeps learning, you need smaller datasets that simulate updates. They show whether the model is developing in the right direction and still matches reality.

Chatbots that were released into the wild and turned into rather interesting personalities within a short time show what happens when this case is ignored. How does a self-learning model behave under update conditions? If nobody tests that, exactly this kind of thing can go wrong.

There is no silver bullet that tells you whether you need this. But ask the question. And remember that it is not only about the initial test setup, but also about smaller update test datasets for ongoing model changes.

A Lot of It Is Old Wine in New Bottles

Nondeterministic systems existed long before AI. “Eventually consistent” systems in microservice architectures have forced testers for years to shape test cases so they can cope with nondeterministic behavior.

Statistics in testing isn’t new either. A good non-functional requirement for time behavior reads something like this: in 95 percent of cases, the web page is delivered in under 300 milliseconds. Anyone who takes performance testing seriously has been dealing with statistics for a long time.

“Just because you now have an AI system, you hopefully still have requirements. Look at your requirements and go back to your normal test methods.”

(Marco Achtziger)

What changes is the output of the AI component. The interface no longer simply returns B for A, but B with a probability of x percent, along with metrics such as accuracy and precision. The rest of your toolbox comes with you.

Classification Problems Can Be Tested with Familiar Methods

Many AI tasks are classification problems, and those you can cover with established test methods. You feed in an image and want to know whether it is 90 percent cat or dog.

Apply the methods you already know. What are the functional tests? Are there boundary values to consider? Can you identify equivalence partitions? These questions lead you straight to a clear test design.

The more common blocker is not missing technique but shock and fear of the buzzword. People hear the big word “AI”, think of human-like intelligence and believe they need a psychologist rather than a tester. That giant has to be broken down into something tangible: your requirements, your components, your familiar methods and the one central switch between determinism and nondeterminism.

Frequently Asked Questions

Do AI components inevitably change the test strategy?

No. The key factor is whether the system always behaves the same way given the same input. With deterministic behavior, almost everything stays the same: There are functional requirements and a subcomponent that does its job, so fixed input-output pairs can be defined. Whether machine learning or linear regression is behind it is then of secondary importance.

Can an unchanged model still produce different results?

Yes. The same model can produce different results on different computer architectures without any changes being made to the model. Among other things, this is caused by the algorithms used for matrix operations on the GPU, where race conditions can creep in between processing units. If the system runs on-premises on fixed hardware with a single-core CPU, reproducibility is often restored.

What skills should testers acquire for AI systems?

A basic understanding of statistics is paramount, because models only deliver the expected output with a certain probability. “Stats 101” is often sufficient to interpret metrics such as accuracy and precision and to ask targeted questions about them. Combinatorial and exploratory testing are gaining in importance. The difference between machine learning and deep learning is nice to know, but not a required skill.

What are the benefits of close collaboration between testers and data scientists?

Both roles have overlapping activities but identify different gaps. A data scientist evaluates their model mathematically and considers it complete. The tester brings a view of the overall system and identifies boundary conditions of the data that weren’t considered during training. This either leads to a better training dataset or to the realization that the previous one was unusable.

Does a tester need to understand what’s happening inside a model?

You can treat most of it as a black box. Fully understanding what happens inside a model is an open area of research; otherwise, there would be no need for Explainable AI. How deeply you can delve into it manually depends on its readability: You can walk through a decision tree depending on its breadth, but you can’t do the same with deep neural networks that involve pages of formulas.

Where do testers get test data for AI systems?

Some comes from the data scientist, who generates data for model training anyway. Determine what has reusability and what additional data is needed. Even if the datasets don’t fit perfectly, they provide a sense of how to generate your own test data later on. Whether testing must always be done on the same dataset is a separate question.

Why do model updates need their own test datasets?

Because a system that is retrained through continuous learning can evolve in an undesirable direction. Smaller datasets simulate such updates and show whether the evolution aligns with reality. Chatbots that were released and quickly developed into interesting personalities are a well-known example of this overlooked scenario.

Are non-deterministic systems and statistics in testing really an AI topic?

No, both existed before AI. “Eventually consistent” systems from the microservices environment have long required test cases to be designed to handle non-deterministic behavior. Statistics are also nothing new: a non-functional requirement such as “under 300 milliseconds in 95 percent of cases” is part of any serious performance test. What’s new, above all, is the output of the AI component.

Share this page