Skip to main content

Search...

Systematic Testing: Exit Criteria Indicate When It’s Enough

Systematic testing picks risky and everyday inputs, and exit criteria show when you have tested enough. Complete testing is mathematically impossible.

• • Updated: • 12 min read
Cover of the expert talk on 'Systematic Testing: Exit Criteria Indicate When It’s Enough' with Andreas Spillner and Richard Seidl.

Systematic software testing means picking the relevant inputs out of a practically infinite number of possibilities: the ones with high risk and high everyday importance. Testing doesn’t break software. It shows where the software is already broken. Sufficient test coverage comes from techniques such as equivalence partitioning and boundary value analysis, not from blind trial and error.

Key Takeaways

  • Testing is not destructive: no test case ever breaks software. It makes visible where the software is already broken.
  • Exhaustive testing is mathematically impossible. Three 16-bit inputs at 100,000 tests per second would take around 89 years, and with 32 bits you are looking at 10 to the power of 15 years.
  • A Foundation Level certificate is a driver’s license, not proof that you can drive: it lets you take part in traffic, but real testing skill only comes from years of practice.
  • Agile has only partly torn down the wall between development and testing, and in many people’s heads the wall is still standing.
  • Developers who meet clear coverage criteria with equivalence partitions and boundary values know when they have tested enough. Anyone who runs 300 automated test cases without a criterion doesn’t.

Testing Doesn’t Break Software, It Reveals Defects

Testing shows where software is already defective rather than destroying anything, and clear exit criteria tell you when you have looked closely enough. The distinction between breaking and revealing may sound like hair-splitting, but it shapes how an entire discipline sees itself.

No test case damages a program. If a piece of code is faulty, it was faulty before the test ran. The test only brings it to light. Calling testing a destructive activity turns the logic upside down.

Testing is constructive because it takes real thinking. Testers work out which test data and which test environment will expose a problem the developers didn’t think of. They have to think further: was the negative case covered? Did anyone think of February 29? Gaps like these persist decades after they were first described.

“Testing doesn’t break anything. It shows where the software is broken. Not a single test case breaks any software.”

(Andreas Spillner)

Why Complete Testing Is Impossible

No program in the world will ever be run through every possible path. That is a question of mathematics, not discipline.

A quick calculation makes it tangible. Take an old 16-bit system with three numbers to enter and a test automation setup that runs 100,000 tests per second. Trying every combination would take around 89 years. With 32 bits, you end up in the range of 10 to the power of 15 years.

That doesn’t make testing pointless. It means you have to choose. Which workflows run every day? Where does a failure carry a high risk? That is where systematic testing starts, instead of reaching in blindly somewhere.

Once this one number has sunk in, you’ve gained a lot as a tester: no program is ever exercised in every way it could be. So the question isn’t whether you select, but how.

Sufficient, Not Exhaustive

The idea of testing something “completely” is misleading. To claim something is fully tested, you would have to have run every possibility once. As the bit example shows, that never happens.

A better goal is a set of test cases that is sufficient for the problem at hand. Not extensive, not exhaustive, but appropriate to the risk. A highly critical system calls for many techniques and great test depth. A vacation planner for employees doesn’t.

Test design techniques help with that choice. Equivalence partitioning lets you boil many input values down to a few representatives. That gives you a solid set of test cases without drowning in combinatorics.

A spider’s web captures the idea. A good web is evenly woven so that things get caught. It isn’t tight in one place and full of holes in another. That kind of even coverage is exactly what test design techniques provide.

What Exit Criteria in Software Testing Do

Exit criteria decide whether you can head home at the end of the day feeling good about your work. If you know the equivalence partitions are covered, the boundary values are checked and the decision tables have a green tick, you have a reliable signal that coverage is sufficient.

The opposite is common. An automated suite runs through 300 test cases, and nobody knows exactly what those 300 test cases actually check. Then the criterion that tells you “you have tested enough” is missing.

A structured technique gives developers exactly what they otherwise lack: a point at which they’re allowed to stop. Without that criterion, testing feels like a bottomless task that’s hardly worth starting.

Agile Has Partly Torn Down the Wall Between Testing and Development

The biggest shift in testing over the past two decades has been agile. The Agile Manifesto dates from 2001, long before the topic found its way into syllabi and standard textbooks in a big way.

There used to be a wall between development and testing. Developers threw the product over the wall, testers threw the bug report back, and nobody talked. Agile has broken up that separation, because it puts the team at the center instead of fixed handovers.

The demand itself is older than agile. People had long argued that developers and testers should work with each other, not against each other. What’s new is a method that turns this collaboration into a principle.

One wall still stands, though: the wall in people’s heads. You see it wherever collaboration would be possible but old reflexes keep kicking in.

Communication Is Productive, Even If Management Doesn’t See It

Talking to each other is the core of agile work, and that’s exactly what companies often don’t recognize as work. There’s frequently no time set aside for exchange and cooperation, because nothing seems to come out of it.

In reality, it’s the other way around. Communication often produces the most, you just don’t see the result on paper right away. That gap between impact and visibility is the real problem when you try to sell it.

One practical approach is to disguise the space for communication. A workshop or a short coaching session creates a formal frame in which exchange is allowed to happen. Once the first effects show, tolerance for it grows.

Testers and developers both gain from this. The developer picks up testing know-how, the tester picks up development knowledge. An agile team today needs both, for test automation, for example.

There Is No Favorite Technique, Only Suitable Ones

No universal test technique exists. Which technique, or which combination, is right always depends on the system and its environment.

There is no silver bullet, least of all in testing. The rule instead: the more critical the system, the more intensive and varied the techniques. The lower the risk, the less effort is justified.

That mindset takes pressure off. If you stop looking for the one right technique and look for the one that fits the problem in front of you, you make better decisions. Choosing a technique is a question of context, not fashion.

Error Culture Decides Whether Defects Lead to Learning

A good error culture means owning your mistakes and learning from them. In Germany, that culture is often weak, and not only in software development.

One image makes the contrast clear. When a rocket explodes at launch and everyone cheers because they learned so much from the failure, that’s an attitude that is hard to imagine in Germany. People are more likely to duck, hide and sweep things under the rug.

Celebrating isn’t the goal. Learning is. Some companies live a culture where mistakes are discussed openly so they don’t happen again. Others duck. That difference decides whether a defect becomes a reason to learn or a taboo.

Personal openness is part of it. Working through your own mistake in public instead of slinking away mostly helps you, even if nobody else can learn from it directly.

The Foundation Level Certificate Is a Driver’s License, Not a Master’s Degree

A testing certificate at Foundation Level is like a driver’s license. From now on you’re allowed on the road, you know the rules and the signs, but that doesn’t mean you can really drive yet.

The knowledge behind it isn’t new. Many of the techniques were described back in the 1970s and are still valid today. The merit of a good foundational textbook lies not in innovation but in clear terminology and a solid presentation.

Real competence grows afterward. Only after years of working with real systems do you learn to test, just as it takes years before you really know how to drive. The certificate is the entry ticket, not the destination.

A syllabus as a guide also makes writing about the subject much easier. When the content is set, the hardest part falls away: the constant question of what belongs in and what doesn’t. Embedded topics and real-time systems, for instance, are deliberately left out of the foundation level.

Why Developers Struggle to Get Excited About Systematic Testing

Convincing developers to test systematically is still hard. The idea of exhaustive testing runs deep: testing looks like an endless task that is hardly worth starting.

This is exactly where the lever is. Structured techniques hand developers exit criteria. They turn a task that seems limitless into one with a clear finish line.

You can see the difference in two ways of ending the workday. With equivalence partitions covered, boundary values checked and decision tables complete, you go home satisfied. With 300 test cases and no idea what they’re for, you’re left with the uneasy feeling that you have no criterion for having tested enough.

In the end, it takes passion. Whoever approaches a topic with conviction produces solid results, in almost any field.

Predictions About the Future of Testing Rarely Hold

Predictions about technology often miss, and testing is no exception. Expectations such as heavy modularization, where systems would be assembled only from thoroughly tested building blocks with verified interfaces, haven’t come true in that form.

Hype waves reinforce the pattern. A few years ago, blockchain was supposed to make contracts unnecessary. That hasn’t happened on the promised scale. Technologies like that often settle on a plateau of real use, far from the original promise.

AI is under similar scrutiny right now. Whether it’s a longer hype or a lasting change, nobody can say with confidence. Only one thing is certain: the expectation that AI will make testing unnecessary because you just tell it what to do repeats an old pattern. People used to say formal specification and code generation would make testing obsolete. It didn’t turn out that way.

Your own misjudgments teach humility, too. Watching TV on a mobile phone was once dismissed as pointless, and today almost everyone does it. Looking into the future remains hard.

Frequently Asked Questions

Why isn’t it enough to simply test a program by feeding it a very large number of inputs?

Exhaustive testing fails due to combinatorial complexity. An old 16-bit system with three numbers to enter and an automation tool capable of running 100,000 tests per second would be busy for about 89 years; for a 32-bit system, the figure is on the order of 10 to the 15th power years. That’s why we have to be selective: processes that run every day, and areas where a failure poses a high risk.

What distinguishes “sufficiently tested” from “thoroughly tested”?

“Thoroughly tested” would mean having executed every possible scenario at least once, and that is never achievable. A sensible approach is a test case set that matches the risk: A highly critical system requires many methods and deep test coverage, whereas an employee vacation scheduling system does not. The equivalence class method condenses many input parameters into a few representative values. The goal is uniform coverage, comparable to a well-formed spider web.

How do you know when you can stop testing?

The signal comes from the exit criteria of the methods: covered equivalence partitions, verified boundary values, and fully processed decision tables. Once you’ve achieved this, you can end the test with a clear conscience. If, on the other hand, 300 automated test cases run without anyone being able to say what these test cases are actually checking, this very criterion is missing, and that uneasy feeling remains.

Has agility truly eliminated the separation between development and testing?

Only partially. The old handoff (where developers would “throw the product over the wall,” testers would send the error report back, and no one would talk) has been broken down. The call to work with each other rather than against each other predates agility; what’s new is that a methodology has made this collaboration a guiding principle. What remains is the wall in our minds.

Is there a testing method that works for all systems?

No, there’s no “silver bullet” in testing. Which method (or combination of methods) is appropriate depends on the system and its environment: The more critical the system, the more intensive and varied the testing; the lower the risk, the less effort is warranted. Those who look for the right method for the situation, rather than “the one right method,” make better decisions.

What does a Foundation-level testing certificate say about a person’s skills?

It’s like a driver’s license: You know the rules and traffic signs and are allowed to participate, but that doesn’t mean you can actually drive yet. True competence only develops over the years through working with real systems. Many of the methods taught were already described in the 1970s and remain valid today. Embedded aspects and real-time topics are deliberately left out at the foundational level.

Will artificial intelligence make testing obsolete?

This expectation follows an old pattern. It was once claimed that formal specification and code generation would make testing unnecessary, and that didn’t happen. Likewise, the idea that systems would soon consist solely of intensively tested building blocks with verified interfaces has not materialized. Whether AI is just a passing fad or a lasting change cannot be said with any certainty.

Share this page