AI-supported test design means using language models such as Copilot or ChatGPT to generate equivalence partitions, boundary values and test cases in a structured way. The tools produce usable results within minutes, but they also get things wrong. To judge what they deliver, testers need a working command of the classic test design techniques themselves.
Key Takeaways
- Test design techniques such as equivalence partitioning and boundary value analysis are rarely used in practice, even though almost every tester learned them in a certification course.
- AI tools return a test design in 30 to 40 seconds, and after two or three minutes of follow-up prompting the quality reaches a level Michael Fischlein has rarely seen in client projects.
- Anyone who wants to assess AI-generated test designs needs method knowledge of their own, or they will not recognize the errors and gaps in the output.
- Boundary value analysis offers the widest practical benefit, because boundary errors occur systematically and even a vague word like “between” leads to wrong test derivations.
- In a suite of tens of thousands of test cases, equivalence partitioning showed that more than half were redundant, while whole value ranges had never been tested at all.
Most Testers Design Tests by Gut Feeling
Plenty of experienced testers never use formal methods, even though they once learned them. Their test data, requirements and test cases come from experience and instinct rather than from equivalence partitioning or boundary value analysis. That gap is exactly where AI-supported test design comes in, and also where it runs into its limits.
You can see it in every Foundation Level class. People who have been testing for ten or twenty years and are only now getting certified end up relearning the basics. They may have grasped the idea of an equivalence partition intuitively, but they have never heard the term. Boundary values ring a faint bell. The more advanced techniques that follow are largely unknown.
That is not automatically a problem. A team that is happy with the quality of what it ships and delivers on time is doing a lot right. Often this works because one person on the team has a real feel for the software, a kind of natural intelligence that finds bugs without any formal technique.
Why Test Design Techniques Rarely Make It into Practice
The main reason is lack of practice. In a four-day Foundation Level course, only half a day to a full day goes into practicing test design techniques. That is not enough to make them stick.
Real projects are also messier than the textbook examples. A coffee machine, a soda vending machine or an ATM is easy to grasp. A decision table with five conditions has 2 to the power of 5 rows, which is still manageable. With twenty conditions it turns into a huge table, and many people back away from it.
“Then you’re staring at a blank sheet of paper. And we all know the fear of the blank page.”
(Michael Fischlein)
Agile makes things worse when it is misunderstood. “We document less” turns into “we don’t write anything down.” Testers think they have to be fast and skip test design. Under time pressure, few are willing to invest in practicing a method first.
On top of that, the organization often does not ask for it. Test managers with an Advanced Level certificate have the paperwork, but not necessarily the knowledge to demand test design from the team and model it themselves.
Exploratory Testing Needs Methods Instead of Replacing Them
The usual way out is exploratory testing: checking things on the side, without formal design. Yet good exploratory testing only works if you have the techniques down.
When you explore, you should be applying equivalence partitions, boundary values and decision tables implicitly. That knowledge is what guides you as you move through the software: which price category, which value ranges, which combinations matter. The methods do not go away. They run in the background.
Simple Techniques Deliver the Biggest Payoff
Boundary values and equivalence partitions are the two techniques that pay off almost every time. They are simple and cover most of the typical defects.
Boundary values go wrong again and again because requirements are written in prose. A sentence like “the shortest board is between one meter and five meters” is ambiguous. Does “between” exclude the limits, so the valid range runs from 1.01 to 4.99 meters? That is clearly not what the author meant, but this is exactly where defects creep in.
Equivalence partitions cut the number of test cases dramatically. One client had tens of thousands of test cases. After the analysis, more than half could be deleted because they were duplicates from an equivalence partitioning point of view: different values, same partition. At the same time, entire value ranges were not covered at all.
A simple question from job interviews shows how thin the basic understanding often is. If you are asked to test a closed interval from 9 to 99, you should pick one value below it, one inside and one above. Instead, some candidates wrote one, two, three, four and so on up to 100 on the flip chart without noticing how pointless that was.
State transition testing also works well day to day, as long as the system has states. A ticketing system is easy to visualize and understand. What matters is that none of this is rocket science. These are simple methods you can apply hands-on, even under agile time pressure.
AI Closes Half of the Method Gap
Large language models produce usable test designs in seconds, even though nobody trained them for testing. Tools like ChatGPT or Copilot process language statistically and generate test cases, acceptance criteria, boundary values and equivalence partitions on request.
How accurate are they? Not flawless. Hallucinations and nonsense show up in the output. As a rough estimate, about 80 percent of the results are correct. Give the same task to experts and they land in a similar range of 80 to 90 percent, but they need one to two days. The AI takes 30 to 40 seconds, and with a bit of prompt tuning you have a draft design in two or three minutes.
So AI bridges the gap between “learned the method” and “never used the method” in a very practical way. A second, harder gap remains, though: someone has to evaluate the result. If you have not mastered the techniques, you will not spot the mistakes in the AI output.
That is where the real training challenge lies. An experienced trainer sees the flaw in a complex test design right away. A junior consultant fresh out of Foundation Level does not. Building that ability to judge the output is the task ahead.
Getting Started with AI-Supported Test Design
Talk to these tools the way you would talk to a person. They are chat systems, and a “please” in the prompt does no harm. Say clearly what you want, for example: please write out the equivalence partitions for this state.
A workable sequence starts with the business context and moves toward the method:
- Set the context: “I am an insurance company and want to create new customers. Which variables and fields do I have, and what are their attributes?”
- Refine the variables: add fields, remove fields, adjust until the list is right.
- Apply the method: “Build the equivalence partitions from this.”
- Sharpen: drill into individual areas and have missing points added.
This also works in pairs. Bring the AI to the table as a third participant and push back on its answers together. In training, a mistake in the output becomes a learning moment: if you can tell that an AI answer is wrong, you have understood the method.
Be careful what you paste in. Customer data and sensitive information do not belong in a chat tool without a check. A requirement for a standard online shop system with no names in it, on the other hand, can be copied in to generate equivalence partitions and reveal the first gaps.
It is striking where these tools get their domain knowledge. For a requirement from the automotive sector covering reversing and safety systems, Copilot produced usable requirements complete with source links. The sources were the manufacturer’s own public web pages. A week of planned work was almost done, apart from the details that were still missing.
Playing Around Is the Fastest Way to Learn
The best way in is to try it, not to read about it. By playing with the tools you learn what prompt engineering means in practice and how a system responds to roles, facets and wording.
You only get a feel for the effect of a phrasing by doing it. Ask the system to act in a specific role or add another facet, and the answer changes. Sometimes you get nonsense, sometimes exactly the hit you were looking for. That instinct only develops when you poke at the system yourself.
The same principle carries over to professional work. If you build a story or a scenario with an AI, you practice the same mechanics you will later use when writing test cases and requirements.
Know-How Remains the Bottleneck
AI tools will take over the simple, repetitive parts of test design. What stays is the need to frame a problem correctly and to check the result with common sense.
That shifts the weight back toward requirements engineering. You have to make clear to the system what you want and then review the result with real expertise. Both depend on know-how you have built up over time.
A comparison shows what is still open. Hardly anyone today knows how a compiler works, yet everyone trusts it to do the right thing. AI-supported testing may well go the same way. The unanswered question is how testers will build the know-how they need to evaluate the results at all.
Frequently Asked Questions
Why do experienced testers so rarely apply test design methods in their day-to-day work?
The main reason is a lack of practice. In the four-day Foundation-level course, only half a day to a full day is devoted to practicing the techniques, and that’s not enough to truly internalize them. On top of that, real-world projects are more complex than a coffee maker or an ATM: A decision table with five conditions is manageable, but many people shy away from one with twenty conditions. Often, there’s also no organizational mandate to require test design.
Is it a problem if a team derives test cases based on experience and gut feeling?
Not necessarily. If a team is satisfied with the quality of its deliverables and meets deadlines, it’s doing a lot right. This is often made possible by one person on the team with a keen intuition for the software, who can find bugs without relying on formal techniques. This approach only becomes costly when no one is left who can assess which value ranges aren’t covered at all.
Can exploratory testing replace formal test design methods?
No. Good exploratory testing requires a solid grasp of these methods. Anyone who explores the software implicitly applies equivalence partitions, boundary values, and decision tables: which price categories, which value ranges, and which combinations are relevant. The methods don’t disappear; they work in the background. Exploratory testing is therefore not a workaround for a lack of methodological knowledge.
Why do boundary values so often go wrong in practice?
Because requirements are formulated in text, and language is ambiguous. A sentence like “the shortest board is between one meter and five meters” leaves open whether the limits are included or whether it actually means 1.01 meters and 4.99 meters. What is usually meant is the opposite of what is written there, and this is precisely where incorrect test derivations arise.
How much testing effort can be saved with equivalence partitions?
For a client with tens of thousands of test cases, the analysis revealed that more than half could be eliminated: different values, same class, in other words, duplicates that provided no additional insight. At the same time, the analysis revealed that entire value ranges were not covered at all. So the method not only reduces the volume but also uncovers blind spots.
How reliable are test cases generated by a language model?
Roughly 80 percent of the results are correct, though they include hallucinations and nonsense. Experts achieve 80 to 90 percent reliability on the same task, but it takes them one to two days. The language model delivers results in 30 to 40 seconds; with follow-up prompts, a draft is ready in two to three minutes. Evaluating the output remains a human task.
Is it permissible to copy requirements from customer projects into a chat system?
Customer data and sensitive information should not be entered into a chat system without being reviewed first. A requirement for a common e-commerce platform (one that does not contain any names) can, however, be copied and pasted to generate equivalence partitions and identify initial gaps. The line, therefore, is not drawn at the tool itself, but rather at the content of the documents entered.
What skills must testers have if AI takes over simple test design work?
Two things: the ability to formulate a problem precisely and to review the results from a technical perspective. Both require established expertise, shifting the focus back toward requirements engineering. An experienced trainer can spot an error in a complex test design immediately, whereas a junior tester who has just completed the Foundation Level cannot. How this ability to evaluate will develop in the future remains an open question.


