Human in the loop testing starts from natural intelligence: the human ability to find bugs through curiosity, intuition, and context that no AI system would uncover without being pointed at them. Exploratory testing, checking, and a newly proposed third category called digging together make up a complete test strategy. Human judgment remains necessary to assess AI-generated test results on their merits.
Key Takeaways
- Exploratory testing finds bugs that come from human curiosity, intuition, and context, bugs that systematic test case generation alone does not produce.
- AI-generated unit tests often land in the same equivalence partition and miss the input cases that really matter, so accepting them unchecked is a quality risk.
- Besides testing and checking, AI activity needs a third category: digging through training data without actually understanding the task.
- Junior testers or developers who never assess AI output themselves never build the expertise to judge that output at all.
- The deciding question with AI is not what is technically possible but whether it makes sense to hand an AI this particular task.
A Bug No AI Would Have Gone Looking For
Some bugs only surface because a person is curious, and that is the case for human in the loop testing. On his first project, Jonas Poller was exploring a piece of software with complicated business logic. He changed several parameters and clicked back and forth between two states. At some point a price went up by one cent for no reason. The bug was reproducible and turned out to be a serious problem.
Christian Brandes later called this move the “gross/net flick-flack”: two toggles that someone keeps switching back and forth until something tips over. You don’t get an idea like that from a script. You get it from wanting to understand someone else’s software.
A second example shows the same pattern. While testing input validation, Jonas spent half an hour watching the software catch everything he tried. Only when the cursor sat at the far left, blinking, did Ctrl+V suddenly work. The clipboard happened to hold a decimal number. When he pasted it, the whole browser crashed. A chain of coincidences that nobody would have written down as a test case in advance.
Why Exploratory Testing Remains a Human Task
Exploratory testing runs on curiosity, intuition, and context, and no model currently brings those on its own. That foundation is what separates a planned test case from an observation nobody asked for.
An example from years of exploratory testing training makes this concrete. The test object was a learning laptop for preschoolers. Across all those sessions, exactly one participant asked whether the device could be used without a mouse, because his child held the laptop on their lap in the car with nowhere to put a mouse. That one person, with that one context, found a lot of bugs.
An AI could well produce many of these test cases as candidates. But you would have to steer it there first. You would have to prompt it: think about other environments, other usage situations. The impulse does not come from the model.
Is AI Creative, or Does It Only Simulate Creativity?
In the view of both guests, current models are not creative. They only look that way. Asking whether an AI can bring real creativity to test design quickly leads to two deeper questions: Is the AI actually creative, or is it simulating? And what is creativity in the first place? That second question is in the same league as asking what intelligence is.
Test design doesn’t need a final answer to that philosophical question. Even if a model delivers something that feels creative, there is no reason to give up human curiosity and intuition. Exploratory testing belongs to experience-based testing, and experience cannot simply be rationalized away.
One point stands regardless of the creativity debate: an AI cannot reliably tell you what is exactly right. In many cases a person knows for certain whether an implementation or an output is correct. With AI it stays an estimate. That estimate can be very reliable, but in critical areas the question is whether you want to depend on it.
Testing, Checking, Digging: A Third Category
What an AI does in test design fits neither the testing box nor the checking box. The distinction between testing and checking is surprisingly little known among testers. When the audience was asked, barely more than five hands went up.
Christian and Jonas propose a third term for what the AI does: digging. As things stand, the model does not understand what it is doing. It has training data and uses probabilities to carry learned information over to a task, hoping for a hit. Side by side, the three terms separate cleanly:
| Term | Basis | Character |
|---|---|---|
| Testing | Intuition, curiosity, experience | Human, exploratory |
| Checking | Script | Mechanical, automatable, repeatable |
| Digging | Data, training, probabilities | Neither one nor the other |
Blindly rummaging through test ideas from other projects is the core of the term. A colleague suggested “puzzling” as an alternative, and the word isn’t set in stone. “Assisting” is explicitly ruled out, because it sounds far too competent for something that doesn’t know what it is doing.
Why Generated Unit Tests Often Hit Only One Equivalence Partition
Tests generated by AI often look clean and still cover only a fraction of what matters. In one example, the generated set contained six unit tests that all landed in the same equivalence partition. Five representatives of the same case, and the very cases an experienced tester would have thrown in immediately were missing entirely.
That is exactly the trap for beginners. If you don’t know what an equivalence partition is, you can’t assess the output. The code looks good, the tests go green, and the coverage is still weak. Assessing it takes domain knowledge and test design knowledge, and you can’t borrow either from the AI.
This raises an uncomfortable point for training. Juniors who hand everything to the AI never reach the point where they know enough to judge the results. If you want to learn, there is no way around doing some of the work without AI.
“I wouldn’t have a problem with a junior developer saying: unit tests, annoying, give me a few ideas and I’ll take it from there. The problem starts when someone says: I don’t care, do you have unit tests? Yes, there they are.”
(Christian Brandes)
Ask Whether It Makes Sense Before Asking Whether It Works
Before you put an AI on a task, ask whether that makes sense at all. Joseph Weizenbaum made the point early on in “Computer Power and Human Reason”: what matters is not what a computer could do, but whether it makes sense for it to take on a particular task. The same question is on the table with AI today.
Acceptance criteria are a concrete case. Plenty of tools promise to generate perfect user stories, acceptance criteria included. But acceptance criteria are supposed to express what a product owner needs to see in order to be convinced. How is an AI supposed to know what would convince a person? And if someone can’t come up with anything they would accept, the real problem lies somewhere else.
The market is heading the other way. Vibe coding and headlines about AI-generated code feed the idea that entire developer roles can be cut. The industry may have to learn this the hard way: crash into the wall once with untested, reused code, then take two steps back.
Human in the Loop Testing: Keep Testing, Keep People Involved
Two principles sum up the stance. First, don’t throw out exploratory testing. The subconscious and intuition produce test ideas that no script and no model will come up with on their own. Second, keep people in the process.
A workflow where someone sketches an idea and then several AI agents pass the ball around, one coding, one testing, one reviewing, carries too much risk. A human brain needs to come in somewhere along the way. Do you really want one AI to quality-assure the output of another?
In practice, that means a clear division of labor. The AI can take the busywork, such as generating bulk test data or producing a first draft so you don’t start from a blank page. As soon as things get more complex, your own head has to be involved to assess the output. The thinking stays with humans.
Frequently Asked Questions
What kinds of bugs does exploratory testing uncover that planned test cases miss?
Bugs that arise from unusual user sequences. In one example, a price increased by one cent after someone had switched back and forth between two states multiple times, a reproducible issue with significant consequences. In another case, after half an hour, the input validation suddenly allowed pasting via a keyboard shortcut: There was a decimal number in the clipboard, and the browser crashed.
Can AI replace exploratory testing?
No. A model can generate many such test cases as candidates, but you have to steer it there first: think of different environments, different usage scenarios. The impetus doesn’t come on its own. During a training session on a preschool learning laptop, a single participant asked if the device could be operated without a mouse because his child uses it on their lap in the car. He found many bugs.
What does the term “digging” mean in software testing?
“Digging” refers to a third category alongside “testing” and “checking”: what an AI model does in test design. Testing relies on intuition, curiosity, and experience; checking relies on a script. Digging is based on training data and probabilities: blindly sifting through test ideas from other projects without understanding the task. “Puzzling” was suggested as an alternative, while “Assisting” was explicitly rejected.
Can an AI determine whether a result is technically correct?
Only approximately. In many cases, a human knows with certainty whether an implementation or an output is correct. With a model, it remains a probabilistic statement, which can have very high reliability. In critical areas, however, the question remains open as to whether one wants to rely on it.
How reliable are AI-generated unit tests?
They often look clean but still provide limited coverage. In one example, all six generated unit tests hit the same equivalence partition, so they were five representatives of the same case. Ironically, the inputs that an experienced tester would have entered immediately were missing. Tests that pass therefore say little about actual coverage.
Why is using AI risky for those new to testing?
If you don’t know what an equivalence partition is, you can’t evaluate the generated output. Getting ideas from the AI and then continuing the work yourself is not a problem. It becomes problematic when someone accepts tests without reviewing them and simply reports that they exist. In that case, the knowledge needed for evaluation is never gained.
Which software testing tasks can be meaningfully delegated to AI?
Routine tasks. These include generating bulk test data and creating an initial draft so you don’t start with a blank page. As soon as things get more complex, you need to use your judgment to evaluate the output. Workflows in which multiple agents pass the ball back and forth (one codes, one tests, one reviews) carry too much risk.
Can AI formulate acceptance criteria for user stories?
It can generate suggestions but easily misses the mark. Acceptance criteria are meant to express what a product owner wants to see in order to be convinced. A model cannot determine what convinces a human. And if someone can’t even think of what they should accept, the real problem lies elsewhere.


