Exploratory ensemble testing, formerly known as mob testing, is a method in which a group tests software together at one computer: no predefined test cases, but clearly rotating roles (driver, navigator, ensemble) on a fixed timer of eight to ten minutes. The starting point is a charter, a rough idea of the area to explore. Each experiment shapes the next.
Key Takeaways
- Ensemble testing only works when everyone sticks to their role: the driver makes no decisions, the navigator never touches the computer, and the timer forces a rotation every 8 to 10 minutes.
- The concrete trigger for introducing exploratory ensemble testing was late stakeholder feedback shortly before release dates, not methodological idealism.
- The stated goal of a session is not to find bugs but to learn something about the system, guided by a charter. Bugs are a by-product, not the yardstick.
- Session notes are deliberately kept neutral and only reviewed one or two days later with product and project management, which keeps the “bug or not a bug” debate out of the testing flow.
- Developers build testing knowledge implicitly through repeated sessions. You can tell when they predict the edge cases the group is about to try, because they already tried them themselves.
What Is Exploratory Ensemble Testing (Mob Testing)?
In exploratory ensemble testing, which many still know as mob testing, a group of people sits together at one computer and tests software in clearly assigned roles, without predefined test cases. Instead of ready-made scripts, there is a charter: a rough idea of what to explore. The first experiment gives rise to the next ones.
The name ensemble testing replaces the older term mob testing. The switch came once it was clear that “mob,” with its closeness to “mobbing” in the sense of bullying, was a poor choice of word. Ensemble also describes the practice more precisely: a group working on one task together.
The method brings two practices together. The role model comes from collaborative work on code, the open-ended approach from exploratory testing. The two mesh because the group decides together what to try next.
How the Driver and Navigator Roles Work
An ensemble has three fixed roles: driver, navigator, and the observing ensemble. That separation is the heart of the method.
The driver sits at the computer and operates it, but makes no decisions. The navigator decides what happens next, but never touches the keyboard. If both stay silent, nothing happens. A test step only happens when they talk to each other. That forced communication is the point.
The rest of the group watches and chips in ideas. Still, only the navigator decides what gets done. If you are set on trying an idea, you wait until you are the navigator. Then it happens.
The roles rotate on a fixed timer, in practice every eight to ten minutes, so everyone gets a turn in every position. A full session usually runs about an hour and a half: at least one hour of pure testing, the rest for the intro and wrap-up.
Start With a Clearly Named Problem
Ensemble testing needs a concrete reason, or it never gets going. For Tobias Geyer, the trigger was a recurring pattern: problems surfaced in late test phases, and feedback from internal stakeholders kept arriving too late to make it into the current release.
The release cadence made it worse. With two to four releases a year, there is no quick follow-up patch. Late feedback often means months of delay, and that is exactly what frustrated many people on the team.
The method was not introduced as a mandatory process. It was proposed as an experiment, and that framing helped. Instead of enforcing a new rule, the question on the table was whether it would solve the problem. The feedback was positive, and it has run regularly ever since.
If you want to introduce ensemble testing, this route works: name a real pain point, call it an experiment, invite people widely, and let them decide for themselves whether it is worth it.
How Often and With How Many People: A Four-Week Rhythm
How often you run it decides whether ensemble testing becomes a habit or a chore. A four-week rhythm tied to the sprint has worked well.
The interval does two things. First, at least one sprint has passed, so there is a new feature or at least a new angle, such as a familiar feature seen through a different persona. Second, the session keeps some novelty, because it doesn’t happen all the time.
The hour and a half is also a deliberate break from day-to-day work. People see it as a benefit, not a burden. One sign: canceled sessions lead to complaints, not relief.
About five people per ensemble is a sensible starting point. Once there are ten or more, split the group so that rotation and participation still work. These numbers aren’t a law. They are experience you should adapt to your own context.
What Kinds of Bugs Ensemble Testing Finds
The range goes from simple click problems through exceptions to usability weaknesses. In practice, usability made up a surprisingly large share of the findings.
Above all, the sessions expose what small-scale, feature-driven development misses. When a charter follows a larger user workflow through the whole product, and the participants actually launch the software instead of relying on unit tests, the fault lines between features show up.
One example makes it concrete. A team had tested an input field for formatted text, including moved lines and deleted passages. Everything worked. Then a product manager, acting as navigator, suggested moving the text with drag and drop. Everyone thought that was impossible. It worked, and it promptly threw an exception.
Findings like this happen because people with different perspectives sit together in the ensemble. What one person considers unthinkable, another simply tries.
The Real Gain Is Testing Know-How in the Development Team
Ensemble testing doesn’t just find bugs. It builds testing expertise among developers, and that effect lasts well beyond a single session.
In the sessions, participants pick up test ideas and test design methods. They develop a feel for what needs checking beyond the happy path, and they carry that into the next sprint, explicitly or not.
One moment captures the learning: the ensemble thinks it has found an edge case that will take the software apart. The person who built the feature is sitting right there and says they already tried that. That is the moment test thinking has moved into development.
For many, the method also clears away a barrier: testing can be fun. People who used to see exploratory testing as a nuisance experience it in the ensemble as something shared and lively.
Keep Evaluation and Testing Strictly Separate
During the session, nobody debates whether something is a bug. That separation keeps the flow going and prevents long, sticky arguments.
Testing borrows an attitude from improv: “Yes, and.” There are no bad suggestions. Ideas get picked up so the group keeps moving. Notes stay as neutral as possible. An exception or a crash is obvious. Everything else is only recorded at first, not judged.
The evaluation happens one or two days later in a small group with product and project management. Only then does anyone decide what is a real bug, what can be accepted, and what might just be a misunderstood product that needs better documentation.
A clear goal matters too. The point is not to break the software or to find as many bugs as possible. Bugs turn up anyway, because all software has them. The goal is to learn something about the system, guided by the charter.
What Are the Pitfalls of Ensemble Testing?
Two problems hit almost every team: note-taking and role discipline.
Note-taking is underestimated. Caught up in testing, the group easily forgets to write down what it did and found, and at the end of the session it is hard to reconstruct the findings. When you are in the flow, you don’t think about the record, which is exactly why you need a dedicated note-taking role or a reminder.
Role discipline slips most when new people join or an ensemble starts from scratch. One pattern is creative rule-bending. Someone says to the navigator: “Tell the driver to click button A, then enter XY, and select feature Z.” Technically correct, and entirely beside the point. The only fix is a friendly reminder of the rules of the game.
The timer is another trap. In the middle of an action it is easy to miss, and suddenly you are ten minutes over. Rotation only creates discipline if people stick to it.
Internal Stakeholders Need Their Own Format
Bringing internal customers into the ensemble yields valuable feedback before the real customer sees the product, but it can tip the group dynamic. They tend to dominate.
The answer is two types of session. One runs with the development team. The other brings in internal stakeholders, joined by a few people from development. That creates an exchange of knowledge without one group taking over.
Participation remains the open issue. The development team is much larger than the circle that joins regularly. Some never came, others dropped out over time. Bringing more people back is an explicit goal.
“In the end, it’s a meeting invitation, maybe you need a meeting room. An hour and a half flies under any management radar. Just do it and be surprised by how much people like it.”
(Tobias Geyer)
Frequently Asked Questions
Why was the term “Mob Testing” replaced with “Ensemble Testing”?
The term was changed because the similarity between “mob” and “mobbing” makes it an inappropriate choice. “Ensemble” also describes the concept more precisely: a group working together on a task. The renaming doesn’t change the process. The roles of Driver, Navigator, and Ensemble remain, as does the rotation at fixed time intervals.
Do you need prepared test cases for exploratory testing in a group?
No. Instead of pre-written scripts, there is a charter, a rough outline of what needs to be investigated. The next steps emerge from the first experiment, and the group makes decisions as they go. The stated goal isn’t to find as many bugs as possible, but to learn about the system based on the charter.
How do you convince a team to try out a new testing practice?
Identify a real pain point and frame the approach as an experiment, not as a new mandatory process. For Tobias Geyer, the trigger was feedback from internal stakeholders that regularly arrived too late to be included in the current release. With two to four releases a year, that quickly adds up to months of delay. Invite a wide range of people and let the participants judge for themselves.
How often should joint testing sessions take place, and how large should the group be?
A four-week cycle, aligned with the sprint, has proven effective: there’s a new feature (or at least a new perspective), and the session retains its novelty. About five people per group is a good starting point; if there are ten or more, split them up. A session lasts about an hour and a half, with at least one hour of pure testing time.
Why does a group find bugs that individual testers overlook?
Because people with different perspectives sit side by side. One team had thoroughly tested an input field for formatted text, including moved lines and deleted passages. Everything worked. Then a product manager, acting as the navigator, suggested moving text via drag-and-drop. Everyone thought this was impossible, but it worked, and it promptly threw an exception. This particularly highlights the gaps between individual features.
When is it decided whether a finding is actually a bug?
Not during the session. During the session, observations are recorded as objectively as possible: An exception or a crash is clear-cut; everything else is simply noted for now. One to two days later, a small group including product and project management evaluates what constitutes a real bug, what can be accepted, and what is simply a misunderstood product with incomplete documentation.
How do developers benefit from participating in testing themselves?
They build testing knowledge that extends beyond the individual session. During the sessions, they learn about testing ideas and methodologies and develop a sense of what needs to be tested beyond the “happy path.” The effect becomes apparent when someone predicts a supposed edge case in the test suite because they’ve already tried it out themselves while building the feature.
Should internal stakeholders participate in joint testing sessions?
Yes, but preferably in a separate format. Internal customers provide valuable feedback before the actual customer does, but they tend to become dominant within the group. Dividing the sessions into two types has proven effective: one with the development team and one with internal stakeholders, supplemented by individual representatives from the development team.


