Risk-based testing means focusing test effort where a defect would do the most damage. Three factors determine a risk: damage potential, detectability and complexity. Multiplying them gives a risk score, and the higher that score, the more test coverage the area needs.
Key Takeaways
- At the Bison Group, risk-based testing made it possible to drop 50 to 60 percent of the old end-to-end tests without replacement, because they were moved to lower test levels or turned out to be unnecessary.
- Risks are easiest to break down by domain, such as sales, goods receipt or master data, and then to name specifically what must never go wrong in each one.
- Each risk is rated on damage potential, detectability and complexity, each on a scale of 1 to 3, and the multiplied result is a score that makes prioritization decisions more objective.
- Risk-based testing builds quality awareness along the whole development chain, from requirements analysis to implementation, because everyone involved can see where defects do the most damage.
- A risk catalog is never a one-off document: it has to be reassessed with every new requirement and after every bug report to keep its steering effect.
Risk-Based Testing Starts with a Headline
The quickest way into risk-based testing is one simple question: which headline would you not want to read in tomorrow’s newspaper? Turn that headline into a risk, and the risk tells you where testing needs to happen.
The example behind it is easy to picture. An invoice run produces 4,500 invoices, and somewhere along the way the system adds up wrong or applies the wrong tax rate. The customer has to recall the invoices, and the damage is high on both sides. Cases like this mark the point where test effort pays off.
The strength of this question: the customer knows best where it hurts most. Ask them directly and you have often pinned down your damage potential before a single test case has been written.
How a Risk Is Assessed
At the Bison Group, risk assessment rests on three components: damage potential, detectability and complexity. Each one is rated on a scale of 1 to 3, and the ratings are multiplied. The higher the resulting score, the more closely you need to look.
Uwe Paesch walks through the logic with the invoice example. The damage potential is high because the customer has to recall thousands of invoices. Detectability is rated high because the error isn’t visible at first glance. Complexity is high because the error is hard to track down. Three high values add up to a high risk.
A high total isn’t the only signal. If you rate everything high, you automatically get high numbers. So also keep an eye on risks with high damage potential, even when they are noticed quickly and easy to fix. If a defect like that reaches the customer, they will still ask more than once how that could happen.
Slicing Risks at the Right Level in an ERP System
The first step isn’t the assessment, it’s the slicing. Define topic areas first; after that, capturing the risks underneath gets much easier.
For an ERP system, domains work well: sales, ordering, goods receipt, delivery, interfaces, master data. For a mobile app, it is more the app’s own functionality and its communication with surrounding systems. Only once these clusters are in place do you ask, area by area, what must never go wrong.
The most common trap is the wrong level of detail. “Sales doesn’t work” is too coarse as a risk; individual edge cases are too fine. Miss that level and the assessments become useless, some too low, others far too high. For an ERP system, you end up with around 150 risks, about a dozen of them in sales.
If you are just starting out, start small. Subject areas first, then break them down. The first steps count for more than a perfect catalog.
Why a Migration Is the Ideal Occasion for Risk-Based Testing
A migration is the perfect moment to switch to a risk-based approach. Bison faced exactly that: replacing a rich client with a web client, with around 1,000 end-to-end tests that were supposed to be carried over.
Instead of rebuilding every old test one-to-one, the decision was to put test coverage where it would hurt the customer most. The team drew up a risk catalog, identified the highest risks and brought business people, developers and architects into the assessment. Only then did they check which tests already existed, whether they needed extending and where new test cases were required.
The result was a noticeable relief. Depending on the system, between 50 and 60 percent of the old end-to-end tests were dropped, because the risk coverage showed they simply weren’t needed.
Why End-to-End Tests Aren’t the Default Answer
End-to-end tests are expensive to build, expensive to maintain and expensive in feedback time. Before you write one, ask yourself: do I really need this at this level?
If an end-to-end test ends up checking only one small detail, it belongs two levels lower. Move the test down one or two levels and you get your feedback much faster and save on maintenance.
At Bison, automation already runs on a tight schedule. All unit and module tests run with every build, and twice a day full test runs follow, end-to-end tests included. So the real work shifts to finding the right test level, not to whether to automate at all.
Risk Assessment Is a Living Process, Not a Law
A risk catalog is never finished. The ratings change with the feedback that comes back from the field.
That feedback comes from several sources: requirements engineers, the customer and bugs. At Bison, each bug was tagged with the risk it violated, first in Confluence catalogs, later directly in Jira so that the links would last. That revealed a correlation: the higher the risk index, the more severe the bugs.
New requirements trigger a reassessment as well. When an idea comes from the customer, you check whether it introduces a new risk or violates an existing one, and you run through the assessment again.
Over three years, this produced a clear picture. Blocker and high-severity bugs, the ones that hurt the customer, show a clear downward trend.
How to Agree When Assessments Differ
When one person sees a major risk and another sees none, the team discusses it instead of voting. Each side hears the other’s arguments, and then the group looks for consent.
This usually goes quickly and stays factual. In the end, someone says they can live with how the others see it. Only a genuine veto forces a longer debate. Friction mostly comes up around detectability, meaning how fast an error becomes visible, and around complexity, where opinions differ.
The Biggest Side Effect Is Quality Awareness
Risk-based testing reaches beyond the test team. It builds awareness of where to look closely, and that awareness runs through the entire chain.
“It’s not just the developer who sees that three risks are assigned to a story and then tests especially carefully. It runs through the whole chain, all the way to the requirements engineer who talks to the customer.”
(Uwe Paesch)
In practice, the requirements engineer can tell the customer that a certain implementation will break the pricing. Developers get a better sense of how complex a function really is for the customer, instead of just working through their own story in their own bubble. The risk discussion turns into aha moments about what is happening out there.
Where Risk-Based Testing Goes Next
Spreading the approach across the company takes time. At Bison, it started in the ERP teams, driven by the legacy portfolio, and other teams joined step by step. It will be a while before it applies everywhere.
The road there wasn’t smooth. In 2022, the goal was to practice risk-based testing across the entire group, which proved too ambitious. In 2023, it was made official and backed by a course that teaches the procedure, purpose and method.
The next milestone is company-wide rollout, followed by more automation. That means not only tests that run automatically, but tests that are generated from use cases and risk descriptions or proposed as suggestions. This is exactly where AI support is expected in the coming years.
Frequently Asked Questions
How do you figure out where testing efforts are actually worthwhile?
A quick way to get started is to ask the customer what headline they wouldn’t want to see in the newspaper tomorrow. This headline is used to formulate a risk, from which the testing requirements are derived. A typical example: An invoicing run generates 4,500 invoices; the system totals them incorrectly or applies the wrong tax rate, and the customer has to recall everything.
Is the multiplied metric sufficient as the sole criterion for prioritization?
No. If you assign high ratings across the board, you automatically end up with high numbers, and the metric loses its meaning. Therefore, you should also keep an eye on risks with high potential for damage that can be quickly identified and easily resolved. If such an error reaches the customer, they’ll still ask three times whether it’s still acceptable.
How many risks should a useful risk catalog include?
For an ERP system, there are about 150 risks; in sales, about a dozen. This number results from the level of detail: “Sales aren’t working” is too broad, while individual, detailed cases are too narrow. If this balance is missed, the assessments become useless: some will be too low, others far too high.
Should existing end-to-end testing be carried over one-to-one during a system migration?
No. When the Bison Group replaced a rich client with a web client, there were approximately 1,000 end-to-end tests to be transferred. Instead of recreating them, a risk catalog was first created and evaluated in collaboration with business stakeholders, developers, and architects. Afterward, depending on the system, 50 to 60 percent of the old tests were eliminated without replacement.
How can you tell if a test runs better at a lower test level?
One indication is when the end-to-end testing ultimately checks only a minor detail. In that case, it belongs one or two levels lower. End-to-end testing is expensive to create, maintain, and provide feedback on. At a lower level, feedback comes much faster, and maintenance costs decrease.
How do you keep a risk assessment useful over time?
The assessments are updated based on real-world feedback: from requirements engineers, from the customer, and from bugs. At Bison, the risk violated by each bug was noted (initially in Confluence, later in Jira) to ensure the links remained valid. This revealed a correlation between a high risk index and a higher bug level. New requirements also trigger a reassessment.
How does risk-based testing benefit roles outside the testing team?
It highlights exactly where closer scrutiny is needed and has an impact throughout the entire development chain. A developer who sees three risks associated with a story will test more carefully. The requirements engineer can tell the customer that a specific implementation will break the pricing model. Developers also gain a better understanding of just how complex a feature really is for the customer.


