Risk-based testing means ranking test cases by their risk value, so that the most important tests run first when time gets tight. A multidimensional approach rates five levels separately: business logic, test history, release scope, tester assignment and code changes. The values from each level are then combined into one aggregated risk score per test case.
Key Takeaways
- Many projects practice risk-based testing intuitively, but without a structured risk metric, the decision about which test cases go into a release is nothing more than gut feeling.
- A level model with five perspectives (business logic, test history, release scope, tester assignment and code changes) produces an aggregated risk score for each test case instead of one rough high, medium or low rating.
- Using Fibonacci numbers as the rating scale means a single high risk value dominates the overall score, and low values on other levels can’t cancel it out artificially.
- Test case complexity, measured by the preconditions and the number of steps compared with the other test cases in the project, feeds directly into the probability that a defect occurs.
- A risk cut-off set before the test run determines which test cases are included at all, so when time runs short, the cases that drop off at the end are demonstrably the ones with the lowest risk.
What Multidimensional Risk-Based Testing Means
Multidimensional risk-based testing rates the risk of a test case from several perspectives at once instead of boiling it down to a single value, and that makes test case prioritization far more precise. The classic split into high, medium and low is one-dimensional. It tells you that a risk is high, but not where it comes from.
Richard Hönig at MGM has developed a level model for risk analysis that breaks up this oversimplification. The basic idea: a complex world is hard to capture on a single scale. If you only rate risk globally, you lose the information about why a test case is critical.
Instead, the model looks at every test case through several separate lenses and only brings the results together at the end. What does the business logic say? What does the test history say? What matters for the current release? These individual ratings add up to an aggregated risk score per test case.
The Trouble with Gut Feeling
Many projects do risk-based testing without calling it that. Teams decide intuitively which test cases matter for a release and which don’t. Usually they decide on gut feeling, with no metric and no scale.
That is the weak spot. The gut feeling of experienced stakeholders is valuable, but nobody else can follow it or reproduce it. As soon as the project grows or people change, there is no shared basis for prioritization.
The answer isn’t to replace gut feeling but to back it up with data. At the start of a project, when there is hardly any data, the experience of developers, architects and testers is still the best source. Every test run adds data that sharpens that assessment.
The Five Levels of Risk Analysis
Richard’s model works with five perspectives, each rating every test case on its own. Each level produces its own value, which later goes into the overall score.
| Level | What it rates |
|---|---|
| Business logic | Probability and potential impact of a defect, including test case complexity |
| Test history | Failure rate of past runs, defect tickets raised and their priority |
| Release scope | Links to requirement tickets that are part of the current release |
| Tester assignment | Expertise and available time of the person running the test case |
| Code changes | Which code areas a test case is linked to and what has changed there |
The business logic level builds on what risk-based testing has always used: probability times impact. What’s new is that test case complexity enters the probability as a measurable factor.
Tester assignment is a level many models overlook. A tester who is spread across several projects, or who barely knows the subject area, adds risk. That has nothing to do with the test case itself, but it does affect the outcome.
The code changes level closes a typical gap. A tester reads a requirement and knows exactly what to check. Even so, defects show up in places nobody had on their radar: a refactoring in the background, an updated library. Changes like these affect the risk without showing up in the requirement.
How Test Case Complexity Affects the Probability of Failure
A complex test case is more likely to fail, because more steps mean more places where something can go wrong. The model looks at how extensive the preconditions are and how many test steps a test case contains.
What counts is how a test case compares with the other test cases in the project. A test case with 60 steps isn’t complex in absolute terms. It is complex compared with the small test cases next to it. If every test case in the project is a monster like that, the value tells you nothing.
This relative rating keeps a single heavyweight test case from dominating everything automatically. The end-to-end business case that touches a thousand requirements and has often been unstable doesn’t stand out just because it is big. It gets put in proportion.
Why the Fibonacci Sequence Is the Right Scale
The model rates risks with Fibonacci numbers because they grow exponentially by themselves, so high risks carry more weight in the overall score. The rough three-way split into low, medium and high offered too little nuance and too little flexibility.
The ranges follow the number of digits. The single-digit Fibonacci numbers from 1 to 8, that is 1, 2, 3, 5 and 8, stand for low risk. Two-digit values mean medium risk, three-digit values high risk.
You see the effect when the values are aggregated. If a test case scores a three-digit value on one level and 5 or 8 on the others, the low values don’t simply offset the high one. The average stays high, because the exponential spread carries the high risk all the way through.
“If I have a value of, say, 610 there, and the other levels show a low risk of 5 or 8, you still end up with a very high number.”
(Richard Hönig)
Risk Analysis Must Not Create Extra Work
A risk analysis only holds up if it uses data that already exists. Once it demands additional documentation, the benefit turns into a burden.
MGM’s test management tool already brings the necessary information together: linked requirements, linked defect tickets, previous test runs. For projects that have their documentation reasonably under control, linking is standard practice and costs testers no extra time. The backend calculates the risk values from this data automatically.
No system gets every assessment right. That’s why the automatic calculation can be overridden for each test case. If you see the risk differently from the algorithm, you set the value or individual parameters yourself.
How the Risk Score Drives Test Case Prioritization
The aggregated risk score becomes a tool for steering the test run, not just a nice overview. Managers see the risk distribution across the project at a glance. For testers, the practical benefit is bigger.
When you put a test run together, you can set a cut-off: only test cases above a certain risk value go in, and the rest stay out for now. Within the run, the test cases are sorted by risk, with the most critical at the top.
That pays off when time runs short. If you test from the top down, you have always covered the highest risks first. If something drops off at the end, it’s the test cases with the lowest risk. Annoying, but you can live with it.
Next Step: Industry Context and Weighting
Different projects and industries weight the risk levels differently, and a rigid model can’t do that justice. One team sees complexity as the deciding factor, another the business logic. Freely adjustable weighting of the levels is still in MGM’s backlog.
Timing can shift the risk too. In car insurance, contracts have to be renewed at the end of the year, and the rush on the forms starts in the fall. A defect in the same area weighs less in January than in the fall. Seasonality of this kind is meant to go into the model as well.
Further levels, such as security, are on the horizon. The real step toward maturity goes further, though: giving test managers a framework to define their own levels. They know best where their risks are, and a tool vendor can’t think that through in advance for every project.
Frequently Asked Questions
Why Is a Classification into High, Medium, and Low Insufficient for Testing Risks?
It remains one-dimensional. It states that a risk is high, but does not indicate where it comes from. A multidimensional model evaluates business logic, test history, release scope, tester assignment, and code changes separately and only combines the values into a single score at the end. This makes it clear why a single test case is critical.
How can risk be assessed when no test data is available at the start of the project?
In that case, the experience of developers, architects, and testers is the best source. The goal is not to replace gut instinct, but to support it with data. With each test run, failure rates, defect tickets, and correlations are added, which refine the initial assessment. Without metrics, prioritization is neither traceable nor reproducible once the project grows or personnel change.
Does a test case with 60 steps automatically constitute a high risk?
No. Complexity is always assessed relative to the other test cases in the project. Sixty steps are complex compared to small test cases, but lose their significance if all cases are of this scope. The scope of the preconditions and the number of test steps are evaluated, and the result is factored into the probability of a defect occurring.
Why do code changes play a role in prioritizing test cases?
Because defects occur in places that aren’t visible in any requirement: a refactoring in the background, an updated library. A tester may have fully understood the requirement and still miss the critical area. The “code change” level therefore checks which code areas a test case is linked to and what has changed there.
Which rating scale is suitable for risk values per test case?
Fibonacci numbers, because they grow exponentially on their own. The breakdown is based on the number of digits: single-digit values from 1 to 8 represent low risk, two-digit values represent medium risk, and three-digit values represent high risk. When aggregated, low values at other levels do not offset a three-digit value. The spread carries the high risk through to the overall score.
What happens if a tester assesses the risk differently than the automated calculation?
The calculated value can be overridden for each test case, either completely or for individual parameters. No system gets every assessment right. The automated system derives its data from information that’s already available: linked requirements, associated defect tickets, and previous test runs. If a risk analysis requires additional documentation, its benefits are outweighed by the extra workload.
How does a risk score help when test time is running short before a release?
Test cases are sorted by risk during the test run, with the most critical ones at the top. Working from top to bottom ensures that the highest risks are addressed first. If cases are omitted at the end, they are the ones with the lowest risk. In addition, a cutoff can be set in advance, below which test cases are not included in the run at all.
Can the risk levels be adapted to the industry and project context?
That is precisely where the next step in the maturity process lies: Test managers should be able to define their own levels, because a tool provider cannot anticipate the risks for every project. A freely configurable weighting of the levels is still on MGM’s backlog. Seasonality should also be factored in, for example in auto insurance, where an error in the fall carries more weight than one in January.


