Skip to main content

Search...

Test Coverage: Process Mining Shows What Really Runs

Test coverage needs the right denominator. Process mining shows which workflows run in production, and the top 10 of 600 runs cover about 80 percent.

• • Updated: • 9 min read
Cover of the expert talk on 'Test Coverage: Process Mining Shows What Really Runs' with Sven Braxein, Athanasios Kallinikidis and Richard Seidl.

Process mining gives test coverage a solid basis: production data shows which software workflows actually run and how often. Teams can derive targeted test cases from that instead of relying on assumptions. Knowing the most frequent runs lets you justify coverage with data and focus regression tests on what matters.

Key Takeaways

  • Process mining shows which workflows actually run in production, giving teams a data-based foundation for choosing regression test cases instead of relying on gut feeling.
  • A single workflow can have more than 600 different runs, yet the top 10 already cover around 80 percent of all production runs.
  • Workflows that occur often in production but aren’t covered at all in testing expose concrete risk gaps that stay invisible without analyzing production data.
  • A manual analysis of production data loses its value as soon as the person doing it leaves the project. Only automation keeps the approach sustainable.

Test Coverage Starts with the Denominator

If you claim test coverage is too low, you first have to say what that claim refers to. At its core, test coverage is a percentage: numerator divided by denominator. Without a defined denominator, any verdict of “too little” or “enough” is just an opinion.

That is exactly where Sven Braxein’s approach starts. In a project replacing a long-lived legacy application with a new system, someone complained that coverage in regression testing was too poor. His first question back: what exactly is the denominator you’re measuring against?

The question is more than rhetorical. It forces a team to put its reference point on the table. Only once it’s clear what should be fully covered can you put a number on how much of it has actually been tested.

What Process Mining Makes Visible

Process mining combines classic process analysis with data mining and shows how processes really run in production. Not the target diagram someone drew, but what actually happens.

Every click and every action in a digital system leaves a trace. A process mining tool collects those traces and links them together. The result is a view of what really happened in the system, instead of an assumption about what should have happened.

The technology has been around for a little over ten years. Athanasios Kallinikidis came across it during an internship and later introduced it at his company for the new software. The prerequisite was knowing the system’s data model well enough to connect the right data points.

Soccer Analytics as a Model for Test Prioritization

The idea came from amateur soccer. Athanasios coaches a men’s team in a local league and records their matches with a camera on a tall tripod. Software then produces statistics from the footage: heat maps, running paths, where a shot came from.

The guiding principle: more knowledge, less opinion. If a player claims he ran a lot, the analysis shows how far he actually ran. From what really happened in Sunday’s game, the coach works out what to practice on Tuesday.

That logic carries over to software development. If a local amateur team can learn from its match data for training, a large IT project should be able to learn from its production data for testing. You analyze what really happens and align your testing with it.

“We look at what happens in reality, derive what we do in testing from that, and end up with a more stable production system.”

(Athanasios Kallinikidis)

Why Workflows Beat Business Processes as the Denominator

In this project, workflows turned out to be the more useful denominator, not business processes. Business processes are used in production, but in testing they are often underestimated and hard to pin down.

At its core, the legacy application is one big workflow engine, similar to the status workflows in a ticketing system. The project counts around 160 different workflows: creating a customer, creating a contract, changing it, ending it or settling it early. Each of these 160 workflows is modeled and can run in different ways.

That defines the denominator: 160 workflows. The numerator is the number of workflows actually run in testing. An opinion about coverage becomes something you can count.

Why the Top 10 Runs Beat Full Coverage

Risk-based testing means covering the few runs that matter instead of every possible one. Process mining shows which runs occur how often, and that makes the prioritization defensible.

One example from the project shows the ratio. The most complex workflow has more than 600 different runs. The most frequent one alone accounts for about 20 percent. The top 10 runs together reach around 80 percent of all cases.

That leads to a clear decision: if you have test cases for the top 10 runs, you can skip the remaining 590. Not out of convenience, but because the data shows that’s where the relevant share of real-world activity lies. And you can prove why you picked exactly those ten.

Production Data Reveals Blind Spots in Testing

Comparing production with testing brings up workflows nobody had on their radar. Some are harmless, others are real gaps.

In the project, the three most frequently run workflows weren’t on the expected list at all. The most frequent one is a blacklist check via an interface that decides whether a customer gets a contract. That very workflow causes regular trouble in testing, because the interface often doesn’t work. In production, it’s the one triggered most often.

The comparison produces three kinds of findings:

  • Green: A workflow runs often in production and is adequately covered in testing. For very simple workflows with only two runs, two good test cases are enough, not three thousand.
  • Yellow: A frequent production workflow barely shows up in testing. Worth a closer look to see whether that can be right.
  • Red: A workflow is among the top 10 in production but isn’t triggered even once in testing. That can’t be right and needs investigating.

When you get a red finding, you send someone to look into it. Either the workflow runs in production for no reason and nobody needs it anymore, or it has to run and is missing from testing, which is a serious problem.

Without Automation, the Benefit Fades

The biggest lever isn’t the one-time analysis but the repetition. If you collect and aggregate data by hand, you depend on one person continuing to care enough.

In the project, the analysis currently runs manually at regular intervals. That works as long as someone keeps it going. If that person leaves the project, the analysis will probably stop too. The manual effort is the real weak spot.

Lasting value only comes once collecting, aggregating and correlating the data is automated. Then a continuous improvement process kicks in: the system flags anomalies regularly, and you only follow up where a red finding appears.

Balancing Effort and Benefit

More data doesn’t automatically mean more value. This approach is subject to diminishing returns too, and the project deliberately drew a sensible line.

The analysis covers the central system, with other systems connected around it. You could include all the surrounding systems and get more complete measurability. The question is how much effort that takes and how much extra information it yields.

The current scope strikes a good balance between effort and benefit. The goal isn’t maximum completeness, but getting as much usable information as possible with as little manual effort as possible.

Frequently Asked Questions

How can we determine whether a project’s test coverage is sufficient?

At its core, test coverage is a percentage calculation: numerator divided by denominator. Without a defined denominator, any judgment about whether coverage is “insufficient” or “sufficient” remains merely an opinion. In a legacy software replacement project, the counterquestion to the accusation of poor regression test coverage was therefore: What exactly is the denominator? Only when the reference value is clear can coverage be quantified.

What does process mining reveal that a process model does not?

It reveals the actual, real-world workflow as it is performed, rather than the theoretical diagram. Every click and every action in a digital system leaves a trace. A process mining tool collects these traces and establishes relationships between them. This requires that someone knows the system’s data model well enough to link the correct data points.

Are business processes or workflows better suited as a benchmark for test coverage?

In the project described, workflows proved to be a more suitable common denominator. While business processes are used in production, they are often underestimated in testing and remain difficult to grasp. The legacy software was essentially a workflow engine with around 160 modeled workflows: create customer, create contract, modify, terminate. The count is based on the workflows actually run during testing.

Do test cases need to cover every possible execution of a workflow?

No. The most complex workflow in the project had over 600 different executions. The most frequent one alone covered about 20 percent of production cases, and the top 10 together covered about 80 percent. If you have test cases for these ten, you can omit the remaining 590 and justify the selection using production data.

How do you identify workflows that are missing from testing but are important in production?

By comparing production frequency with the workflows triggered in testing. Green means: frequent in production and adequately covered. Yellow means: frequent, but rarely represented in testing. Red means: among the top 10 in production, but not triggered even once in testing. If it’s red, the workflow is either unnecessary or missing from testing.

Is a one-time analysis of production data sufficient?

No, the key lies in repetition. If data is collected and aggregated manually, everything depends on it remaining a priority for one person; if that person leaves the project, the analysis usually ends as well. Only automated collection and correlation support a continuous improvement process in which you need only follow up on red findings.

Should you include as many connected systems as possible in the analysis?

Not necessarily, because the law of diminishing returns applies here as well. In the project, the central system is analyzed, with other systems connected to it. Including all surrounding systems would ensure complete measurability, but it requires significant effort. The goal is not maximum completeness, but rather a wealth of actionable information with minimal manual effort.

Share this page