Selecting regression test cases means that on each code check-in, only the test cases actually affected by that change run automatically. The basis is a mapping built every night: which files, classes or functions does each test case load? When one of them changes, only the matching tests run. This cuts runs of several hours down to often 10 to 15 minutes.
Key Takeaways
- Combining a nightly analysis of full test runs with C++ dependency tracking at compile time makes fully automatic test case selection possible across language boundaries, with no manual intervention.
- With selection, a test suite of 25,000 test cases and an original runtime of 6.5 hours gets feedback down to 10 to 15 minutes for more than 50 percent of commits.
- Granularity decides how efficient the selection is. Selecting at DLL level is not enough: only analysis at function level brought C++ the same breakthrough that class level had brought Java before.
- Several weeks of validation in parallel showed that the selection did not miss a single real test failure. Every discrepancy was caused by infrastructure problems.
- Skepticism in the team at the start of the project only faded once the Java part was done and the benefit could be measured.
Why Long Regression Test Runs Hold Up Development
A full test run of six and a half to seven hours slows down any pipeline. IVU has around 25,000 test cases, and running all of them takes that long. Developers who check something in don’t want to wait for hours to find out whether their commit was clean.
The problem grows with the number of branches maintained in parallel. At IVU there are up to ten: several releases are in use at customers, and a bug fix may have to travel through four or five branches. A single check-in can therefore trigger six and a half hours of tests on several branches.
Even ten large machines running around the clock didn’t provide enough capacity. Some of the test cases had to be cut on purpose, because otherwise a backlog built up. The alternative, testing everything once a night, has its price: if a test fails in the morning, the hunt for the commit that caused it only starts then. That costs extra analysis time.
What Test Case Selection Means for Regression Testing
Test case selection means picking exactly those test cases from a large test suite that are affected by a specific code change. Instead of running all 25,000 test cases, the system only runs the ones the check-in actually touches.
At IVU the selection is fully automatic. Nobody has to step in manually. Based on the files that were checked in, the system decides which test cases are relevant and triggers only those.
The approach was developed together with the Technical University of Munich. At Professor Alexander Pretschner’s chair, doctoral student Daniel Elsner worked on the project for three years, researching how to optimize automated regression testing. Today’s procedure grew step by step out of several experiments.
The Challenge: Two Languages and a Database
IVU’s software mixes two major technologies. Large parts are written in C++, others in Java. The components build on each other: trips generate duties, duties generate the concrete work instructions for employees, and the individual stages of this chain are sometimes in C++ and sometimes in Java.
On top of that, the tests depend heavily on data in the database. Without data, meaningful testing is hardly possible, and mocking everything is too complicated, especially for the edge cases. So at the start of a test run, the system loads data into a database, which usually requires the C++ components.
Which binaries are needed differs from test case to test case. A Java program first calls a C++ binary to generate its test data and then performs actions on it, such as creating a printout or triggering an interface to a connected system. Not every Java test case needs the same C++ binaries, and some need none at all.
How Automatic Test Case Selection Works Technically
At the core of the procedure is a link between checked-in files and the test cases that actually use them. The first step was to record, on the Java side, which files each test run loads.
At first this wasn’t about the code level in detail but about loaded artifacts. Which Java class from a JAR is actually used? Which C++ library is loaded to generate data? And beyond code altogether: which XML, CSV or YAML file is opened, for example to define a test oracle or apply settings? Files like these influence a test run too.
There are two pipelines for this. The first builds and tests continuously. The second runs regularly, usually at night, executes all test cases and logs which files they open. From this collected data, the next check-in immediately shows which test cases a changed file affects.
C++ libraries are a special case, because developers don’t check in DLLs but source and header files. Here the C++ build itself helps: compilation produces the information about which files go into which DLL. This link is added as well, so that every change can be traced to the affected test cases.
The data collected every night forms a kind of index. The actual selection only looks up this index and doesn’t analyze anything again. That keeps the selection fast when tests run.
The Results: From Hours to Minutes
The selection saves 50 to 60 percent of all test cases. The team looked at a development branch and at a released branch that only receives bug fixes.
On the Java side, runtime dropped significantly. It used to be two and a half hours; on average it is now one hour. In many cases developers learn within 10 to 15 minutes whether their commit is fine.
The jump on the C++ side was even bigger after a second expansion stage. Instead of selecting only at DLL level, the system now checks inside the DLLs, at function level, which C++ function is actually called. If a function has changed, only the test cases that pass through it run. That took the C++ side from around four to four and a half hours down to usually 10 to 15 minutes.
The long full runs don’t disappear entirely, they just become rarer. For major changes to the core, all necessary test cases still run, and then two and a half hours are acceptable again. Especially in released branches, whose bug fixes go out to customers shortly afterwards, you want the safety of a full run.
| Before | After (on Average) | |
|---|---|---|
| Java side | 2.5 hours | 1 hour, often 10-15 minutes |
| C++ side (function level) | 4-4.5 hours | 10-15 minutes |
| Test cases saved | 50-60 percent |
Trust Comes from Validation, Not from Promises
For a selection procedure to be accepted, it has to prove that it doesn’t miss real failures. For four to six weeks the full suite ran in parallel, while the team logged what the selection would have picked.
The result: not a single missed failure. The few failures that only the full run found were all caused by infrastructure, such as a database that was briefly unreachable because of network problems. Glitches like that don’t count as missed test failures.
This validation is what the trust in the selection rests on. Without a solid comparison against the full run, there would always be doubt about whether the system filters out the right test cases.
Faster Feedback Changes How Developers Work
Small commits pay off once the selection is in place. If you check in little at a time, you change less code and trigger fewer tests. Feedback comes back faster. If you collect changes for days and dump them all in at once, you start a big suite again.
The most noticeable reaction was that the complaints stopped. There used to be a lot of criticism of the slow pipeline that didn’t work. Those complaints are gone, even if effusive praise rarely arrives directly.
More remarkable is how the mood in the project turned. Before the start, some people doubted the effort would pay off, because earlier attempts to analyze the source code in more detail had failed due to its size and complexity.
“When we had finished the Java part, there were suddenly a lot of appreciative voices. I would never have thought we’d get so much out of this project.”
(Silke Reimer)
How to Prioritize Selected Test Cases in Long Runs
The biggest remaining lever is prioritizing the long runs that are left. Today the system runs all selected test cases in no particular order. For long-running parts, you could control which test cases go first, so that likely failures show up early.
There are two ways to prioritize. With test coverage, you pick each additional test case so that it adds as much new coverage as possible. Alternatively, you use historical data and run first what failed most often.
The Jenkins pipeline has to play along technically. The run shouldn’t stop at the first failure, because further failures are still of interest. At the same time, developers should get an early warning: a preliminary message right away, the full report later.
The Java side is also due for the same refinement as C++. Right now the system selects at the level of Java classes there. The next step would be to go down to individual methods and run only the test cases that hit the changed method.
The Scale: Millions of Lines of Code, a Nightly Index
The procedure holds up on a large code base. The C++ part has almost 10 million lines of code, the Java part around 4 million. The belief that this much code couldn’t be analyzed in any meaningful way was one reason for the early doubts about the project.
The nightly index run, which rebuilds the links between files and test cases, takes about five to six hours. IVU has since cut the number of pipeline machines from ten to five, because few check-ins happen at night and that free capacity can go to the index runs.
With ten branches and five machines, each machine manages roughly one run per night. On average, the index is rebuilt about every two days, sooner when a machine frees up early. The Java part has been in production for about one and a half to two years, the more precise C++ analysis at function level for three to four months.
Frequently Asked Questions
Why Isn’t a Nightly Full Run of the Test Suite Enough?
A nightly full run merely postpones the problem instead of solving it. If a test fails in the morning, the search for the commit that caused it only begins then, and that requires additional analysis time. With around 25,000 test cases and a runtime of six and a half to seven hours, a developer won’t find out until the next day anyway whether their check-in was clean.
Why does regression testing become more expensive the more branches a team maintains in parallel?
Any bug fix may end up passing through four or five branches, and on each one, the check-in triggers the full test run all over again. At IVU, up to ten branches were maintained in parallel. Even ten large machines running continuously weren’t enough for this; some of the test cases had to be deliberately cut to avoid a backlog.
Can test case selection be used when an application combines multiple languages and a database?
Yes, but it must then bridge language boundaries. A Java test case first calls a C++ binary to load test data into the database and then executes its actual actions. Which binaries are required varies by test case; some don’t need any at all. Mocking everything would be too complicated, especially for edge cases.
At what level should a code change be mapped to test cases?
The finer the granularity, the greater the effect. Selecting at the DLL level was not sufficient for the C++ components. It was only by testing at the function level that the runtime was reduced from four to four and a half hours to, typically, 10 to 15 minutes. On the Java side, testing at the class level previously made the biggest difference; testing at the method level is still pending there.
Do configuration and data files also count as dependencies of a test case?
Yes. The system logs which artifacts a test case opens, and these include not only Java classes from JARs and loaded C++ libraries but also XML, CSV, or YAML files, for example for a test oracle or for settings. Such files influence test execution just as much as source code and must therefore be included in the mapping.
How can you verify that a test selection does not overlook any real defects?
Through parallel operation. The complete suite continued to run for four to six weeks while logging which test cases the selection would have chosen. Not a single real failure went undetected. The few discrepancies were due to infrastructure issues, such as a database that was temporarily unreachable due to network problems.
How much effort is involved in setting up the mapping between files and test cases?
The index run executes all test cases and logs the files that were opened; it takes about five to six hours. It runs at night when few commits are made. With ten branches and five machines, each machine manages roughly one run per night, meaning the index is, on average, about two days old.
What are the benefits of prioritizing test cases in addition to selection?
It helps in cases where runs remain long despite selection. If likely failures are executed first, the developer can see sooner whether something is broken. There are two approaches: selection based on additional test coverage or based on historical data regarding tests that frequently fail. The pipeline must not abort at the first error.


