Skip to main content

Search...

Test Selection with Teamscale: A CI Pipeline in Ten Minutes

Test selection with Teamscale runs only tests that cover changed code, so 6,000 PyTest cases fit a five-minute budget in a ten-minute CI run.

• • Updated: • 11 min read
Cover of the expert talk on 'Test Selection with Teamscale: A CI Pipeline in Ten Minutes' with Lars Kempe and Richard Seidl.

Test selection in a CI pipeline, often called test impact analysis, means running only the tests that actually cover the changed source code on every commit. Teamscale analyzes code coverage at the level of individual tests and prioritizes the selection so that test execution stays under five minutes. The target is a total pipeline runtime of no more than ten minutes with around 6,000 test cases.

Key Takeaways

  • The CI pipeline at Dolby has a hard limit of 10 minutes, with at most 5 minutes for test execution and the rest for the build, the environment and unit tests.
  • Teamscale automatically ranks the returned tests by proximity to the changed source code, failure frequency and recency, so the most relevant tests run first.
  • PyTest parameterization quickly turns a single test function into 50 to 100 test cases, which is why not all 6,000 test cases fit into CI and a time-based selection becomes necessary.
  • Integrating PyTest with Teamscale took two to three weeks, because there was no native Python support at first and the team had to write its own plugin to capture code coverage per test.
  • The code coverage upload runs every night, so Teamscale is never more than a day behind and can be updated manually for individual branches when needed.

Why Fast CI Pipelines Break Down as Test Suites Grow

A CI pipeline is supposed to give feedback in minutes, not hours. Test impact analysis, selecting only the tests that touch the changed code, is how one team keeps it that way. Without it, feedback slows down as soon as the number of test cases grows. Many teams know the end state: 6,000 tests, a run that takes all night, and during the day nothing but the hope that the important tests are already included.

Lars Kempe, QA Lead at Dolby, explains the core of the problem using his own team’s test setup. The team automates 100 percent of its tests with PyTest and builds libraries that ship to customers. The full suite of around 6,000 test cases runs overnight on Linux and takes about an hour and a half. That’s too long for CI.

The real driver of the test count is parameterization. A single test function turns into many test cases across different configurations. For audio, the range goes from mono through stereo and 5.1 to Dolby Atmos with additional height channels. Combine that with different bit rates and frame rates, and one test quickly becomes 50 to 100 test cases.

Manual Test Selection Goes Stale When Nobody Maintains It

If you pick tests by hand with markers, you have to keep updating that selection. That’s the uncomfortable truth behind every fast pipeline that tries to get by without tooling.

In practice, the upkeep falls by the wayside. Someone would have to sit down every two weeks and adjust the markers to the newly added code. Instead, a base selection keeps running, and nobody knows for sure whether it really covers the latest changes.

That creates a false sense of security. The pipeline is green because “at least the important stuff” runs. Whether the freshly changed code is covered remains an open question. This is exactly where the team decided against manual upkeep and in favor of tool-supported selection.

How Test Impact Analysis Works with Teamscale

The approach depends on one piece of information: which test covers which source code? Teamscale needs that link to pick the right tests for a code change.

The basics are quick to set up. The Git integration takes about ten minutes. After that, Teamscale knows the branches, the commits and the code base. The hard part is linking tests to the code they cover.

That link comes from code coverage at the level of each individual test. The team instruments the binaries with GCOV and measures coverage. The problem with overall coverage is that it doesn’t show which individual test touched which lines. Doing it by hand, you would have to run each test on its own, save the coverage, reset everything and start the next test.

A PyTest Plugin as the Bridge to Coverage

The solution was a custom PyTest plugin that hooks into PyTest. A hook fires after every single test execution, not just after every test function. At that point, PyTest already provides the test name, the status (pass, fail, skip) and the runtime.

What’s missing is the covered code. The plugin parses the GCOV log files and extracts the covered files along with line numbers. That data goes into a log file in the format Teamscale expects and is uploaded.

This fine-grained integration took two to three weeks. Teamscale was strongly focused on Java at first and didn’t support the PyTest environment directly. So the team connected it through the API, working closely with the vendor.

Coverage Data Updates Nightly, Code Updates Continuously

Source code and new tests reach Teamscale automatically, while the coverage data is refreshed once a night. That split is a deliberate choice with a clear consequence.

Thanks to the Git integration, Teamscale always knows which code and which tests exist. The test-to-code mapping, however, is only built during the nightly run. That creates a lag of at most one day: a new test written during the day alongside the code only becomes part of the selection the next day.

For longer-lived or larger branches, the update can be triggered manually. An update on the branch is enough, and the mapping is correct again. In daily work the one-day lag isn’t a problem, because the key mechanism is automatic: when someone changes code, the matching tests are selected the next day.

Time Is the Hard Selection Criterion, Not Test Count

The pipeline doesn’t select tests by number but by the time available. The limit is five minutes of pure test execution, within a total CI run of ten minutes.

For a code change, Teamscale returns the relevant tests already prioritized. The ranking considers how close a test is to the changed code, how often it has failed before and how new it is. Because of parameterization, a single change can still return 800 tests.

Nobody wants to run all 800. At first, the team added up the runtimes of the prioritized tests and cut the list off at five minutes. Later, at the team’s request, an API function was added that takes the available time as a parameter. Teamscale then returns exactly the tests that fit that time window.

The time limit prevents a slide back to where things started. Without it, a large change could, in the worst case, pull thousands of tests back into the pipeline. The full suite still runs every night, so broad regression testing doesn’t get lost.

What the Pipeline Run Looks Like

Build and test preparation run in parallel, which saves precious minutes. While the build is running, the test environment is installed and the tests are requested from Teamscale.

The request itself is an API call based on the commit ID. One small hurdle: in the API version the team uses, Teamscale doesn’t expect the commit ID itself but its timestamp. A one-line Git command handles the conversion.

Teamscale responds with a JSON file that converts directly into a Python dictionary and includes the test runtimes. From the selected test cases, the team generates a text file with a PyTest plugin and passes it in on the command line. PyTest then runs exactly those tests, parallelized through the appropriate plugins.

Once the build finishes, the tests run in five minutes at most. Including the surrounding tooling, the entire CI run takes about ten minutes. If it fails, the developer who triggered it sees the status in GitLab and gets an email. With a team of five or six people, the cause is usually found quickly. Otherwise, they look into it together.

A Small Team Can Shape Its Tools Instead of Just Using Them

The integration came together in just a few weeks partly because the team is small. When one person owns both architecture and implementation, things move faster.

Lars describes his dual role as a QA Lead who also writes code and does the QA work. Large organizations often split test architecture and implementation into separate roles, which has its own advantages. For a quick integration, being close to the system helped: sometimes two hours of pair programming were enough.

The open toolchain helped as well. PyTest and Python come with lots of plugins, and there was already a ready-made way to generate the test list from a text file. Where Teamscale didn’t cover the PyTest world, the team filled the gaps through the API, sometimes with new features the vendor added.

“If I do this, I want to do it properly. We start with a fast CI from the very beginning, even if we don’t have that many tests or that much code yet.”

(Lars Kempe)

What’s Next on the List

The CI setup is considered solved and has been running reliably for about a year and a half without major changes. The next lever is merge requests, which are supposed to run more tests.

At the same time, the team wants to roll out Teamscale to other projects that don’t use it yet. The goal there is the same as in the original project: optimize test selection and keep feedback fast.

Test data isn’t a bottleneck in this environment. The input files are available and the tests fetch them live; about 80 percent of them are audio. What you do have to watch is length: decoding a four-minute PCM file takes about four minutes. With a five-minute budget, every runtime like that counts.

Frequently Asked Questions

Why do test suites grow so rapidly with parameterized tests?

Parameterization quickly generates 50 to 100 test cases from a single test function. For audio, the range extends from mono to stereo and 5.1 all the way to Dolby Atmos with height channels, combined with various bit rates and frame rates. This is how the team described ended up with around 6,000 test cases. Running all of them overnight on Linux took about an hour and a half, too long for a CI pipeline.

What are the drawbacks of a manually maintained test selection using markers?

It becomes outdated because maintenance gets neglected in day-to-day operations. You’d have to sit down about every two weeks and update the markers to reflect newly added code. Instead, a base selection defined once continues to run. The result is a false sense of security: the pipeline is green because at least the most important tests ran, but it remains unclear whether the recently changed code was included.

What information does a tool need to selectively choose tests for a commit?

It needs to know which individual test covers which source code. Overall coverage isn’t enough, because it doesn’t show which lines a specific test has touched. In the project described, the binaries were instrumented with GCOV. Without tool support, one would have to run each test individually, record the coverage, reset it, and start the next one.

How time-consuming is it to integrate test-specific coverage tracking?

The Git integration itself took about ten minutes; linking the tests to the covered code took two to three weeks. In 2023, Teamscale was heavily focused on Java and did not directly support the PyTest environment. The team therefore wrote its own plugin and integrated it via the API, working closely with the provider.

How up-to-date does the coverage data need to be for automatic test selection?

A one-day delay is manageable. Source code and new tests were continuously delivered to the tool via Git integration; however, the mapping of tests to code only occurred during the nightly upload. A test created during the day alongside the code is therefore not taken into account until the following day. For longer or larger branches, the update could be triggered manually.

Should test selection in CI be limited by number or by time?

By time. At Dolby, the limit was five minutes of pure test execution, embedded within a total CI run of ten minutes. A single code change can still trigger 800 tests through parameterization. The tool selection was returned as a prioritized list, weighted by proximity to the modified code, previous error frequency, and the age of the test; the list was truncated based on the time budget.

What role does test data play in the runtime of a fast pipeline?

Acquiring the data is rarely the bottleneck; the length of the data is. The input files were available and were retrieved live by the test; in about 80 percent of cases, they were audio files. Decoding a four-minute PCM file also takes about four minutes. With a budget of five minutes, a single test of this kind fills the window almost completely.

Share this page

Related Posts