Skip to main content

Search...

Test Intelligence: Finding More Bugs in Less Time

Test intelligence finds 80% of errors in 1% of the runtime with test impact analysis, and test gap analysis shows untested changes before release.

• • Updated: • 12 min read
Cover of the expert talk on 'Test Intelligence: Finding More Bugs in Less Time' with Elmar Jürgens and Richard Seidl.

Test intelligence is the umbrella term for analysis techniques that mine data from your own development and test process to find more bugs in less time. The concrete methods are test gap analysis (showing untested changes before a release), test impact analysis (running only the tests a change affects) and test suite minimization (calculating the best subset of a large test suite).

Key Takeaways

  • Test gap analysis combines version control data with coverage profiling to show untested changes before a release, so teams can make a deliberate call.
  • Coverage profilers record manual tests just as completely as automated ones, and testers don’t have to change how they work.
  • Test impact analysis finds eighty percent of the bugs the full test suite would find in one percent of the total runtime, which makes feedback much faster.
  • Test suites that have grown over the years test too much and too little at once: many tests overlap heavily while real gaps go unnoticed.
  • Even a well-staffed, experienced development team can’t reliably test all of its changes without dedicated tools, as an internal study at CQSE showed.

What Test Intelligence Means

Test Intelligence covers a family of analyses that turn data your project already produces into better testing decisions. It taps two sources that exist in every project: the version control system with every change to the code, and test coverage as measured by coverage profilers.

Elmar Jürgens groups several concrete analyses under this term, and they complement each other. None of them is an end in itself. Each one is about a better basis for decisions: what needs testing, what is worth it, what can go.

What they share is the combination of a static and a dynamic view. A static analysis reads the code without running it; a coverage profiler records what a test run actually executes. Only when you bring both data sources together do the results become reliable.

How Test Gap Analysis Shows Untested Changes

Test gap analysis tells you, before a release, which code changes haven’t been tested yet. The underlying assumption: in long-lived software, most bugs sit in the places that have changed since the last release. In a complex team, though, it is hard to keep track of whether every change was really caught by a test.

So two data sources are laid on top of each other. The changes come from version control and show what has changed since a reference point, such as the last release or the last test run. The coverage comes from coverage profilers, which record every test phase without gaps.

The key point: this records not only automated unit and end-to-end tests, but manual tests as well. A coverage profiler doesn’t care whether it is watching an automated or a manual test. Testers don’t have to change how they work; the recording runs in the background.

The goal isn’t to test everything. The goal is a conscious decision. Sometimes an untested change is allowed into production because the feature it touches won’t be needed for months. Often that isn’t the case, and then the analysis shows in time where the gap is.

“For whatever reason, coverage profilers have mostly been used for automated tests over the past decades. But we have been doing this with customers for manual tests too for more than ten years, and it works very well.”

(Elmar Jürgens)

Coverage Profilers Work Differently Depending on the Technology

There is a suitable profiler for every common technology, commercial or open source. Broadly, they fall into three categories.

  • Profilers on the virtual machine: used for languages such as Java, C# and Python, and for SAP ABAP. They attach to the VM and watch from there what gets executed.
  • Instrumenting profilers: change the code itself by inserting their own instructions between statements that report when a spot has been executed. This is the route for C and C++, where there is no virtual machine to attach to.
  • Hardware profilers: for embedded systems. Some processors have dedicated pins that a hardware profiler connects to. They give a hardware guarantee that they don’t affect timing and that the processor behaves exactly as it would without profiling.

The choice depends on what works best for the technology at hand. For some languages, dedicated profilers were written because there was no usable alternative.

Why the Performance Impact Stays Small in Practice

Very fine-grained instrumentation can slow the runtime down considerably, but test gap analysis doesn’t need that level of detail. You don’t have to know whether every single path through a long method was taken. It is enough to know whether a method was entered at all.

Recording whether a method was entered takes a single bit. That produces less data and far less performance load. Benchmarks with various profilers often show a slowdown of around one percent, sometimes five percent.

In manual testing you don’t notice it at all. The system is usually waiting for the network or the database anyway, and profiling doesn’t slow anything down at those points.

Even Good Teams Don’t Fully Cover Their Changes

Even disciplined development teams don’t reliably manage to test all of their changes. It isn’t a lack of will. Without dedicated tools, it is simply too complicated.

A study inside the company backs this up. The starting conditions were favorable: a small, stable group of trained computer scientists, plus a code peer review for every change, where a second person checks each change before it may go into the release. The study looked at whether people could correctly cover someone else’s freshly reviewed code with tests. Even under those conditions, they couldn’t.

If gaps remain in a small, well-organized team, they are the rule in large systems that have grown over the years. That is where the analysis steps in, instead of relying on gut feeling.

How Test Impact Analysis Brings Back Fast Feedback

Test impact analysis picks exactly those tests from a large test suite that are affected by your latest changes. That shortens the feedback loop dramatically.

Many teams know the problem behind it: over time, more and more automated tests are added and the total runtime keeps growing. One customer has 80,000 automated end-to-end tests that take around 400 hours when run one after another. Even heavily parallelized, it often takes several days to get results.

Late feedback loses its value. If something breaks today and you hear about it right away, you know what caused it. If the feedback arrives a month later, you have forgotten what you did, and many other changes in between could just as well be to blame.

The solution: each test case is measured once on its own. After a change, you can then say precisely which tests need to run now. The results are clear: in about one percent of the total runtime, the selection finds 80 percent of the bugs the whole suite would find, and in two percent of the time it finds 90 percent.

A small share slips through, so an occasional full test run is still needed. But most new bugs show up very quickly. It plugs into the continuous integration pipeline without rewriting existing tests. That helps teams in particular that don’t have a clean test pyramid and no easy way to get one.

Legacy Test Suites Do Too Much and Too Little at Once

Test suites grow the way the software does, and that leads to a paradox: a lot stays untested, while at the same time many tests check almost the same thing as others. Redundancy often comes from copy and paste with small changes, especially in end-to-end tests.

Pareto optimization tackles this by extracting a small test suite from a large one that finds a big share of the bugs in a fraction of the time. That small suite works well as a quality gate before you start a more expensive test run.

The benefit becomes concrete where testing is expensive. Hardware-in-the-loop tests with machine tools can’t be parallelized at will, because the expensive machine has to be physically there. The software has to be good enough first before it is allowed onto that run. Otherwise a single bug that makes 1,000 of 5,000 tests fail hides all the others.

Instead of putting an acceptance suite together by hand, the optimization calculates the suite that finds the most bugs within a given time window. In studies with customers, the calculated set found twice as many bugs in the same time as the hand-picked selection.

It pays off in resources too. If you run tests in the cloud, on AWS for example, you pay for every execution, and the bill grows with every new test and every extra run. A smaller, targeted suite cuts those costs directly.

Acceptance Decides, Not the Tool

An analysis tool only has an effect if the team actually uses it, and that takes change management. This applies equally to static, dynamic and hybrid analyses, the latter combining the static and the dynamic view.

Measuring alone doesn’t improve anything. Every team needs a feedback loop where it is clear what the tool can and can’t do. That includes raising concerns about personal performance monitoring proactively, before anything is measured at all. That fear often hangs over analyses.

Integration into daily work matters just as much. The information should show up in the tools people already work in: the test management tool for testers, the IDE and the pull request for developers. A tool that is thrown over the fence falls flat.

Delivering tool, rollout and support from a single source also cuts out the game of telephone between users and the development team. When the support people help build the tool and use it themselves, feedback comes back more directly, and they notice quality problems in their own product first.

Where Test Intelligence Is Heading

The same technology that measures in testing can also run in production to show what users actually use. In software that has grown over the years, this often uncovers functionality nobody has needed in years.

Deleting that dead code saves twice: you save effort in testing and in development. If a static analysis finds a security vulnerability in an area nobody needs anymore, deleting it is cheaper than a painstaking fix.

More data sources are on the roadmap. Requirement and test smell analyses read natural-language requirements the way a static analysis reads code, looking for passive constructions or vague words, for example. Wording like that leads to underspecified test cases that different people interpret differently, and so to non-deterministic tests. An ambiguous requirement leaves open how it will be implemented in the end.

Another field is requirements tracing. A tool that sits in all the systems anyway can generate verification matrices from requirements, manual and automated test cases and code. Every pull request shows when a requirement has changed and what that means for code and tests. The matrix is no longer built after the fact for an auditor; the artifacts stay consistent all along.

Frequently Asked Questions

What data do you need to find untested changes before a release?

Two sources are sufficient, and they’re generated in every project anyway: the version control system and measured test coverage. The changes made since a reference point (such as the last release or the last test run) are overlaid on the recorded coverage. The underlying assumption is that, in long-lived software, most defects are found where code has been modified.

Can the coverage of manual tests even be measured?

Yes. A coverage profiler doesn’t care whether it’s observing an automated or a manual test; the recording runs in the background, and testers don’t have to change their workflow. In 2023, Elmar Jürgens reported on over ten years of practical experience with clients in which profilers were also used for manual testing phases.

Does coverage profiling noticeably slow down test execution?

Generally, hardly at all. Very fine-grained instrumentation does increase runtime, but for test gap analysis, it’s sufficient to know whether a method was entered at all. A single bit is enough for that. Benchmarks often showed a slowdown of around one percent, sometimes five percent. In manual testing, this isn’t noticeable because the system is usually waiting for the network or database.

Is a comprehensive code review sufficient to ensure that all changes are tested?

No. An internal study at CQSE examined precisely this favorable starting point: a small, stable group of trained computer scientists and a peer review for every change before release. The study examined whether code written by others, which had just been reviewed, was correctly covered by tests. Even under these conditions, this was not achieved with high reliability.

What is the benefit of running only a portion of the tests after a change?

Time, above all. If each test case is measured individually, the affected subset can be executed specifically after a change. In about one percent of the total runtime, this subset found 80 percent of the defects that the complete suite would have found; in two percent of the time, it found 90 percent. Occasional full runs remain necessary.

How can you tell if a large test suite has become redundant?

Redundancy typically arises from copy-and-paste with minor changes, especially in end-to-end testing: Many cases check almost the same thing, while real gaps go unnoticed. A Pareto optimization calculates the suite with the highest defect yield for a given time window. In studies with customers, the calculated set found twice as many defects as a manually curated selection.

Why do analysis tools fizzle out in teams, even though the technology works?

Because acceptance determines the impact, and measurement alone doesn’t improve anything. Change management is essential: clarify what the tool can and cannot do, and address concerns about personal performance monitoring before any measurements are taken. This includes integrating the tool into existing workflows, such as test management tools for testers, IDEs, and pull requests for developers.

What does coverage measurement in production have to do with testing effort?

It shows which features users actually use. In software that has evolved over time, functionality often appears that no one has needed for years. Deleting such code saves effort in two ways: in testing and in development. If a static analysis finds a security vulnerability in an unused section, deleting it is cheaper than fixing it.

Share this page