Skip to main content

Search...

Is Mutation Testing Worth the Runtime? Where It Pays Off

Every mutant reruns your test suite, and that gets slow fast. How incremental analysis in PIT and a focus on core logic make mutation testing pay off.

• • Updated: • 10 min read
Cover of the expert talk on 'Is Mutation Testing Worth the Runtime? Where It Pays Off' with Birgit Kratz and Richard Seidl.

Mutation testing checks the quality of a test suite, not the production code. A tool changes the code on purpose with so-called mutants, for example by turning a greater-than into a greater-than-or-equal, and then checks whether at least one test fails. If no test fails, the mutant survives, which points to a gap in the tests.

Key Takeaways

  • Mutation testing evaluates the test suite itself: a deliberately injected change in the code counts as killed as soon as at least one test fails.
  • The Java framework PIT keeps runtimes down with incremental analysis. It compares hash codes of code and tests and reruns mutations only where something has actually changed.
  • Mutation testing pays off on the core business logic. Applying it to the whole codebase, including UI code or database access, drives up the effort for little gain.
  • People who use mutation testing regularly end up writing better tests and better code, because they already anticipate the typical mutants while writing.

Mutation Testing: Testing Your Tests

Mutation testing is a way to test your tests. It doesn’t ask whether the production code works. It asks whether your test suite is able to find errors at all.

The principle is more than 50 years old. Richard Lipton described it in a paper in 1971, and the basic idea has barely changed since.

The process is simple. You have a green test suite and you assume everything is fine. Now you deliberately break something in your code and see whether the suite notices. If it does, the suite is at least good enough for that spot. If it doesn’t, you’ve found a gap.

These deliberate changes aren’t called bugs. They’re called mutants, hence the name mutation testing. A mutant is a targeted change to the production code.

Is mutation testing worth the runtime? It is if you keep it in check: run it incrementally, point it at the core business logic and start with a few mutators. The payoff is better tests and, over time, better code.

Which Mutants Mutation Testing Uses

The kinds of mutants depend partly on the programming language. In the Java world you can tell several categories apart, and some of them work regardless of language.

Conditional boundaries are a common category. A > becomes >=, a < becomes <=, and the boundaries of a condition shift. Other mutators negate entire conditions, for example by turning == into !=.

Increment mutators swap ++ for --. Arithmetic mutators replace addition with subtraction or multiplication with division.

You can also change how entire methods behave. A void method that does something but returns nothing simply isn’t called by the mutator. Methods with a return value return null instead of an object, and for primitive types they return zero or an empty string.

There’s hardly any limit to what you could come up with. Applying every conceivable mutator to an entire codebase by hand would keep you busy for a very long time, which is exactly why frameworks do this work.

How a Mutation Testing Tool like PIT Works

A mutation testing tool creates a mutant, runs all the tests and checks whether at least one of them fails. A failing test is good news here: the test suite has killed the mutant.

The vocabulary is martial. A mutant either survives or gets killed, and killing it is the result you want.

In the Java world, PIT is a widely used framework. You plug it into your build, it applies the mutators to the source code and then logs how the code behaves. You evaluate that log afterwards.

PIT ships with preselected sets of mutators at different levels, from a basic set up to a level that applies every mutator it comes with. You can switch individual mutators on and off, because not every mutator makes sense for your code, and more mutators mean longer runtimes.

Mutation Tests Run Long Unless You Limit Them

Mutation tests can run for a very long time. For every mutant, the entire relevant test suite is executed. Let that loose on a whole codebase without any control and you’ll block your build.

PIT is heavily optimized at this point. It keeps a history of hash codes for code and tests. On the next run, it checks which tests are affected by a change and runs only those. Where nothing has changed, the result doesn’t change either.

This incremental approach is what makes integration possible. If every run took hours, you could hardly fit mutation testing into a daily build. When it runs incrementally, you have far more options for adding it to your regular builds.

In the build process, mutation testing always comes after compilation and the normal test run. The mutation analysis only starts once all tests are green.

Start with the Core Business Logic

When you introduce mutation testing, don’t turn it loose on the entire source code right away. Start with the core business logic.

Running mutation testing on UI code or database access is overkill at first. Frameworks like PIT let you limit the analysis to specific packages or individual classes. That way you decide where mutation testing actually adds value, and you avoid runtimes that get out of hand.

The same goes for the number of mutators. Switch on just a few at first and work your way up while you analyze the results closely and adjust tests or code.

What the Results Tell You

Mutation testing uncovers three kinds of weaknesses. It shows whether your test data is any good, for instance whether it covers the boundaries. It shows what you haven’t tested at all yet. And it brings logic problems to light.

If a mutation doesn’t make any test fail, look at both sides: the tests and the code. Sometimes it turns out the tests are green but check the wrong logic or miss the point entirely.

That’s where the real value lies. You improve your tests, and you improve your code along with them.

Equivalent Mutants and Their Pitfalls

Equivalent mutants are a well-known problem in mutation testing. They are mutants that don’t actually change the logic of the code. The code behaves exactly as it did before the mutation, so no test fails.

PIT tries to detect and avoid equivalent mutants in advance, but it can’t rule them out completely. You have to filter out those spots yourself.

You spot them through the basic principle: a mutant was created, yet all tests stay green. Whenever a mutant survives, it might be an equivalent one.

Depending on the code, though, equivalent mutants tend to be rare. Sometimes you can rewrite the code so the problem goes away. Often they also hint that something is off with the code itself: if a different construct produces the same result, something may have gone wrong in the implementation.

The Tool Becomes a Coach for Better Code

The more you use mutation testing, the more its effect shifts. At first it mainly gives you new test ideas. The more often you run it, the stronger the learning effect.

After a while you know the mutants that get applied and account for them while you’re still writing the tests. Birgit Kratz describes how hard it was for her to deliberately write sample code for a talk in which a mutant survives. Once you know the tool well, you write tests and code that let almost nothing slip through, practically without thinking about it.

In a team, the learning effect can be considerably bigger. In Birgit’s view, a joint session where the team goes through the findings of the mutation tests is well worth it.

“It’s good for better tests, and it’s just as good for better code.”

(Birgit Kratz)

In practice, this team approach meets resistance. Many people react to the idea of testing their own tests by asking where it’s supposed to end. That’s why mutation testing often stays with individual developers.

Two Prerequisites: Tests, and Green Tests

Before you can use mutation testing, you need two things: tests, and green tests.

That sounds obvious, but it isn’t. In practice, these are exactly the two things that often fall short. Either there are too few tests to begin with, or the existing tests aren’t green.

Only when both conditions are met does mutation testing make sense, whether you run it locally or as a stage in the build pipeline after the normal test run.

Frequently Asked Questions

Is mutation testing a new approach?

No. Richard Lipton described the method in a paper in 1971, and the basic idea has hardly changed since then. What’s new, above all, is the tool support. The process remains the same: A bug is intentionally introduced into code with a passing test suite, after which it becomes clear whether at least one test flags it.

Can mutation testing be done manually without tools?

Practically speaking, no. In Java alone, there are numerous categories of mutants: shifted comparison thresholds where “greater than” becomes “greater than or equal to,” negated conditions, swapped arithmetic operations, uncalled void methods, or return values that become null or an empty string. Anyone who applies all of this manually to a codebase will be busy for a long time. That’s why frameworks take care of this work.

Why do mutation tests take so long to run?

For every mutant generated, the entire relevant test suite is executed. If unleashed unchecked on the entire codebase, this blocks the build process. PIT limits the overhead by maintaining a history of hash codes for the code and tests: on a subsequent run, only the parts that have changed are mutated and subjected to testing.

Where does mutation testing fit into the build process?

After compilation and the normal test run. The mutation analysis only follows once all tests pass, because passing tests are one of the two prerequisites for the process. Because PIT works incrementally and only reruns affected tests, this step can also be integrated into a daily build rather than being executed locally.

Should mutation testing be applied to the entire codebase?

No. A sensible starting point is the core business logic. Mutating UI code or database accesses is overkill at first and drives up runtime. Frameworks like PIT allow you to limit the analysis to specific packages or individual classes. It’s also worth taking a cautious approach with mutators: enable a few, analyze the results, then expand.

What weaknesses does mutation testing uncover?

Three types. It shows whether the test data is adequate and covers edge cases, for example. It reveals what hasn’t been tested at all so far. And it brings logic problems to light. If a mutation doesn’t cause a test to fail, both sides need to be scrutinized: Sometimes the tests pass, but they’re checking incorrect logic or missing the actual point.

What are equivalent mutants?

Equivalent mutants do not actually change the logic of the code. The program behaves exactly the same after the mutation as it did before, which is why, quite rightly, no test fails. PIT attempts to identify such cases in advance, but they cannot be completely avoided. They are often an indication of a problem with the code itself: If a different construct yields the same result, there may be something wrong with the implementation.

Is mutation testing more beneficial for a team or for individual developers?

The learning effect is greater when done as a team. According to Birgit Kratz, a joint session to review the findings is very beneficial. In practice, however, this approach faces resistance: Many people react to the idea of testing their own tests by asking where it will ever end. As a result, it often remains the domain of individual developers.

Share this page

Related Posts