Branch coverage is a form of code coverage that requires every possible branch of a condition to be tested. A single if statement with an “and” condition therefore needs at least three test cases. Reaching 100 percent branch coverage prevents many bugs, but it doesn’t guarantee bug-free code: loops that run several times can still hide defects.
Key Takeaways
- 100 percent branch coverage doesn’t protect against every bug: an off-by-one error in a copied method survived full coverage because the loop was only tested with zero and one iteration, never with several.
- Roger Butenuth found a security hole in an authentication check, where one variable was compared with itself instead of with the stored value, only because branch coverage forces all three branches of a compound condition, not just the line.
- Testability comes from changing the code: dependency injection and clear interfaces made it possible to trigger I/O errors and capture standard output in tests without running into hardware limits.
- Applying “Don’t Repeat Yourself” consistently, driven by the test effort for copied two-line checks, makes code easier to read and maintain because changes happen in one place only.
- A coverage number as a target doesn’t work, because developers can hit any measurable threshold without writing meaningful assertions, and then the number means nothing.
100 Percent Branch Coverage: A Self-Experiment with a Clear Lesson
Full branch coverage is achievable, but it won’t find every bug. Roger Butenuth covered every single branch of a self-built interpreter with tests and still found bugs long after the metric had reached 100 percent.
Branch coverage goes further than the usual code coverage metrics. It doesn’t count methods or lines. It counts every possible branch in the control flow. An if statement with an “and” condition therefore needs three test cases instead of one, so that every combination of the branch gets exercised. That is the main difference from line coverage, which is satisfied as soon as the line has run once.
This rigor is what sets the experiment apart from everyday practice. Many teams agree on 70 or 80 percent coverage, usually measured at line or statement level. Hardly anyone takes branch coverage this far, because the effort is high.
The Project: A Lisp-Like Interpreter in Java
The starting point wasn’t a testing project at all. It was an interpreter for a Lisp-like language, written in Java. Lisp stands for List Processing, and lists are at the heart of it: immutable, or persistent. When you call an operation, you don’t change the existing list. You get a new list back that looks like the old one plus the change.
The naive way would be to copy the whole list every time. Nobody wants that. What you need is an implementation that reuses as much as possible, and the existing options only partly fit.
- Java’s
ArrayListisn’t immutable. - Scala’s singly linked list is slow at reaching the nth element, because you have to walk through it n times.
- Scala’s index tree was efficient, but even a list with a single element took more than 200 bytes, too much for lots of small lists.
So Roger built his own structure: a two-dimensional array with exponentially growing subarrays, with logarithmic cost for most operations. The price was a complicated implementation with a high risk of off-by-one errors. That’s exactly why the list got tested so thoroughly, and it soon reached 100 percent coverage. From there grew the ambition to extend full branch coverage to the whole interpreter.
Pure Functions Are Easy, Side Effects Are the Real Work
The easy part is functions without side effects. Arithmetic and other pure functions, in the functional programming sense, behave like abstract data types. They have no outside dependencies, and their tests stay manageable.
It gets hard with anything that touches the outside world. File reads and writes need prepared test data. The built-in HTTP client and HTTP server each need a counterpart to talk to. It helped that the language shipped both client and server anyway, both built on standard Java libraries.
Completeness is the most expensive part. If you want 100 percent, you also have to test the exception handler of a file operation. And nobody wants to fill up the hard drive just to run a test.
Reaching Full Branch Coverage in Java Often Means Changing the Code
The way to reach these awkward branches was dependency injection through interfaces. The language’s scanner is simply handed a Java Reader. In a test, you can pass in a Reader that throws an IOException at a defined point. That lets you hit the exception path on purpose without provoking real failures.
Standard output works the same way. The interpreter normally writes to System.out, but the PrintStream behind it can be swapped out. In a JUnit test, a custom PrintStream captures what the interpreter prints and makes the output checkable.
In several places, the production code itself was changed to make this possible. That’s the uncomfortable truth behind high testability: sometimes it’s better to rework the code so you can test it than to hold on to a structure that’s hard to test.
In a real project, that’s harder to push through. Here there was no outside product owner; the author was his own. In practice, someone will ask why code should change just for the sake of testability. The answer stays the same: testable code is often better code.
Testable Code Often Becomes More Readable and Maintainable
The rework brought more than green metrics. In one spot, a method was added to the list functions that prints the list’s internal structure. Nobody needs it in production. In tests, it checks whether the list has the expected internal layout.
Dependency injection makes code more flexible, even without a framework like Spring Boot behind it. The second lever was applying “Don’t Repeat Yourself” consistently. Small checks like “if the list is empty, throw an exception” tend to get copied into many functions as two-liners. Every one of those copies needs its own test cases.
If you pull these mini-checks out into their own functions, the number of test cases drops, and readability often improves too. Changes then happen in one place only. The flip side: whoever changes that one place can break many callers at once. High test coverage catches exactly that, which keeps even larger refactorings under control.
In the end, the ratio of production code to test code was roughly one to one.
Full Coverage Catches the Trivial, Not the Overlooked
Full branch coverage finds bugs where you least expect them. The worst find was a security hole in the Basic Authentication of a web service. The method was supposed to check the username and password passed in against the known values. On one side, the this was missing before user and password, so the user passed in was compared with itself. Anyone would have got through.
That line had already been covered by a first test. Only the requirement to hit every branch of the nested if forced the two missing test cases, and that’s when the bug showed up. Without branch coverage, it would have stayed hidden.
Even so, 100 percent doesn’t catch everything. During Advent of Code, a series of daily programming puzzles in the run-up to Christmas, another bug turned up in the list implementation. It sat in a while loop that had been tested with zero iterations and with one. The bug was an index that was shifted late in one iteration and only used in the next.
The cause was a classic copy-and-paste problem. The code existed twice, once for changes at the front of the list and once for the back. In the copy, there was a plus where a minus belonged, so the index ran in the wrong direction. Branch coverage structurally can’t catch errors like this that span loop iterations.
“You make mistakes in exactly the places where you think: this is too trivial, there can’t be a bug here.”
(Roger Butenuth)
The remaining bugs were few, but they were there. So the takeaway is humility rather than the promise of bug-free code. Where you used to just write code down, it’s worth giving it a second thought.
Why a Coverage Target Alone Doesn’t Drive Quality
A fixed percentage target doesn’t reliably steer quality. Any metric you set can be gamed, and that’s exactly what happens. A number as a goal produces test cases that meet the number, not necessarily test cases that find bugs.
Coverage is also only as good as the assertions behind it. There is code with decent coverage and pointless checks. High coverage only helps you with sensible assertions; otherwise it just measures that the code ran.
Choosing a coverage level depends on context. What is the software for? How critical is it? How expensive is a failure in production? The answers decide how far you should go.
| Coverage type | What it checks | Effort |
|---|---|---|
| Line / statement | Whether a line was executed | Low |
| Branch | Whether every branch was taken | High |
Exception handlers show where the limit lies. They are hard to cover. But if a codebase is littered with scattered handlers, the better question isn’t how to test them all, but whether the structure should look like that in the first place.
In real projects, 100 percent isn’t reachable, and that’s fine. Instead of relying on one metric, it makes more sense to trust developers’ judgment. Where the debate matters also shifts with the test level: in integration projects, the problems don’t sit in unit tests but between systems, where meaningful coverage looks completely different.
Frequently Asked Questions
How does branch coverage differ from standard line coverage?
Branch coverage counts not the number of executed lines, but every branch in the control flow. An if statement with an “and” condition therefore requires three test cases instead of one to ensure that every possible combination of the branch is covered. The effort involved is significantly greater than with line or statement coverage. However, this helps uncover errors that are hidden by a line that has already been covered, such as an incorrect comparison in a compound condition.
What coverage targets do teams typically set?
It is common to agree on 70 or 80 percent, usually measured at the line or statement level. Hardly anyone consistently pursues branch coverage with this level of rigor because the effort involved is high. In the interpreter project described, the ambition to achieve full coverage arose only when dealing with a particularly error-prone data structure: a two-dimensional array with exponentially growing subarrays.
Which parts of the code require the most work when aiming for high test coverage?
Everything that interacts with the outside world. Pure functions, such as arithmetic, have no dependencies and remain manageable in testing. File I/O requires prepared test data; an HTTP client needs a server, and vice versa. Completeness is the most costly aspect: Even the exception handler for a file operation must be run through without a test filling up the hard drive.
Is it justifiable to refactor production code solely for the sake of testability?
Yes, because testable code is often better code. Dependency injection using interfaces allows you to target hard-to-reach branches specifically: The scanner is provided with a reader that throws an IOException at a defined point in the test, and the PrintStream behind System.out is replaced in the test to make output verifiable. In real-world projects, this refactoring is politically more difficult to implement.
Why does duplicated code increase the testing effort?
Every duplicate creates its own branches and thus its own test cases. Small checks like “if list is empty, throw an exception” often end up as two-line blocks in many functions. If you move them into their own function, the number of test cases decreases, and changes occur in only one place. The downside: A bug there affects many callers, which in turn is mitigated by high code coverage.
What errors does even complete branch coverage overlook?
Errors that span loop boundaries. In the list implementation, an off-by-one error survived full coverage because a while loop was tested only with zero and a single iteration. The misaligned index wasn’t used until the next iteration. The cause was copy-paste: In the copied version for the end of the list, a plus sign was used instead of a minus sign. The metric does not structurally cover such cases.
Does a fixed coverage target in a project lead to better quality?
No. Any measurable threshold can be met without writing meaningful assertions, and that is exactly what happens. A numerical target generates test cases that meet the target, not test cases that find bugs. It makes more sense to consider the context: how critical the software is and what a bug costs during operation.


