Skip to main content

Search...

Continuous Everything: How to Measure the Quality Gain

Does Continuous Everything really improve quality? Six years of data show trunk-based development smooths out the defect peaks around each release.

• • Updated: • 13 min read
Cover of the expert talk on 'Continuous Everything: How to Measure the Quality Gain' with Marco Achtziger, Gregor Endler and Richard Seidl.

Continuous Everything means improving software quality in a measurable way through automation, continuous integration and early defect detection. Two metrics decide whether it works: how quickly a change reaches the customer and how many defect reports come in. Both go down demonstrably when teams build autonomy and cultural change step by step instead of following central mandates.

Key Takeaways

  • Trunk-based development measurably flattens the defect spikes around releases: instead of large batches of defects at release time, teams see a steady, lower defect level during ongoing operation.
  • Cultural change in software development takes years, not quarters: even with targeted craftsmanship programs, it takes two to three years for new ways of working to take hold in teams.
  • Autonomy beats mandates: teams that decide for themselves how many open defects they tolerate accept quality rules far more readily than teams forced into a zero-defect policy.
  • Two metrics are enough to measure software quality meaningfully: lead time to the customer and the number of incoming defect reports reliably show whether a process change is working.
  • A value stream analysis before any process optimization keeps you from starting in the wrong place: what feels like the problem rarely is, and the lost time often sits in waiting for other teams.

Continuous Everything Only Pays Off If You Measure It

Sooner or later, every team that rolls out continuous integration, delivery and deployment hears the same question from management: is this actually worth it? Thousands of automated builds, constant test runs, continuous releases. All of it costs time and money long before any effect shows up.

Continuous Everything means automating as many steps of software development as possible and turning them into a continuous flow. Conferences are packed with tool vendors and experience reports on how to build that technically. The harder question usually goes unanswered: can you prove the quality gain in numbers?

Marco Achtziger takes a clear stance on this. If you claim a change makes the software better, the burden of proof is on you. The good news is that the data already exists. Source control systems, test execution logs and ticket systems all produce timestamped records that you can analyze over years without collecting anything extra.

Why Technical Change Is Mostly a Culture Problem

The technology is the smallest part of the problem. The real obstacle is habit. Developers who have spent years working with a heavyweight version control system won’t change how they work just because a new tool is available.

Moving to Continuous Everything is therefore a change of mindset first and a change of tools second. Marco Achtziger builds on Dan Pink’s model of motivation with its three levers: mastery, purpose and autonomy. Address all three, and people start taking part on their own initiative.

In his case, that happened through craftsmanship programs. Taking part was entirely voluntary: autonomy. The programs were organized in levels that made skill growth visible: mastery. And every measure came with an explanation of why it mattered: purpose.

Making Voluntary Participation Work When Not Everyone Joins

There is no silver bullet. Every change process leaves the last ten percent behind, the people who simply won’t be convinced. Marco Achtziger openly admits that some colleagues still think the old version control system was better. The goal is to reach most people, not every last one.

The lever that works is the opinion leaders. Every developer community has a few influential people others listen to. Win them over, make them pilots, and they become multipliers. They talk about what they’re doing, others get curious, and curiosity turns into participation.

This takes time. It took two to three years for the craftsmanship program to become part of the culture. By the end, almost every team was involved, and about half of them actively. Anyone expecting cultural change to happen overnight will be disappointed.

How Levels Target Real Pain Points

The level structure of a craftsmanship program does two jobs at once. It defines the skill set expected of developers and testers, and it picks up the pain points that are already hurting.

Continuous integration is a good example. Setting it up went quickly, but most builds were red and nobody cared. A familiar story. The levels went after exactly that, one step at a time:

LevelBuild Requirement
BasicAt least one test runs in the build
IntermediateThe build runs the tests but doesn’t have to be green yet
AdvancedThe build is green most of the time, and the team shows how it gets there

Flaky tests were a pain point of their own: tests that behave differently depending on the phase of the moon or the day of the week. An unstable test isn’t a disaster in itself. Ignoring it is. The key measure was a quarantine build. If a flaky test can’t be made green within a few hours, it moves out of the normal run and into quarantine. That keeps the regular build green.

Metrics Teams Choose Themselves Work Better Than Imposed Ones

At many levels, the programs didn’t prescribe a solution. They asked the teams to come up with their own. One level, for example, required metrics the teams would use to improve their code. The first question back was predictable: which metrics should we use?

The answer: you tell us. Think about which metric actually helps you. Many teams ended up with established measures such as code coverage or McCabe complexity, but they could explain why. That was the point. A team that understands what a metric is for uses it sensibly instead of chasing a number handed down by a central department.

Flaky Tests Don’t Disappear, You Learn to Handle Them

A stable build isn’t a goal in itself, and a build that is one hundred percent green is actually suspicious. Marco Achtziger sums it up:

“If a build is one hundred percent green, congratulations, you have stable tests. The bad news: apparently nobody is working anymore.”

(Marco Achtziger)

Where people work, mistakes happen, and that leaves a certain baseline of instability in any system. Trying to stabilize every single test at any cost gets you nowhere. What matters is not to ignore instability and to find a way of dealing with it, so that most builds stay green.

You can see this learning process directly in the data. Shortly after the craftsmanship programs started, the share of flaky tests went up, because more people were writing tests. Then, as awareness of how to handle unstable tests grew, the number dropped again and settled at a low level.

Two Metrics Are Enough to Assess Software Quality

When it comes to judging software quality, only two metrics really tell you something: how long it takes for a change to reach the customer, and how often the customer complains about the software.

In other words: lead time and customer feedback, usually in the form of defects or tickets. Both can be reconstructed from data you already have and tracked over years. The effect of a change doesn’t show up right away. It shows up across several releases.

The effect on lead time was clear. After moving from a branch-based approach to trunk-based development, the time it took for a change to reach the customer dropped from days to hours.

How Trunk-Based Development Flattens the Defect Curve

Trunk-based development changes more than speed. It changes the shape of the defect curve. In the branch-based setup, customers were served from different version branches. A defect had to be fixed on several branches, and figuring out at which branch point the problem had crept in was often guesswork.

In that old model, defects came in spikes, each one around a release. For developers, that meant recurring surges of work that seriously disrupted the normal flow. With trunk-based development, where all customers have been served from the same mainline for two years now, the spiky curve turned into a steady, flat line at a low level.

Put the two curves on top of each other and you see more than the smoother shape. The overall level is lower too. The continuous load in the trunk-based model sits below what was left over at the tail end of the big release defect batches.

Why Linkable Data Comes First

The most important lesson from the data work sounds trivial: the data has to be captured and it has to be linkable. Often a small tweak to the tooling is enough to bring separate data sources together.

One concrete example shows how this works. A database recorded which tests had run in which build. But to link that to changes in the source control system, it lacked a field for the version number. That small gap was fixed, and from then on the analysis could even be reconstructed for the past.

Only that link made it possible to detect flaky tests. A build as a whole is coarse: red or green. Linked down to the level of individual commits and test cases, though, you can check exactly which test gives different results on the same version of the software.

How a Changed Mindset Shows in Daily Work

The most visible change shows in how developers think, even more than in the numbers. There used to be a central test team and the classic divide between development and testing. The most recent switch of the version control system to Git, however, happened because the developers asked for it. Their condition, roughly: if continuous integration and the tests work, go ahead and roll it out.

Customers notice the change as well, in the falling number of tickets. As a platform supplier, the team hears that the platform is easier to integrate and that each new version needs less adaptation. Defect reports are measurably down, while the platform is demonstrably still in use.

Why the Zero-Defect Policy Failed and Bug Jail Worked

Not every experiment worked. A zero-defect policy, which required teams to drop everything the moment a defect came in, went down badly. Marco Achtziger would never introduce it again.

The next iteration turned the principle around. Under the name Bug Jail, teams decided for themselves how many defects they would allow before stopping development. More cautious teams set the limit at two, others at five. Despite these differences, the approach worked much better.

The difference was autonomy. Once teams thought about the threshold themselves, they dealt with incoming defects sooner. The lesson applies to any change: name the goal and the problem, and leave the path to the team.

Where the Next Lever Is: Optimizing Test Execution

The current bottleneck is the sheer volume of test execution. A useful metric here is test execution time per calendar month. Ideally, that value would be one: a single machine running all the tests around the clock. In practice, it is approaching the theoretical maximum of the available machine capacity, because in a regulated domain many tests are mandatory.

The approach being pursued is a machine learning system that predicts, for a given source code change, which tests are most likely to fail. That could speed up a gated check-in, for example: the run can stop at the first failing test or once a set threshold of failing tests is reached. That brings the number of machines needed back down to a reasonable level.

Where to Start: Value Stream Analysis

Every change should start with a value stream analysis, a tool from lean. Before you change anything, write down the numbers and look at where your time actually goes. Experience shows that what you think the problem is at the start is almost never the real problem.

Often it’s something mundane, such as waiting for another team, that eats most of the time. Once the value stream is in front of you, tackle the biggest time sink you can actually influence.

Plan for patience and iterate. Doing things more often and faster changes the problems themselves, so after a while it pays to look at the old problem list again. Some steps even get slower at first. Trunk-based development asks individual developers to think harder, about backward compatibility, for example.

So always look at the whole chain, not the single step. A sub-step that slows down for a while doesn’t mean the whole chain slows down. Quite the opposite: it gets faster.

Frequently Asked Questions

Can the benefits of continuous integration and delivery be quantified?

Yes, using data that’s already available. Source control systems, test execution logs, and ticket systems provide timestamped information that can be analyzed over the course of years without having to collect it separately. Anyone who claims that a change will improve the software has a burden of proof. The prerequisite is that the separate data sets can be linked, for example, via a common version number.

How long does it take for new ways of working to take root in a development organization?

Two to three years. That’s how long it took in the case described for a craftsmanship program to become part of the culture. In the end, nearly all teams were involved, with about half of them actively participating. Anyone expecting a cultural shift to take hold within a few quarters will be disappointed. The real obstacle is habits, not technology.

Do you have to convince all developers to adopt a new way of working?

No. In any change process, about ten percent will lag behind and remain unconvinced. Some still consider the old version control system to be the better one. The goal is to reach the majority. The most effective lever is the community’s opinion leaders: if you make them pilot the new system, they become multipliers who spark curiosity in others.

Is a build that’s always green a good sign?

Not necessarily. A build that’s 100 percent green is actually suspicious, because it suggests that no one is working on the software anymore. Wherever people work, there’s bound to be some underlying instability. A single unstable test isn’t a problem, but ignoring it is. Tests that can’t be made to turn green within a few hours are better off in a quarantine build.

What changes when switching from version branches to trunk-based development?

Speed and defect distribution change. In the project described, the time to delivery of a change dropped from days to hours. Previously, defects peaked around each release because bugs had to be fixed across multiple branches. Afterward, a constant, flat trend emerged at a lower overall level.

How many open defects should a team tolerate before halting development?

It’s best for the team to set this threshold itself. Under the name “Bug Jail,” more conservative teams set a maximum of two defects, while others set five. Despite this variation, the approach worked significantly better than a mandated zero-defect policy, under which all other work would have to be put on hold immediately. As soon as teams began to think about the limit themselves, they addressed incoming defects sooner.

Can a process change slow down individual work steps?

Yes, and that’s not a counterargument. Trunk-based development requires individual developers to think more carefully, for example about backward compatibility. What matters is the entire chain, not the individual step: a sub-step that is temporarily slower does not slow down the overall process. Moreover, doing things more often and faster changes the problem itself, which is why it’s worth revisiting the old list of problems later on.

Share this page