Software metrics, from software quality metrics to lead time, are targeted measures that put decisions in software development on a basis of data instead of opinion. Useful metrics fit the context, show trends rather than absolute numbers and are accessible to the person who needs them. Quality, productivity and performance are closely linked, and trends reveal problems in all three early on.
Key Takeaways
- If you collect metrics but don’t use them to make decisions, you can save yourself the effort: data without consequences is wasted capacity.
- Software quality and development speed are directly linked: poorer code quality makes changes more expensive and stretches the time it takes to get work done.
- Lead times are not fixed numbers but random variables: they can be shown as a histogram and used for short-term forecasts with Monte Carlo simulations.
- Comparing teams on productivity metrics creates internal competition that weakens teams instead of strengthening them, because different contexts make fair comparisons impossible.
- Metrics can be thought of as concentric circles: some are only useful to the individual, others to the team and others to management, depending on who knows which context.
Metrics Are a Decision Aid, Not an End in Themselves
Metrics are meant to help you make faster, better decisions in your day-to-day work. That is what software quality metrics and every other metric are actually for, not filling a dashboard. In many companies, discussions about approach and priorities rest mostly on opinions. As software development grows more complex, that is no longer enough.
Maik Wojcieszak sums it up with a line from W. Edwards Deming:
“Without data, you’re just another person with an opinion.”
(W. Edwards Deming)
This is exactly where metrics come in. They turn a claim into a statement you can check.
The sticking point is not whether data is available. Today’s tools capture plenty of numbers and display them, too. What stays weak is how that data gets used. It ends up in dashboards that nobody consults for concrete decisions.
Why Overloaded Dashboards Hurt Decisions
Too much information does as much harm as too little. There are two ways to keep someone from acting: give them too little information, or bury them in too much. Dashboards that keep growing and getting more colorful fall into the second group.
The job is to filter out exactly the information that matters for a particular person’s work. Everything else is noise. A flood of information overwhelms people instead of making things clearer.
An everyday comparison shows what a good metric looks like. A pilot who had to read the instrument data from an Excel sheet would not have quick access to it. The speedometer in a car, on the other hand, no longer catches your attention, yet you use it all the time without thinking. That is how intuitive metrics should be in everyday work.
There Is No Universal Standard Set
Which metrics make sense depends on the specific case, not on a one-size-fits-all list. A small project needs fewer metrics than a large one. A single team needs fewer than an organization with many teams. A metric that works for one use case can be worthless for another.
Maik gives a clear example. The DORA metrics from the book “Accelerate” by Jez Humble, Gene Kim and Nicole Forsgren are popular in DevOps circles. But if you don’t do DevOps at all, for instance when building mobile apps, deployment frequency tells you nothing. The metric simply doesn’t fit the context.
That leads to a simple rule: understand the use case first, then pick the metrics that fit. Not the other way around.
Software Quality Metrics: Reliability, Recovery, Speed of Change
Software quality has many aspects, and several of them can be measured directly. Reliability shows up in failure rates. Mean time to recovery also gives you useful signals and trends.
One aspect many teams underestimate is how long it takes to make a change to the software. Most of what counts as internal quality, meaning clean code and clean architecture, is aimed precisely at making changes faster.
Ideally, measurement starts with the requirements. Requirements have a huge influence on how many rounds a team has to go through before a specification turns into what the customer actually needs.
Quality Data Belongs in the IDE, Not in Reports
Quality metrics only show their value when developers see them right as they write the code. Static analysis and linter results should appear in the IDE immediately and change behavior on the spot.
The reason is simple. If you never build poor quality in, you save yourself the expensive repair later. If you do build it in, things get gradually worse until you end up in a situation that costs a lot of money and a lot of time.
Reality often looks different. Tools produce mountains of findings that nobody looks at, because there are too many and there is no time. That is the opposite of intuitive use. The lever: find places where things could obviously be better, and start the conversation from there with concrete data. That turns metrics into a means of communication that also gives managers a solid basis for their decisions.
Measuring Performance Through Trends, Not Single Values
Performance metrics prove their worth when you look at them over time, from build to build. Classic performance tests run directly in the pipeline. Plot the results across several builds and you can spot trends and react right away when performance degrades. It is also how you check whether an optimization had any real effect.
Trends often say more than absolute numbers. A single value can jump in an individual case without reflecting a real trend at all. The direction of the arrow, better or worse, is the information you can act on.
Why Team Comparisons Are Misleading
Absolute metrics for comparing teams are rarely useful and hard to collect cleanly. If you measure productivity by how long a team takes for a certain number of requirements, a lower number says nothing about speed or quality. The context may be completely different.
Comparisons like these create internal competition, and that does damage. Maik compares it to crab pots. With a single crab, fishermen can leave the lid open and it will climb out. Put two in, and one keeps pulling the other back down. Lots of movement, no progress.
Lead Time Is a Random Variable, Not a Deadline
Lead time is easy to measure: from the start of a requirement to its completion. But anyone using it has to avoid two mistakes.
First, software development must not be treated as a linear system. If you do that, you are wrong from the outset.
Second, you cannot predict in advance when a requirement will be done; you can only determine it afterward. That gives lead time the character of a random variable. And a random variable is not represented as a single number but as a histogram.
Once you do that, several things become visible. You get a skewed distribution, because some tasks take unexpectedly long. That makes long-term predictions difficult or even impossible. For short-term forecasts, however, the metric works very well, for example combined with a Monte Carlo simulation.
One thing matters here: micromanagement gets you nowhere. Nobody needs to log their minutes. Lead time gives you the information you need without sliding into nitpicking.
Three Things a Metrics Initiative Needs
If you want to approach metrics deliberately, start by clarifying the purpose: What goal are you pursuing? Only then come the three building blocks.
| Building Block | What It Involves |
|---|---|
| Find the right metrics | Choose the metrics that fit the use case and provide them correctly |
| Learn to evaluate them | Understand relationships, for example that poorer quality increases the time it takes to get work done |
| Apply them | Metrics have to influence decisions, otherwise the effort is wasted |
Performance and quality are closely related. When quality drops, the time it takes to get work done goes up. So one metric lets you draw conclusions about the other. Performance, however, depends on many other factors too, which is why you need to look at several metrics together to see the connections.
The most common mistake lies in applying them. If you collect metrics but don’t let them influence your decisions, you can spare yourself the trouble. With a clear goal, choosing the right metrics becomes much easier.
Start With Yourself
The best way into metrics is your own work, not a big team program. You don’t need to launch an initiative for the whole organization to benefit. As a developer, tester or whatever your role, you can use and automate metrics for yourself first.
Maik describes this as concentric circles with the individual at the center. Some metrics are just for you. They help you see where you stand, and nobody else needs to see them. Next come team metrics that help the team assess itself, without management needing them or having to know the context. On that solid basis, you can derive further metrics that management uses for its decisions.
What metrics show best is where you currently stand: in the middle of the field, at the upper limit or at the lower one. Knowing where you stand is motivation enough to improve. Not in competition with others, but for the sake of your own work.
Frequently Asked Questions
Why Do Metrics Programs Often Yield So Little in Practice?
Because the data collected doesn’t influence decisions. Tools collect large amounts of numbers and display them, but they end up in dashboards that no one uses to make concrete decisions. Collecting metrics without basing your actions on them wastes capacity without providing any value in return. The purpose of a metric is to enable faster and better decision-making in day-to-day operations.
Are the DORA metrics suitable for every development team?
No. The metrics from the book Accelerate by Jez Humble, Gene Kim, and Nicole Forsgren are widely used in the DevOps environment but require that context. Those who do not practice DevOps (in mobile app development, for example) gain nothing from a deployment rate. It makes more sense to reverse the order: first understand the use case, then select the appropriate metrics.
How many metrics can a dashboard handle?
Significantly fewer than many dashboards display. Too much information can be just as paralyzing as too little, and ever-expanding, increasingly colorful overviews fall into the latter category. It makes sense to filter down to what matters for each person’s work. The benchmark is the speedometer in a car: unobtrusive, yet used constantly without a second thought.
Which aspects of software quality can actually be measured?
Several. Reliability can be gauged by error rates, and the mean time to recovery provides additional insights into trends. The time it takes to make a change in the software is often underestimated: clean code and clean architecture are aimed precisely at addressing this. Ideally, measurement begins right at the requirements stage, because it determines how many iterations are needed to reach a usable result.
When does static code analysis provide the greatest benefit to the team?
When the results become visible immediately as the code is being written. Linter and analysis results belong in the IDE, where they trigger an immediate change in behavior. Poor quality that isn’t built in to begin with saves on costly repairs later. Mountains of defects in a report that no one looks at, on the other hand, have no effect because there’s no time to clean them up.
Do productivity comparisons between teams make sense?
Rarely. Absolute comparative figures are difficult to collect accurately, and a shorter processing time for a given number of requirements says nothing about speed or quality, because the context can be completely different. Such comparisons create internal competition. Here’s an illustration: two crabs in a basket keep pulling each other down: lots of movement, no progress.
Can you derive a delivery date from lead times?
No, at least not in the long term. The completion of a requirement cannot be predicted in advance; it can only be determined in retrospect. This makes lead time a random variable and should be represented as a histogram, not as a single number. This results in a skewed distribution. The metric works well for short-term forecasts, for example in conjunction with a Monte Carlo simulation.
Where do you start if you want to begin using metrics without an organizational program?
With your own work. Metrics can be thought of as concentric circles, with the individual at the center: Some metrics are just for you and serve to help you assess your own performance, while others help the team gauge its progress without management needing to see them. Only on this basis can metrics emerge that inform management decisions.


