Skip to main content

Search...

Software Metrics: It’s the Context That Makes a Number Useful

Software metrics are often avoided because they reveal inconvenient truths. Yet a simple table can be enough to forecast testing effort and defects.

• • Updated: • 11 min read
Cover of the expert talk on 'Software Metrics: It’s the Context That Makes a Number Useful' with Manfred Baumgartner and Richard Seidl.

Software metrics are measures that make the size, complexity and quality of software measurable, from requirements and architecture to code and test results. Teams use them to steer projects, estimate effort and assure quality. Without them, a project is like a car with no speedometer, fuel gauge or navigation system.

Key Takeaways

  • Software metrics are still used far too rarely in practice, even though simple measures such as defect trends or test case counts can improve project control considerably.
  • A single metric only tells you something once you relate it to another: ten test cases for 500,000 lines of code show at a glance that the ratio is off.
  • Metrics apply not only to code but also to requirements documents, architecture designs and test results, and most projects leave that potential untouched.
  • Combining historical defect data with estimated test case counts lets you forecast defect numbers for later project phases with surprising accuracy, as one project with a deviation of about ten percent shows.
  • You don’t need complex formulas to get started: a simple list of business process, complexity level and expected number of test cases is enough for a solid first planning figure.

Why Software Metrics Are Often Neglected

Many projects barely use software metrics, even though software development is an engineering discipline. There are two reasons for this. Collecting metrics takes work, and the numbers sometimes bring uncomfortable facts to light.

Manfred Baumgartner has watched this pattern for many years. Even simple metrics are often not put to good use, although they would be easy to get. Compared with other production processes and engineering fields, software development measures surprisingly little.

Part of the problem is how the industry sees itself. Many people regard software as something artistic that resists an engineering approach. That self-image is exactly why size, complexity and quality so often go unmeasured.

A Project without Metrics Is a Car without a Speedometer

Metrics are a project’s sensors. Without them, you are driving blind. A car with no speedometer, fuel gauge or navigation system gives you a vague sense of speed but no reliable picture of where you are.

The analogy carries over directly. How much budget is left, how much time, and how far will it get me? Those questions matter just as much during a project as the fuel gauge does on the road. Yet many projects run as if they had none of these sensors at all.

40 km/h is fast in town and slow on the highway. A number only gets its meaning from context. That’s why collecting a metric isn’t enough. You have to know what it stands for.

The Three Dimensions of Software Metrics: Size, Complexity, Quality

Metrics fall along three dimensions: quantity, complexity and quality. This triad brings order to what would otherwise be a confusing mass of figures.

Quantity metrics are the simplest. You count lines of code, test cases, artifacts. Even that tells you something useful, for example when you want to estimate the testing effort.

Complexity is harder to pin down but has a bigger effect on effort. How complex a piece of software is drives test and retest effort more than its sheer size. Quality is the third dimension, and because the term itself has no universal definition, it’s the hardest to measure.

A Number on Its Own Means Nothing

Quantity metrics only become meaningful in relation to something else. Ten test cases for 500,000 lines of code show immediately that something doesn’t add up. The single number is worthless. The ratio is the finding.

There is no universal rule of thumb. Nobody can tell you that 100,000 lines of code need exactly 150 test cases. Every system is built differently. Programming languages produce different amounts of code, and teams that rely heavily on libraries write little code of their own but still need test cases for all of it.

That has a practical consequence for organizations. Metrics work best as a framework or guideline that you adjust iteratively. Values hardly transfer from one application to the next, but they can be refined within a company over time.

Code Is Not the Only Thing You Can Measure

Far more than code can be measured. Requirements, design, architecture, code and tests are each objects of measurement with their own metrics.

Measuring requirements has become harder. Where teams once had detailed requirements and specification documents, agile projects often have nothing but user stories. Deriving size or complexity from plain text is difficult. Where requirements are described in a structured or model-based way, metrics are still possible.

Code metrics are what most people think of first. In the past, people counted GOTO statements, today it’s class calls and lines of code. Alongside those sit test metrics that focus on test deliverables and test results.

Designs are surprisingly easy to check. You can count the tentative wording in a design document. If every other sentence contains a “maybe,” a “could” or an “under certain circumstances,” implementation gets hard, and testing even more so. Developers often mind this less, because vague wording gives them room to maneuver.

Know What a Metric Is For Before You Collect It

Every metric starts with a question: what do you want to do with it? Only the answer turns a number into a useful tool.

Design metrics let you derive effort, such as how much development and testing work a design is likely to need. That belongs to planning a project or part of one. Other metrics serve quality assurance of the test object itself, checking whether standards and criteria are met.

Before a software migration, it helps to know the volume and internal quality of the existing system. Whether a system was built cleanly and straightforwardly or is a complete mess makes a big difference for refactoring. Function point analysis provides a standardized measure of functional size that can be calculated from the existing code.

The Trend Matters More Than the Single Value

For legacy projects, the direction of the trend is often worth more than the absolute value. A static analysis of code that has grown over years quickly produces overwhelming numbers. The more honest and useful target is then: it must not get worse.

That takes pressure off. You don’t have to rework every legacy issue. It’s enough to hold the value at its current level while new code is added and existing code is reworked.

Metrics also help with architecture decisions. When several systems in different programming languages do the same job, complexity and quality figures show which one is fit for further development and which one should be phased out.

What Teams Measure in Practice, and What’s Missing

From a testing perspective, defect counts and defect trends are the classic metrics. They are widespread, but there is often room for improvement in how they are interpreted.

In agile projects, the data gets blurry. Not every defect ends up in a defect management tool, and some disappear into a backlog nobody can analyze. Some companies keep documenting consistently, others lose track.

At the technical level, teams do measure. Metrics from the CI/CD pipeline and code coverage at the unit test level are available. What these numbers say about actual quality and scope, though, rarely gets enough attention.

Productivity measurement often fails for lack of data, too. If business departments are heavily involved in test projects but their hours are never broken down by test activity, you can’t calculate a reliable productivity figure. Burn-down and burn-up charts get set up, then vanish in the day-to-day rush without anyone acting on them.

A Simple Metric That Held Up in a Real Project

A metric doesn’t need to be complex to work. One pragmatic approach from a real project used just three columns: business process, complexity, number of test cases.

The estimation logic was simple. Low complexity meant two to three test cases, medium about ten, high around twenty, and unclear cases potentially hundreds. Added up, that gave a first effort estimate: number of test cases times days per test case.

The key was recalibrating after every sprint. Planned test cases were compared with the ones actually designed, and the estimates per complexity level were adjusted. The plan was set by the numbers and then sharpened by the content.

The result was strikingly accurate. The approach also included an expected number of defects per test case, based on historical data. At the end of the project, the actual defect count was about ten percent off the estimate made months earlier. Thanks to the law of large numbers, the values converged on the planned figure.

“Don’t expect the next sprints to be much better. You can hope for it, but that hope is usually unfounded.”

(Manfred Baumgartner)

That projection into the future is where the real value lies. If you see early on that defects are spread out statistically instead of clustering at the start, you can act: bring in additional testing resources or warn project management in time.

How to Work with a Software Metrics Compendium

You don’t read a compendium cover to cover. You dip into it. The introductory chapters on the three dimensions of quantity, complexity and quality are worth reading in one go. After that, the index gives you targeted access.

For a dimension such as design, several measurement approaches exist side by side, for example those by Tom Gilb or by Card and Glass. No standards body has defined one single, uniform measure. Treat that variety as an invitation to pick the approach that fits your environment.

The first step is always the same. Get an overview, then pick out what you can apply in your project. None of these metrics comes ready-made out of a tool. You have to put in some work yourself.

Frequently Asked Questions

What dimensions should be distinguished when it comes to software metrics?

Quantity, complexity, and quality. Quantity metrics are the simplest: lines of code, test cases, or artifacts can simply be counted. Complexity is harder to grasp, but it is a driver of the effort required for testing and retesting more than sheer quantity. Quality is the most challenging aspect because the term itself isn’t universally defined.

Is it enough to collect a metric to draw a conclusion from it?

No. A number only gains meaning through context: 40 km/h is fast in town and slow on the highway. Metrics are the sensors of a project, comparable to a speedometer, fuel gauge, and GPS. If you don’t know the frame of reference, you may have a measured value, but no reliable insight into your own situation.

Is there a universal rule for how many test cases a given amount of code requires?

No. No one can say that exactly 150 test cases are needed for 100,000 lines of code. Every piece of software is developed differently: Programming languages generate different amounts of code, and those who rely heavily on libraries have little code of their own but still need to cover everything. Such values can hardly be transferred between applications, but they can be adjusted over time within a company.

Can requirements and designs be measured in addition to code?

Yes, requirements, design, architecture, code, and tests are each distinct objects of measurement. When it comes to requirements, it has become more difficult because requirements and specifications are often replaced by simple user stories. In the design document, a simple approach helps: count the number of conditional forms. If every other sentence contains a “maybe” or “under certain circumstances,” testing becomes particularly difficult.

How do you deal with metric values that are overwhelmingly poor in a legacy system?

With mature code, the trend is more important than the absolute value. A static analysis quickly yields numbers that no one can improve in the short term. The more realistic goal is: It must not get worse. This takes the pressure off because not all legacy issues need to be rewritten; the value just needs to maintain its current level.

What simple metric is suitable for getting started with test effort estimation?

A three-column list is sufficient: business process, complexity level, and expected number of test cases. In a real-world project, low complexity corresponded to two to three test cases, medium to about ten, high to around twenty, and unclear cases to potentially hundreds. The total multiplied by the number of days per test case yields an initial effort estimate, which is readjusted after each sprint.

Is it realistic to predict defect counts for later project phases?

Yes, if you combine estimated test case counts with historical defect data per test case. In the project described, the actual number of defects at the end was about ten percent off the estimate made months earlier. According to the law of large numbers, the values converge toward the planned value. The benefit lies in taking early action: allocating additional testing resources or alerting project management.

Why do existing metrics in projects often go unacted upon?

Because there’s a lack of interpretation. Metrics from the CI/CD process and code coverage at the unit-test level are usually available, but there’s little analysis of what they reveal about quality and scope. In agile projects, defects get lost in backlogs that can’t be analyzed; burn-down and burn-up charts are created and then disappear in the hustle and bustle without anyone taking action based on them.

Share this page