Skip to main content

Search...

Sustainable Software Development: How Testing Saves Energy

Sustainable software development starts with the requirements, long before deployment. Performance, maintainability and lean test pipelines all count.

• • Updated: • 12 min read
Cover of the expert talk on 'Sustainable Software Development: How Testing Saves Energy' with Markus Lachenmayr, Florian Krautwurm and Richard Seidl.

Sustainable software development means using energy and hardware efficiently while keeping software durable and maintainable. Sustainability is not a new quality criterion here. It is an additional perspective on existing non-functional requirements such as performance, maintainability, and reliability. The test process itself also consumes resources, through load generators and pipelines for example, and offers concrete ways to save.

Key Takeaways

  • Sustainability in software development is not a new quality criterion but a new perspective on non-functional requirements that already exist, such as performance, maintainability, and reliability.
  • Long-lived, maintainable software is more sustainable than software that gets rewritten often, because every restart burns resources and developer time, and systems that have grown over the years hide large savings.
  • Moving performance tests to time windows with a high share of renewable energy uses power that would otherwise go to waste, at no extra cost from the cloud provider.
  • Unnecessary network traffic in pipelines and staging systems causes measurable CO₂ emissions: a badly built pipeline can equal 600 kilometers driven by car, an optimized one just 600 meters.
  • Where sustainability doesn’t work as an argument, cost does, because less cloud traffic and better resource utilization lower the bill directly.

Sustainable Software Development: A New Lens on Familiar NFRs

Sustainable software development can’t be handled as a standalone non-functional requirement. It is an additional perspective on NFRs that are already part of the project: performance, maintainability, reliability, security. As Markus Lachenmayr puts it, sustainability is not a new NFR but a new way of looking at the existing ones.

The word itself carries the wrong associations. This is not about activism. It is about how you use resources: hardware, energy, and the carbon released along the way. There is also an aspect that many sustainability initiatives overlook: the longevity of the software itself.

Starting over every two or three years burns through resources and people. Software that can no longer be maintained goes in the bin, and the team starts from zero. That is exactly what you want to avoid.

Why Sustainability Belongs in Requirements Engineering

No quality attribute can be tested in at the end of the development cycle. It doesn’t work for performance, it doesn’t work for maintainability, and it doesn’t work for sustainability. If you want sustainability, you have to get involved early, in the requirements phase.

Testers are used to digging up the uncomfortable topics that tend to get forgotten. Right now, sustainability is one of them. Florian Krautwurm argues that testers should join the requirements phase from the start and bring the sustainability perspective into the trade-off analysis with the existing NFRs.

The problem is not that anyone actively rejects sustainability. It simply gets overlooked. Architecture work focuses on the important NFRs, and sustainability quietly drops off the list.

Performance, Maintainability, Security: Where Sustainability Gains and Where It Clashes

Sustainability relates to other NFRs as a set of pluses and minuses. It doesn’t take care of itself. Some quality attributes help, others work against it. Those trade-offs need to be made explicit.

Performance and efficiency are the obvious levers. You can make software fast by building it efficiently, which takes design decisions early on. Or you notice at deployment that it runs slowly and simply open the resource tap wider. That is the unsustainable option.

Maintainability almost always helps sustainability. Security, on the other hand, often clashes with it: encryption and decryption need more computing power. The answer is not to drop security but to ask how far the measures really need to go and where compromises are acceptable. Keeping personal data safe is itself part of social sustainability.

Reliability deserves questioning too. For safety-critical systems, availability is non-negotiable. But does an About menu really need five nines of availability, or can it be down once in a while?

The overview below sums up the trade-offs discussed in the project:

NFREffect on SustainabilityConsequence
Performance / efficiencyPositive if designed for efficiency earlyMake design decisions early rather than adding resources later
MaintainabilityStrongly positiveLongevity reduces rebuilds and resource consumption
SecurityTends to be negativeWeigh measures, don’t cut them
Reliability / availabilityDepends on contextQuestion availability requirements for non-critical parts

Long-Lived Software Saves Hardware, Not Just Code

Maintainability and replaceability are direct sustainability levers. Software that gets rewritten every three years burns through people and resources. Software that keeps running on older hardware extends that hardware’s useful life.

This is where embodied carbon comes in: the carbon released when hardware is manufactured, recycled, and destroyed. The longer your software runs on older hardware, the better that balance looks. Backward compatibility therefore covers old hardware as well as earlier software versions.

Replaceability pays off over time. If a more resource-efficient version of an algorithm appears in a few years, you need to be able to plug it in. In a monolith that is hard. Modular, containerized structures make it possible.

How to Build Sustainability into Code Reviews

Sustainability belongs in the implementation, not just the design. A poorly written loop or missing caching wrecks efficiency at a level no architecture diagram shows. That is why you need code reviews that look not only for bugs but also at how things are implemented.

Proportion matters. Nobody gains from lecturing a developer about four saved bytes. The focus belongs on the hot spots: code paths that run millions of times. Optimizing those changes the curve. For one-off cases, the effort doesn’t pay off.

For the team to come along, every review comment needs a reason. Just giving orders doesn’t work. In terms of method, a sustainability review is no different from a secure coding review: one more aspect you bring to the table and explain.

“It doesn’t help to just look at it and say, do it differently. You have to give the reasons and teach people how. You should program efficiently anyway, for other reasons too.”

(Markus Lachenmayr)

Which Tools Make the Carbon Footprint Measurable

There is rarely a unit test for sustainability that gives a clear pass or fail. Much of the work happens through monitoring rather than test cases. Large cloud providers such as AWS, Azure, and Google Cloud offer their own dashboards that show the carbon footprint.

One tangible KPI is the actual utilization of the resources you pay for. If it sits in the single digits as a percentage, you are probably holding too much capacity and have room to optimize.

For maintainability, static code analysis with metrics such as technical debt, fan-in, and fan-out helps. These metrics are not absolute truth, but they tell you whether things are getting better or worse and make the current state measurable. For web UI efficiency, there are analyzers that point out hot spots at a high level. The actual work afterward is still manual.

One idea still waiting to be built: a custom ruleset for static analysis, with rules specifically tagged for sustainability. It hasn’t been implemented, but the idea holds up.

Sustainable Testing Beats Testing for Sustainability

The bigger lever is not testing for sustainability but testing sustainably. Test infrastructures are often large and consume plenty of resources. Performance tests scale up load generators and the system under test, and pipelines shuffle data back and forth.

Many test tasks aren’t tied to a specific time. Not every test has to run after every commit. Performance tests often run at night anyway, but they could just as well run at midday, when there is a lot of renewable energy in the grid.

For almost every country, there are APIs that return two things: the current electricity mix and current demand. If you find a window with a high share of renewables and low demand, you can move compute-heavy tests right into it. It won’t save you a cent, since the cloud provider bills you all the same. But the green energy you didn’t waste is available elsewhere.

The Cost of Badly Built Staging: 600 Kilometers versus 600 Meters

Test data and staging systems generate heavy network traffic, and that traffic can be converted into carbon. Containers get moved from local machines to the cloud and back again. That traffic adds up.

Florian cites a calculation from a widely discussed Shift report: a badly built pipeline including staging can equal 600 kilometers driven by car. Rebuild the same pipeline, and you are down to 600 meters.

The figures are rough approximations that depend on distance, caching, and other factors. Still, they are eye-opening, because nobody usually thinks about this consumption. In software development, resource use fades into the background, and everything seems to simply be there. Converting it into kilometers driven makes it tangible.

Traffic that never goes to the cloud costs nothing. This is where sustainability and cost meet, and that makes the case to management easier.

The Cloud Ended Bit-Twiddling and Left a Mountain of Debt

In the early days of the cloud, thinking shifted. Nobody had to squeeze every bit into place anymore, because resources could simply be scaled up. Many projects that consume far too many resources were born in exactly that phase.

A lot of them are still running today, some containerized, but following the same pattern: scale up the cloud, done. That is why the biggest untapped potential lies in software that has grown over the years.

Requirements keep changing, and the architecture rarely keeps up. A worthwhile review approach: take two or three steps back, hold today’s requirements against the current state, and ask whether you would write the core of the system differently with what you know now.

Re-architecting is part of that. If you realize years later that an algorithm could run far more efficiently, you should be able to act on it. That only works if the architecture was designed for it from the start. One trade-off remains: some legacy software runs lean precisely because hardware used to be scarce. A rebuild on a modern platform can end up using more. That trade-off has to be made consciously.

How to Start When Sustainability Isn’t on the Project’s Radar

The first step is knowledge, not a tool. Organizations such as Principles.Green and the Green Software Foundation have put far more time and effort into the topic and offer tutorials, articles, and concrete patterns for sustainable software. Learning from them and keeping that knowledge in mind is already worth a lot.

When design decisions come up, you draw on that knowledge and suggest more sustainable options. That is where reading turns into practice.

Speaking the stakeholders’ language is what makes the start work. If a company doesn’t value sustainability, you map the topic onto factors that do count. Cost almost always works, especially in the cloud, where energy and data transfer cost money directly.

So you lead with the argument that resonates with each stakeholder and add the sustainability angle afterward. Management likes to see consumption measured and kept low. That is the lever you use to bring sustainability into the project from day one.

Frequently Asked Questions

Is it worth defining sustainability as a separate non-functional requirement?

No. Sustainability doesn’t work as an isolated requirement, but rather as an additional perspective on the NFRs that are already part of the project: performance, maintainability, reliability, and security. It makes sense to incorporate this perspective into the trade-off analysis. The topic is rarely actively rejected; rather, it unconsciously falls by the wayside during architecture work because other NFRs take precedence.

Can sustainability still be tested into the system at the end of development?

No. No quality attribute can be retroactively tested into the system at the end of the development cycle: neither performance, nor maintainability, nor sustainability. If you want efficiency, you have to start in the requirements phase. Testers are used to uncovering unwelcome and often-overlooked issues, and they can take on this role here as well by being at the table early on.

Does security get in the way of sustainability?

Security tends to have a negative impact on sustainability because encryption and decryption consume computing power. The solution is to weigh the options, not to eliminate them: How far must the measures go, and where are compromises acceptable? Keeping personal data secure is itself part of social sustainability. Maintainability, on the other hand, has an almost universally positive effect.

Why is long-lasting software better for the hardware footprint?

Because embodied carbon matters: the carbon released during the manufacture, recycling, and disposal of hardware. The longer an application runs on older hardware, the better the environmental balance. Backward compatibility, therefore, affects more than just earlier software versions. If an application is rewritten every two to three years, each new release consumes resources and developer time.

What should a code review focus on when it comes to resource consumption?

The hot spots, that is, sections that are executed millions of times. Optimizing these areas changes the trend line; for isolated cases, the effort isn’t worth it. Calculating the savings from four bytes doesn’t help. Typical candidates include poorly implemented loops and a lack of caching. Every comment needs a justification; otherwise, the team won’t buy in.

Is it worthwhile to schedule computationally intensive tests based on the power mix?

Yes, because many test tasks aren’t time-sensitive. For nearly every country, there are API calls that provide the current composition of the electricity grid and demand levels. If you find a time window with a high share of renewable energy and low demand, you can shift performance tests to that window. This doesn’t save any money; the cloud provider still charges for the resources used.

How much do test pipelines and staging systems contribute to CO₂ emissions?

Significantly. A poorly designed pipeline, including staging, can be equivalent to driving 600 kilometers by car; the same pipeline, once redesigned, is equivalent to just 600 meters. These figures are rough estimates and depend on distance, caching, and other factors. The cause is traffic: containers are moved locally to the cloud and back again.

What argument resonates with stakeholders who don’t care about sustainability?

Costs. In the cloud, energy and data transmission directly cost money, so traffic that doesn’t go to the cloud lowers the bill. Management likes to see consumption measured and kept to a minimum. A tangible KPI is the utilization rate of the resources being used: if it’s in the single-digit percentage range, too much capacity is being maintained.

Share this page