Skip to main content

Search...

Software Expert Witness: State of the Art Means Proven

A software expert witness checks whether software is state of the art, a legal term meaning proven, and judges its maturity by defect density.

• • Updated: • 13 min read
Cover of the expert talk on 'Software Expert Witness: State of the Art Means Proven' with Sebastian Dietrich and Richard Seidl.

A court-certified software expert witness examines whether software reflects the state of the art and which quality defects it has. State of the art is a legal term rather than a technical one: a solution has to do more than work. It has to be proven, fit for purpose and suited to the specific field in which it is used.

Key Takeaways

  • Software with a defect density above two bugs per 1,000 lines of code is considered immature and not yet ready for production use.
  • “State of the art” is a legal term, not a technical one: a technology has to be more than new and widespread, it has to be proven in exactly the field where it will be used.
  • Writing more unit tests to raise low code coverage often makes things worse, because it locks in existing misbehavior that the code already depends on.
  • In Sebastian Dietrich’s experience, agile development has raised speed and productivity, but it has not produced higher software quality than earlier, heavier approaches.
  • Lawsuits over software defects rarely reach a courtroom, because public proceedings expose the client to its own customers, so both sides prefer to settle out of court.

What a Software Expert Witness Examines

A software expert witness tells a court whether a piece of software meets the state of the art and how good it really is. The full title is “generally sworn and court-certified expert.” Sworn means the expert took an oath in court once and does not have to be sworn in again for every case. Certified means he passed an exam in his specialty before a judge and two subject-matter examiners.

The idea behind it is simple. When two parties in a civil case both insist they are right and the judge lacks the technical background, the judge proposes an expert from a list. If both sides accept that person, they accept the expert’s assessment. The expert brings clarity where the court cannot judge on its own.

Why Software Disputes Rarely Reach a Courtroom

Most software disputes are settled out of court. The reason is discretion. A company that is unhappy with software it bought does not want its own customers to hear about the problem. Court proceedings are public, so the parties prefer a quiet agreement.

State-owned or state-affiliated companies are the exception. For them, the Court of Audit could question an out-of-court settlement and suspect improper payments. So they deliberately seek a settlement in court, even when the actual dispute was resolved long ago. In a civil case, the judge first asks whether the parties could settle out of court. The answer is no, but the documents are already prepared and just need signatures.

State of the art does not mean the newest and coolest. It means proven. That is the most important correction for any developer who reaches for the latest technology by reflex. The term comes from law and appears in many statutes. If software does not meet the state of the art, that has consequences for warranty claims and damages.

Four criteria define the state of the art. A technology has to be current, it has to be advanced, it has to be proven in the specific field where it is used, and it has to solve the task effectively and efficiently. Being new and celebrated at conferences is not enough. What counts is a track record in real use.

Many JavaScript web technologies are a good example. They are often a poor fit for software meant to run for ten or twenty years, because they change too fast and offer no reliable long-term support. A new version comes out, and six months later the old one is no longer maintained. In a short-lived web context, you can live with that. Over a long lifespan, it eventually blocks even security-critical updates.

“State of the art isn’t a technical term at all, it’s a legal one. It’s not enough that I can solve the problem. I have to solve it efficiently and effectively, and that means I have to think about the next few years.”

(Sebastian Dietrich)

Defect Density as a Measure of Maturity

Defect density tells you how many bugs are still hidden in a piece of software, measured per 1,000 lines of code. Above two defects per 1,000 lines of code, the software counts as immature and not ready for production. Where human lives are at stake, the threshold drops to 0.5 defects per 1,000 lines.

The number of defects the software originally contained can be roughly estimated with formulas. It depends on the size of the software, its complexity and the automated tests. Automated tests that run continuously catch bugs early, and those bugs never show up in a bug tracker because they get fixed on the spot.

From the estimated total, the expert subtracts everything that testers, trial operation or live operation have already found. What is left is the number of bugs nobody has found yet. In practice, that number is sometimes surprisingly high, even for safety-critical software: somewhere between 5,000 and 12,000 bugs in a larger system.

Tools Give You the Overview, Not the Verdict

With millions of lines of code, nobody checks every line by hand. Tools provide the overview, and the best known is SonarQube. Its strength is that it works across programming languages and makes size, complexity and technical quality visible.

But SonarQube stays at a low level. It shows code smells, not architecture or design flaws. What Sebastian Dietrich calls a “code stench” goes further than a smell: it will very likely cause a bug. A typical Java example is a comparison with a single instead of a double equals sign in an if condition. The code runs, but almost certainly nobody wrote it that way on purpose.

Be careful with blind fixes. Cleaning up smells across the board can create new bugs, because parts of the system may depend on the old misbehavior. And no tool can tell you whether a framework will carry a piece of software for the next ten to fifteen years. That takes a look at the framework’s history, the number of committers and the support it actually gets. It is a question for people, not for a tool.

Functional and Technical Quality Are Two Different Things

Software quality has a functional and a technical dimension, and each is assessed separately. Functional quality asks whether the software does what was required: whether all stories are implemented and whether the requirements and specification documents are met. That is the domain of testing.

Here it helps to look at the results backward. There are many types of testing, all of them state of the art, and none is enough on its own. The deciding question is how many bugs each type of testing has found. If one type stops finding anything, that does not mean the software is bug-free. It means another type of testing is due.

Functional quality also covers non-functional requirements, and those are often vague. Usability rarely gets a concrete requirement, and for performance the specification often just says the application must perform well. What that means is left open. A two-second response time is one benchmark, but a year-end financial close can take longer to compute without being slow. The sober question from the tester’s side is often just: have you ever tested usability or performance at all?

Technical quality revolves around the state of the art and has hundreds of criteria. They range from naming and coding conventions to code smells, dependency cycles, dead code, duplicate code and whether the defined architecture was followed at all. One question is whether the car looks good inside and out and gets you safely from A to B. The other is to open the hood and check whether you find a metric bolt or a wood screw.

More Unit Tests Are Never the Right Recommendation

When test coverage is too low, the recommendation is never to simply write more unit tests after the fact. That makes things worse, because it cements a bad state. And code coverage says nothing about how smart a test is anyway.

The better approach starts with the change process. Before every change to the code, someone writes a test that checks exactly that change, and it has to be a smart test. Quality then grows where the system actually changes, instead of coverage being produced blindly.

Quality gates carry this logic through the whole approach. The principle is always the same: it must not get worse, and it has to get better step by step. If coverage sits at 13 percent although 70 percent was promised, it has to rise with every release. The rules apply to new or changed code, and metrics such as coverage or duplicate code must never go down.

From Mount Everest to the Zugspitze

A first quality assessment usually turns up a huge pile of problems. The right reaction is not shock but a first step. Picture yourself standing at the foot of Mount Everest. You climb a mountain by walking and making sure you don’t fall. One measure at a time, keep going, don’t slide back.

The goal doesn’t have to be the summit. At some point the team stands on the Zugspitze, Germany’s highest peak, and realizes that the quality is now good enough, even if the system is still clunky. Negotiating a conditional acceptance buys you time for exactly that: the software goes into use, defined issues have to be fixed within a year, and a large part of the payment only follows after that.

If you are serious about paying down technical debt while you keep developing, expect three to four years before the software reaches the level you would have wanted at acceptance. This is where SonarQube helps again, because it can map the quality gate logic of “better than before.” Flipping the switch that enables every rule at once and produces a report that tears everything apart is the wrong way to go.

Agile Raised the Speed, Not the Quality

Agile development has increased speed and productivity, but not quality. That is the uncomfortable finding from many audits across industries. In the past, huge amounts of time went into mountains of paper, and the result was the same functional and technical quality that teams reach today with less effort.

The mindset has become more agile, often for real and not just on paper. People think in a more agile way and no longer fill out test reports mechanically; they approach the work more intelligently. The output has become cheaper as a result, but not better.

The same goes for heavily regulated fields such as aviation or healthcare. They document more, much of the work is less agile, and everything costs more. The result barely changes. Nobody there has a secret recipe either.

Go Back Years Later and See What’s Left of Your Software

The most honest quality test comes years after go-live. Five, six, seven years after the project ends, go to where your software is used and see what has become of it. Only then do you know whether a cool project delivered real value.

There are two possible outcomes. Either the software is still running and delivering value, and then it was a good project. Or it was thrown out long ago because nobody could work with it. User satisfaction alone does not prove business value. The value only shows once the system is running in production.

Frequently Asked Questions

When does a court call in an expert witness in a dispute over software?

If both parties in a civil lawsuit insist on their respective positions and the judge is unable to assess the technical issue on his or her own, the judge proposes an expert witness from a list. If both sides agree on this person, they accept the expert’s professional assessment. The full title is “generally sworn and court-certified expert”: sworn in through a one-time oath, certified through an examination before a judge and two technical examiners.

Is it sufficient for “state of the art” to use a current and widely used technology?

No. Four criteria must be met: The technology must be current, advanced, proven specifically in the relevant field of application, and it must solve the task effectively and efficiently. Being celebrated at conferences is not enough. Many JavaScript web technologies are unsuitable for software with a ten- or twenty-year lifespan because they lack reliable long-term support, and at some point even security-critical updates are blocked.

At what point is software considered too buggy for production use?

If the defect density exceeds two defects per 1,000 lines of code, the software is considered immature. Where human lives are at stake, the threshold drops to 0.5 defects per 1,000 lines. This value is determined by subtracting the number of bugs already found by testers, during trial operation, and in live operation from an estimated total number of bugs. In larger systems, this can leave 5,000 to 12,000 undetected bugs.

Can analysis tools like SonarQube definitively assess the technical quality of software?

No. SonarQube provides a cross-programming-language overview of size, complexity, and quality and highlights code smells, but it remains at a low level: it does not detect architectural or design flaws. Whether a framework will support a software application for the next ten to fifteen years depends on its history, the number of committers, and the level of real-world support. No tool can provide this assessment.

How do you test performance or usability requirements if they are only vaguely defined?

The first practical question to ask is whether usability or performance have ever been tested at all. If the requirements specification merely states that the application must perform well, there is no benchmark. A two-second response time is a common benchmark, but a year-end financial closing may take significantly longer to compute without being considered underperforming. Non-functional requirements are a matter of functional quality, not technical quality.

How do you improve insufficient test coverage without introducing new bugs?

Not by writing unit tests after the fact. That cements existing flaws that parts of the system already rely on, and code coverage says nothing about the quality of a test. It’s better to focus on the change process: Before every code change, a test is created that verifies that specific change. Quality gates ensure this, for example by requiring that coverage increase by 13 percent with each release.

How long does it take to reduce technical debt in a mature system?

Anyone seriously cleaning up the code while continuing to develop it in parallel should expect it to take three to four years for the software to reach the level that would have been desired for acceptance. Conditional acceptance creates room for this: The system is allowed to go live, defined issues must be fixed within a year, and a large portion of the payment is not due until after that. Enforcing all rules at once is the wrong approach.

Does agile development lead to better software quality?

Based on Sebastian Dietrich’s observations from numerous audits, it does not. Agility has increased speed and productivity: In the past, an enormous amount of time was spent on mountains of paper, yet the result was functionally and technically the same. The output has become cheaper, not better. Even highly regulated fields such as the aviation industry or the healthcare sector produce more documentation, work less agilely, and are more expensive, without significantly changing the outcome.

Share this page