Skip to main content

Search...

Testing in German-Speaking Countries: How Agile Are Teams?

Testing in German-speaking countries has been surveyed every few years since 2011. Agile projects now dominate, while classic techniques lose ground.

• • Updated: • 10 min read
Cover of the expert talk on 'Testing in German-Speaking Countries: How Agile Are Teams?' with Karin Vosseberg and Richard Seidl.

The Software Testing Survey is a long-term study of how software testing is done in German-speaking countries, run every few years since 2011. It covers roles, methods, project models and tools. Among other things, the results show how far the share of agile projects has shifted and how widely specific test techniques such as equivalence partitioning or statement coverage are really used.

Key Takeaways

  • The Software Testing Survey has run every few years since 2011, and with the 2024 edition it offers comparable data across almost 15 years for the first time.
  • Code-centric agile practices rank high, while only about 50 percent of respondents rated collective code ownership as relevant and fewer than half use pair programming.
  • Classic test techniques such as equivalence partitioning and statement coverage are in decline; in 2020, 18 percent of respondents had never heard of statement coverage as a test technique.
  • The new survey asks specifically about AI at several levels, from generating requirements to creating test cases, and sets that use against the use of external contractors.
  • Every result, from the management summary to the technical report with raw data, is public and free, because the participating universities contribute publicly funded working time.

What the Software Testing Survey Measures

The Software Testing Survey tracks how testing actually works in practice, and it has done so every few years since 2011. It looks at how testers, test managers and companies in Germany, Austria and Switzerland work, which techniques they use and how their projects change over time.

There have been three runs so far: 2011, 2015/16 and 2020. The current edition adds a fourth data point, so the series now covers almost 15 years. That long view is the whole point of the project. A single number says little; the shift across several runs says a lot.

Today the survey is backed by the German Testing Board together with several universities. A company that was involved at the start was taken out on purpose after feedback from participants. The data should be clearly independent, and nobody should hold back because they suspect a vendor behind the questions.

Why All Results Are Free to Access

Everything the Software Testing Survey produces is public. That was a deliberate decision from the start, not an afterthought.

The reason is money. The universities put publicly funded working time into the project, and from that follows a simple rule: what the public pays for, the public gets to read.

“As long as universities are involved, there’s really no other way to do it. What is publicly funded has to be publicly accessible.”

(Karin Vosseberg)

The results come out in stages, each for a different audience. First a management summary with the most striking findings, then a shorter brochure with the key results, and finally a technical report running to several hundred pages with all the raw data. For most readers, the summary and the brochure are where the value is.

What the Questionnaire Covers

The survey opens by asking participants which role they see themselves in. That question has become more interesting over the years. Projects have moved from clearly separated roles to agile setups where those lines blur.

One finding from 2020: even though most projects were agile, many participants still placed themselves in classic roles. Asking about skills rather than roles might work better, that is, where people see their main strengths. The answers would then make more sense.

On process models, the numbers have flipped completely. At the start, about 25 percent of projects were agile and about 75 percent followed traditional, phase-based models. Today it’s the other way round. In practice, though, the models often blend, so the current edition lists hybrid approaches as a category of their own.

Calling Yourself Agile Isn’t the Same as Working Agile

The survey doesn’t stop at asking whether people call themselves agile. It also asks which agile practices they consider important for quality assurance. That shows whether an agile mindset has really taken hold or only the label has.

The 2020 results follow a clear pattern. Practices close to the code ranked high. Organizational practices, the ones that make up the mindset, trailed well behind.

  • Collective code ownership, meaning everyone shares responsibility for the code, was seen as relevant by only about 50 percent.
  • Pair programming is apparently used by fewer than 50 percent, even though it is an important part of quality assurance.

The focus on a single goal, getting the code to run, has grown stronger. Other aspects slip into the background. When the whole team keeps an eye on quality, that is a safety net, and code that merely runs doesn’t replace it.

Classic Test Techniques Are Losing Visibility

Classic test techniques have clearly lost ground in practice. Equivalence partitioning, boundary value analysis, and statement and decision coverage play a much smaller role than they used to.

One detail stands out. In 2020, 18 percent of participants didn’t know statement testing, the technique aimed at statement coverage, at all. For an established technique, that is a remarkably high number.

Automation is one possible explanation. In agile projects with short cycles and automated tests, the dashboard reports statement coverage automatically. The technique behind the number disappears into the tool and is no longer seen as a method in its own right.

That leaves an open question for quality assurance. Is coverage now only used to check whether the existing test cases cover everything? Or is it still used to derive new test cases systematically? The shift points to where skills are missing and where training should start.

How the Current Edition Looks at AI

The current survey puts a clear focus on AI. Earlier editions already had one general question on AI, so there is a trend line to compare against. New are specific questions about where AI is used, at several levels from generating requirements to creating test cases.

Three things are measured: whether AI is already in use, how much potential participants see in it, and whether external people handle the same tasks. Putting these answers side by side makes for an interesting analysis.

The underlying question is whether organizations that use AI rely on fewer external contractors. Whether that correlation exists can only be seen with enough responses, which is exactly why the survey needs broad participation.

Every question is reviewed before each new edition. Topics that no longer yield new insights are dropped. Questions on outsourcing, for example, were removed because the last runs showed it was no longer a relevant issue.

Industry Explains Surprisingly Little

Across several analyses, the survey shows no major differences between industries. That runs against what most people would expect.

You’d assume that regulated sectors such as automotive or medical technology test differently from everyone else. The data doesn’t show that gap as clearly as expected. If you measure your own work against industry norms, that is worth knowing.

Using the Results as a Benchmark

The results work as a benchmark. You can take the figures and see where you or your company stand in the overall picture.

For companies, the newer questions offer the most to work with. If many participants see a lot of potential in AI, that’s a reason to look into it yourself. Interest in training also shows which skills are in demand.

For freelancers and the self-employed, the survey gives a sense of the market. It helps answer where further training pays off and which specialization makes sense. On request, the team also runs industry-specific analyses, for example for finance or automotive, as far as the data allows.

Why Participation Decides How Much the Data Says

The quality of the results depends on the number of responses. Without enough participants, there are no reliable conclusions, and then the survey helps neither the people who answer nor the people who analyze it.

There are separate questionnaires for different target groups:

QuestionnaireTarget groupScope
ManagementExecutivesShorter, focused on organizational questions and the overall setting
Operational staffTesters, test managersMore detailed, about half an hour

In each of the last three surveys, around 1,000 people started the questionnaire and about 700 to 800 finished it. That isn’t representative in the strict sense; for that, one percent of everyone working in the field would have to take part. It is still a solid basis for reliable statements.

The current survey runs from September 1 to 30. A first management summary follows soon after, and the first results with a focus on AI will be presented at QS-Tag in early November.

Frequently Asked Questions

How often is the Software Testing Survey published, and what time period does its data cover?

The survey has been conducted every few years since 2011. Surveys were conducted in 2011, 2015/16, and 2020; with the 2024 edition, comparative data spanning nearly 15 years will be available. This long-term perspective is at the heart of the project: A single value is less meaningful than the trends observed across multiple surveys.

Who is behind the survey, and how independent is the data?

It is sponsored by the German Testing Board in collaboration with participating universities. One company that was originally involved was deliberately removed following feedback to ensure the data is recognized as independent. Because the universities contribute publicly funded work time, all results are available free of charge: an executive summary, a brief brochure, and a technical report with raw data.

How has the ratio of agile to traditional project management models shifted?

It has reversed. At the start of the surveys, approximately 25 percent of projects were agile, compared to about 75 percent that used traditional, phase-based models; in later rounds, the ratio was exactly the opposite. Because both approaches are often blended in practice, the 2024 edition breaks down hybrid methods into a separate category.

How can you tell if a team is truly working in an agile way or just using the label?

By the practices considered important for quality assurance. In 2020, practices closely related to program code were highly valued, while organizational practices lagged significantly behind: only about 50 percent considered collective code ownership relevant, and apparently less than half used pair programming. Calling yourself “agile” therefore says little about the mindset.

Is statement coverage still recognized in practice?

Only to a limited extent. In 2020, 18 percent of participants were completely unfamiliar with statement testing as a method, a remarkably high figure for an established procedure. One explanation is automation: The dashboard provides the coverage metric, and the method itself becomes hidden within the tool. It remains unclear whether coverage is still used to systematically generate new test cases.

What is being surveyed regarding the use of AI in software testing?

Three dimensions: whether AI is already being used, what potential the participants see in it, and whether external personnel are employed for the same tasks. Questions cover multiple levels, from requirements engineering to test case creation. The underlying question is whether organizations that use AI employ fewer external contractors. This requires a large number of responses.

Do regulated industries such as automotive or medical technology perform testing differently than others?

The data does not reveal this difference with the clarity that might be expected. Across multiple analyses, no major industry-specific differences were identified. Those who align their work with industry standards should take this into account. Industry-specific analyses are available upon request (for example, for the financial or automotive sectors) to the extent that the data allows.

How reliable are the results of such a survey?

In each of the past three rounds, approximately 1,000 people began the questionnaire, and 700 to 800 completed it. This is not representative in the strict sense; for that, one percent of all those working in the field would need to participate. However, the sample size is sufficient for reliable conclusions and as a benchmark for assessing one’s own position.

Share this page