A lightweight test plan is a modular toolkit of short, linked topic booklets that people pick up for a specific situation instead of writing one big document once. Each module contains a user story, a description of the topic, a self-assessment for measuring maturity, and concrete steps for improvement.
Key Takeaways
- The QA toolkit of Germany’s Federal Office of Administration replaces heavyweight test manuals, some of them 300 pages long, with short modular building blocks that staff can use on their own, as the situation requires.
- Setting up traceability between test cases and requirements from the start costs little effort, while adding it later is a huge amount of work.
- A Community of Practice that meets six to ten times a year, regularly with 70 to 80 participants, has moved more knowledge at the Federal Office of Administration than training courses, because project teams swap concrete solutions directly.
- Outsourcing testing to service providers drained testing know-how from the organization, which is why the toolkit is needed as a way to win that knowledge back.
- The toolkit has been published under a Creative Commons license on the openCode platform, so other public authorities can benefit from it and contribute their own adaptations.
Why Heavyweight Test Plans Fail in Public Administration
Hundreds of pages of test manual go into a drawer and never get read. That is the pattern public authorities keep running into with classic test plans and test strategies, and it is the problem a lightweight test plan sets out to solve. The documents get written because a process demands them, not because anyone needs them.
Germany’s Federal Office of Administration (Bundesverwaltungsamt, BVA) is traditionally bound to the federal government’s V-Modell XT. That works well as long as the requirements are clear, for example when a law or regulation prescribes a particular specialized application. It gets difficult as soon as the situation changes quickly.
For Oliver Kortendick and his team, the so-called refugee crisis of 2015 was the moment to rethink. An application had to be up and running within a short time, even though the law behind it had not been passed yet. They started development and kept adapting the application as the legal situation evolved. With a rigid approach, that would hardly have worked.
A Large Organization Is the Natural Anti-Pattern for Agility
The Federal Office of Administration embodies almost every condition that makes agile work hard. More than 6,000 employees spread across 20 locations in Germany, plus many external service providers, a finely layered hierarchy and business departments that are separate from IT. All of that breeds sluggishness.
On top of that comes the sheer variety of the work: more than 150 specialized applications, from tiny ones to huge projects. Everyone calls for standardization, but with that range, standardization is hard to achieve.
In an environment like this, agile transformation only works as a path of small steps. Simone Mester and Oliver don’t describe a big bang or a framework bought off the shelf. They describe slow change. It still isn’t finished after years, but it is moving.
Testing Know-How Can’t Live Only With External Providers
When external service providers handle most of the testing for years, the organization loses its own knowledge. That was one of the central problems at the Federal Office of Administration. The question became: how do we get testing knowledge back to our own colleagues?
The starting point was a Community of Practice that has existed for more than five years. It meets six to ten times a year, regularly with 70 to 80 participants. For a voluntary format, that is a lot.
At its heart are experience reports. People from the projects talk about what they did, and connections form. Someone listens and realizes that another team has been working on exactly the same problem for months. This pragmatic sharing of knowledge replaces the urge to reinvent everything in every project.
Textbook solutions don’t help much here, because they often don’t fit the authority. What’s needed are approaches that have worked once and can be reused elsewhere.
What the QA Toolkit Is: A Lightweight Test Plan in Modules
The QA toolkit is a modular knowledge model from which you pull exactly the module that solves your current problem. Instead of reading a test manual of 120, 150 or 300 pages, you go straight to a single building block, like taking one booklet off a bookshelf.
The content draws on the ISTQB test process, the ISO software quality characteristics and the team’s own experience. The whole thing is built in Confluence, with the pages linked to each other.
The topics cover the practical side that tends to get lost in day-to-day public administration:
- Requirements management
- Test levels such as component, integration and system testing
- Test design techniques
- Traceability from test cases to requirements
- Review types
- Accessibility
- Test data
Traceability shows the benefit well. If you do it from the start, the effort is small. If you try to add it after the fact, working back from the test cases to the requirements, it turns into a huge job.
How a Building Block Is Structured
Every building block follows the same dramaturgy and so reaches different types of learners. At the top is the ISTQB definition, so everyone involved uses the same vocabulary. Especially with changing service providers, people otherwise end up talking completely past each other.
Next comes a user story. The stories are made up, but they could have happened in the office. Reading one is an easy way into the topic and raises a smile. The hope behind it: whoever smiles once will read the rest.
Then comes the simplest possible explanation of why the topic exists at all. Some content is also delivered through short YouTube videos, because many people find two minutes of video more pleasant than a long text.
Each building block has its own maturity model, tailored to that specific topic. It ranges from “initial,” where things only happen ad hoc, up to controlled levels. That lets you place yourself and see where you stand.
At the end are links for going deeper, from book recommendations to web sources and scientific articles. A small knowledge base catches topics that are too small for a building block of their own. There are also templates that are in high demand, for example for meeting minutes or for the basic question of how to write a test case in the first place.
Explaining Why Matters More Than Rules
People only get better when they know why. That insight came out of many feedback rounds with the office’s own staff, and it shapes the toolkit. People want to understand the context, not follow yet another instruction.
That’s why the knowledge model has several layers. You can stay on the surface and get a quick overview. Or you can follow the hypertext down into the details, all the way to scientific articles.
Behind this is a cultural question. In the classic silo, everyone sat alone, threw their results over the fence and lost sight of the overall process. Agile demands the opposite.
“In agile, what you want is for the team to share responsibility and commit to the highest product quality. That means everyone has to get out of their silo and bring in their knowledge and energy.”
(Oliver Kortendick)
How the Toolkit Works in a Real Project
The toolkit has been in use in the field since July 2024. The team deliberately picked a project in trouble: one with communication breakdowns where some people had stopped talking to each other.
That project had no risk-based test strategy. There were no priorities, and testing was spread evenly across everything. The result was significant delays.
From the toolkit, the team worked out three areas in which the project could improve by the end of November. Most of those points had already been worked through. In parallel, the team measured the mood regularly and simply: how are you feeling right now? Those scores improved, and that is exactly what matters for team building and good collaboration.
Exercises Close the Gap Between Reading and Doing
Some people only learn when they try something themselves. That’s why every building block will get a four-hour exercise. Three are ready, more will follow next year.
An exercise can be very concrete, such as writing a test case with boundary value analysis on paper while the facilitators sit alongside and the participants coordinate with each other. You can explain a technique to someone a hundred times. Some people just have to do it once.
The exercises are organized by role. The BVA works with established roles such as component owner or project manager. You don’t have to do every exercise, just the five or six that fit your role and move you forward.
Reports Have to Be Requested, Not Just Received
Reporting doesn’t only mean writing reports yourself. It also means receiving reports, understanding them and actively asking for them. That is a focus for the operational level.
The counterexample is common: external service providers deliver a stack of HTML pages under the motto “more is better.” That leaves you on your own and barely gets you anywhere. The better way is to say clearly what information you need and in what form.
Why the Federal Office of Administration Published the Toolkit
The QA toolkit has since been published under a Creative Commons license on the openCode platform. You can view it there, download it and adapt it to your own needs.
The reasoning is simple. A lot of work and money went into the toolkit, other authorities have the same problems, and there were already quite a few requests. What works should benefit others too.
It would be great if changes flowed back, because the team wants to keep learning itself. It can’t take on any liability or warranty. And sometimes the gain is simply learning from someone else’s mistakes.
Frequently Asked Questions
What are the drawbacks of a comprehensive test manual as a central document?
In practice, a manual of 120, 150, or 300 pages is rarely read. Such documents are often created because a process requires them, not because anyone needs them. The alternative at the Federal Office of Administration is a modular system consisting of short, interlinked topic booklets. When someone encounters a problem, they can select the specific module they need instead of working through the entire manual.
When does a strict V-model approach reach its limits?
As long as the requirements are clear (for example, when a law or regulation specifies a particular specialized application), the federal government’s V-model XT works well. Things get difficult as soon as the situation changes rapidly. During the 2015 refugee crisis, a procedure had to be in place before the legislative decision was finalized. Development began and was continuously adapted to the legal situation.
What happens to an organization’s testing expertise when service providers take over testing?
It is lost internally. It was precisely this experience that prompted the Federal Office of Administration to reclaim that knowledge. The initiative began with a Community of Practice, which has been in existence for over five years, meets six to ten times a year, and regularly brings together 70 to 80 participants. The focus is on workshop reports from projects, not textbook solutions.
Is it worth establishing traceability between test cases and requirements retroactively?
Hardly. If you maintain traceability from the start, it requires minimal effort. Trying to trace it back from the test cases to the requirements later on becomes a massive undertaking. Traceability is therefore one of the topics that should be clarified early on in government operations, alongside requirements management, test levels, test design techniques, review types, accessibility, and test data.
What is the purpose of a self-assessment of maturity level in a test plan?
It reveals your current status before improvement steps are selected. In the QA toolkit, each module has its own maturity model tailored to its specific topic. The scale ranges from “initial” (where actions are taken only on an ad hoc basis) to controlled stages. This classification determines which specific steps make sense to take next.
Why is it not enough to simply formulate quality requirements as instructions?
People only improve when they understand the context. This insight comes from numerous feedback sessions with the office’s own employees. The knowledge model is therefore structured in multiple layers: On the surface, there is a quick overview; hypertext links lead all the way to scientific articles. Behind this lies a break from the silo mentality, where everyone simply threw their results over the fence.
What are the consequences of testing without a risk-based testing strategy?
Testing is applied indiscriminately across the board because there are no priorities. In a deliberately selected project that was running into trouble, this led to significant delays, accompanied by communication breakdowns, to the point where teams stopped talking to one another. Three key areas were identified from the modular framework, allowing the project to focus its efforts specifically on them.
How do you get useful test reports from external service providers?
By actively requesting them and specifying the format. The opposite is common: the service provider delivers a collection of HTML pages under the motto “the more, the better,” and the client is left to figure it out on their own. Reporting also means receiving and understanding reports. Clearly state what information you need and in what format.


