Product testing at Stiftung Warentest is a structured process that starts about a year and a half before publication. A market analyst, a project manager and external testing institutes define a test program together, buy the products undercover and test them against standards and subjective criteria. Every test ends with a verification step that checks the measurement data for consistency.
Key Takeaways
- Stiftung Warentest plans test topics up to a year and a half ahead, guided by seasonality, user questions and market developments.
- Every product is bought undercover, so manufacturers know neither the timing nor the scope of a test and can’t send in specially prepared units.
- Subjective criteria such as handling or feel are made as objective as possible through blind tests, trained test panels and reference devices that run alongside.
- Every test project is teamwork. Market analysts, project managers, test labs and editors each bring their own perspective, because dry measurement data alone doesn’t tell a story.
- Manufacturers receive the objective measurement data before publication so they can respond to safety-critical findings, but they are not told the publication date.
What Stiftung Warentest Has in Common with Software Testing
Stiftung Warentest and software testers both test from the user’s point of view, not the manufacturer’s. Stiftung Warentest, the German consumer testing organization, describes itself as an independent, perpetual foundation whose goal is to help consumers make informed purchasing decisions of their own. Software testers follow the same aim when they check whether a product does what its users expect.
The shared values are objectivity, independence and transparency. Testers put themselves in the end user’s shoes and ask whether the product keeps its promises. In both worlds, that perspective is where every test starts.
Johannes Stiller brings a picture from consumer protection to this: the testing or scientific view meets the question of what matters to the user out there. Anyone who measures software against real requirements instead of internal specifications knows that tension.
How a Test Comes Together: From Idea to Test Program
A test process starts long before the actual test. At Stiftung Warentest, a year to a year and a half often passes between the first idea for a topic and publication. A test of running equipment published in January was decided about a year and a half earlier, partly because New Year’s fitness resolutions drive demand at the start of the year.
Topics come from several sources. User questions on the organization’s website, market monitoring and suggestions from the teams all feed in. Every idea goes through a critical evaluation against clear criteria.
A topic only moves forward if it meets four conditions:
- It has to make sense for consumers.
- It has to be testable.
- It has to be useful and interesting for readers.
- It has to be measurable.
Once an idea is confirmed, a test program is written. It is based on existing standards, where there are any. Sometimes no standard exists, and sometimes one is being drafted. The project manager studies the product in depth and designs the program, which is then cross-checked several times before it goes to the testing institutes.
How Stiftung Warentest Makes Subjective Criteria Measurable
Soft criteria are tested just as soberly and comparably. How easy a device is to handle, whether it reaches awkward spots, how brushing your teeth feels: none of that can be measured objectively like power consumption. It still can’t be left to chance.
Comparability comes from several mechanisms. Where possible, the same people test over long periods. Reference devices run alongside as fixed candidates, so different tests can be placed relative to each other. That keeps the yardstick stable even when the products change.
Bias is deliberately removed. Product names are taped over, devices are made unrecognizable, and testing happens blind and in varying orders. Testers are not supposed to know which brand is in front of them.
No single person rates a single device. All testers take turns with different devices in different orders, and their impressions are collected and merged. That spreads subjective outliers across the group.
How Test Products Are Selected and Bought
The products are bought undercover and never supplied by the manufacturer. A manufacturer can’t submit its own washing machine. The organization’s own buyers purchase the products on the open market, without manufacturers knowing what is bought when.
Selection follows the consumer’s view. A market analyst works full time on watching the market: who the big players are, who is bringing innovation, what sells well. A niche device that hardly anyone can buy doesn’t help consumers.
For endurance tests, one unit is often not enough. Several units are bought so that wear over time can be tested. If something breaks during the test, it is retested to confirm the result.
Quality Comes from the Team, Not One Person
No test depends on a single person. Behind each project is a group with different roles: the project manager, a market analyst, the testing side and the editorial team, who know what matters to readers. The mix is reminiscent of a cross-functional team in software development.
The diversity is deliberate. If you deliver the sharpest analysis but nobody understands it, you reach no one. Measurement data only becomes a useful test when the dry facts meet a story that can be told: what the highlights are, what the weaknesses are, whether anything is dangerous.
“It’s no use wearing the fanciest analytical glasses if nobody understands what you’re saying.”
(Johannes Stiller)
More eyes don’t automatically produce uniform answers. In the extensive verification process, every tester notices something different. That is exactly what makes the test stronger instead of smoothing it out.
Why Everything Gets Checked Again at the End
Before publication comes a central verification step. The measurement data is checked for errors and consistency. Only then does a result count as reliable.
The results of all tests come together in a central meeting for each individual topic. There, the team discusses what came out and what story it tells. The meeting takes place only a few days to weeks before publication, so the test stays as current as possible while remaining sound.
Every tester knows this logic: a finding is only a finding if it can be reproduced. If a device breaks, it gets retested until the result is confirmed.
How Stiftung Warentest Deals with Manufacturers and Bad News
Manufacturers are informed before publication, but they get no influence over the result. They learn that they are being tested, but know neither when the products were bought nor the publication date. They receive the objective measurement data, including critical findings.
For safety-critical findings, this communication has a direct purpose. When child car seats tore out of their mountings, the manufacturers were informed in advance. The questions behind that: how will you respond, can you offer goodwill, is there a material defect or a problem with the batch?
Reactions vary widely. Some manufacturers engage in a cooperative dialogue, others file for injunctions, even over a second place. Over the years, some product groups, such as mattresses and cordless vacuum cleaners, show a clear development: little works at first, and it gets better over time.
The tester as the bearer of bad news is a familiar role. Whoever reports defects easily becomes a target at first, even though the test serves the product and the user. The position only holds up with verifiable rules that everyone sticks to consistently.
Transparency Is the Foundation of Credibility
A traceable method is what earns the trust readers extend in advance. Stiftung Warentest discloses its test conditions and refers to the underlying standards. The magazine has a compact box titled “How we tested,” and the more detailed version follows online.
This openness is also an obligation. A clear mission has to be upheld, and compliance with the organization’s own rules has to be verifiable. The same principle applies to software testers: a finding only convinces when it’s clear how it was reached.
The line to users stays open. Feedback through email inboxes and forum posts is monitored, and hints about new aspects or changes flow back into the work. If you test for users, you also listen to them.
Frequently Asked Questions
How much lead time does a product test require at a consumer organization?
It takes between one and one and a half years from the initial topic idea to publication. Seasonality plays a role: A test of treadmills published in January was decided upon about a year and a half earlier, because New Year’s resolutions drive demand at the start of the year. This timeframe includes market analysis, the testing program, undercover purchasing, and laboratory tests.
What criteria are used to decide whether a test topic will be covered at all?
Four conditions must be met: The topic must make sense to consumers, be testable, be useful and interesting to readers, and be quantifiable. Ideas come from user questions on our own website, market observation, and suggestions from the teams. Only then is the testing program developed.
Why are test units purchased undercover instead of being requested from the manufacturer?
To prevent manufacturers from supplying special-edition products. Our own buyers source the products from the market, so a manufacturer cannot submit its own washing machine and knows neither what is being purchased nor when. For endurance tests, a single unit is often insufficient, so multiple units are procured to test performance over time.
How can subjective criteria such as ease of use or feel be evaluated in a comparable way?
Through blind tests, fixed test groups, and reference units running alongside them. Product names are covered up, units are made unrecognizable, and testing is conducted in varying orders. No single person evaluates a device: All testers take turns testing different devices, and their impressions are collected and consolidated. Reference devices maintain a consistent standard, even as the products change.
Why isn’t a single expert sufficient for a reliable test verdict?
Because a test requires multiple perspectives: project management, market analysis, the testing lab, and the editorial team, which knows what matters to readers. Raw measurement data alone doesn’t tell a story, and even the sharpest analysis won’t reach anyone if it remains incomprehensible. During the verification process, each tester notices something different, which strengthens the test.
What happens if a device breaks during testing?
Retesting is conducted until the result is confirmed. A finding is only considered reliable if it is reproducible. That’s why multiple units are procured for stress tests. Before publication, a central verification step additionally checks all measurement data for errors and consistency.
Do manufacturers find out about poor test results before they’re published?
Yes, they receive the objective measurement data, including critical findings, but they have no influence over the result and are not informed of either the date of purchase or the publication date. In the case of child car seats that came loose from their mounts, the issues involved goodwill, material defects, or batch problems. Reactions range from cooperative dialogue to injunctive relief.
How can readers understand how a test result was arrived at?
Through the disclosed methodology. The magazine features a concise box titled “How We Tested It,” followed online by a more detailed version with further information and references to the underlying standards. This transparency is also an obligation: Compliance with our own rules must be verifiable; otherwise, we cannot expect readers to place their trust in us.


