Quality Function Deployment (QFD) is a matrix-based method that links customer benefits directly to software functions and test cases. It provides what LLMs can’t: cause-and-effect analysis. Teams that trace their tests back to measurable customer benefit need fewer tests and make better prioritization decisions.
Key Takeaways
- LLMs can’t perform cause-and-effect analysis, because the architecture of neural networks rules it out. That is why hallucinations are unavoidable.
- Quality Function Deployment uses a matrix to link customer benefits directly to specific functions and tests, which makes irrelevant test cases visible so they can be dropped.
- Tracing tests back to customer benefit cuts the number of tests considerably, because priorities follow directly from the cause-and-effect link between a function and what the customer needs.
- The main obstacle for QFD in software development is matrix size: thousands of user stories against thousands of test stories have only become practical to compute since the breakthroughs of around 2014.
Why LLMs Can’t Do Cause-and-Effect Analysis
Large language models recognize patterns, but they can’t explain how they arrived at a result. Quality Function Deployment (QFD) works the other way around: it makes cause and effect explicit. The weakness of LLMs sits in their architecture, not in the details of one particular implementation.
Neural networks have existed as a concept for about 80 years. They are good at recognizing things and poor at logic and at explaining relationships. People work in a similar way: we say something based on gut feeling and only afterward work out why we came to that conclusion.
That is the problem with the term explainable AI. A system built on a neural network can’t simply disclose how it reached an answer. Hallucinations are therefore not a bug you can avoid. They follow from the way these systems are built.
Thomas Fehlmann sees this as a fundamental limit, not something the next model will grow out of. If you need causal traceability, you have to get it from somewhere else.
How Quality Function Deployment (QFD) Works with Matrices
Quality Function Deployment, QFD for short, is a method that uses matrices to examine the links between requirements and functions. The goal is the greatest benefit for the least effort.
If you know how AI works, the principle will look familiar. Both use matrices to represent relationships between many variables. QFD works with measurable contributions and asks: what does it take for a function to support a customer benefit?
A coffee machine is a simple example. If you want a strong Italian coffee, the machine needs the right setting and the function behind it. If all it produces is weak, watery coffee, the customer is unhappy and won’t buy from that manufacturer again.
QFD originated in Japan and reached Germany through a few individual advocates. Volkswagen and Skoda used the method extensively, for example to find out that a car needs a place for the driver’s handbag. BMW uses similar methods under a different name. Much of this stays hidden as a trade secret.
Bringing Causality and LLMs Together
The practical appeal lies in combining the probabilistic logic of an LLM with the cause-and-effect logic of QFD. In theory this works, because both approaches rely on the same mathematical methods.
In this view, data is more than numbers and transactions. It represents knowledge that moves back and forth between objects or modules. These data flows are the way to bring cause and effect in, measurably and traceably.
A reasoning model gives hints about which causal path leads to a result. If causality checks were built in properly, hallucinations would disappear. Only then could an AI be certified for automotive engineering, for example, which is not possible with today’s systems.
The technical hurdle is real. Large, sparse matrices have only been solvable in practice since around 2014. That computing capability powers LLMs today, not QFD. It could be put to that use.
Prioritizing by Customer Benefit Cuts the Number of Tests
If you align tests with customer benefit, you need far fewer of them. Customer benefit runs through every element of the software and gives you a clear yardstick for which test cases really matter.
Tests are expensive, even with today’s AI support. At the same time, test coverage barely pays off in the market. A well-tested car doesn’t sell for more than a poorly tested one. In theory, thorough testing could set a product apart from its competitors. In practice nobody does it, often because of time pressure before the next release.
Prioritizing by customer benefit changes what your tests focus on. If you keep aligning test cases with the benefit, they become more focused over time. You separate good test cases from the ones that add little information.
Security is the exception. Security and privacy are a given, with no compromises. Customers never state this requirement explicitly, but they are entitled to expect it. Anyone who gets into a car assumes it will drive and brake reliably.
Personalized Tests Instead of Mass Testing
Tests can be tailored to individual users. Not every function in a software release matters to every driver, and many are never used at all.
If the machine knows which functions you actually use, it can test exactly those. A new release could run through a test series at home in your garage that covers only the functions that matter to you.
That doesn’t fit the classic idea of mass production. It fits an Industry 4.0 that produces for the individual. The yardstick shifts from general coverage to personal relevance.
Why QFD Is Rarely Used in Software Development
QFD is unpopular in software, even though everybody talks about customer benefit. There are two reasons for that.
The first is a side issue, but a powerful one. QFD requires functional models, and developers hate functional models. Managers keep trying to derive a pay scale from function points, along the lines of more functions, more money. That is not how software development works.
The second reason is matrix size. Even comparing 20 user stories with 25 test stories pushes the method to its limit. Real projects look different: 1,000 user stories and 5,000 test stories no longer fit on a matrix like that.
“If I use customer benefit as the reference for which tests really matter to the customer, I end up with far fewer tests to run.”
(Thomas Fehlmann)
The computing power for large matrices has existed since 2014. Today it goes into LLMs, not into Quality Function Deployment. Thomas knows of academic research on QFD only in Aachen and Stuttgart, for example in building small electric cars. Even there, funding remains an open problem.
Transfer Functions in Six Sigma and Software: Measuring Causes You Can’t See
A transfer function describes how an effect results from a cause. It is the common tool behind Six Sigma, behind software functions and behind learning systems.
Music on your phone is an everyday example. A transfer function turns the digital information in an MP3 or MP4 file into sound. The same logic applies on a cosmic scale: you can’t measure exoplanets directly, only their effect. To detect them, you need to know the laws of gravity, the law behind the effect.
Software is no different. You need to know the laws by which functionality gets into the product. In Six Sigma projects, transfer functions help minimize variation in production.
With learning systems, things get uncomfortable. Whether it is the training, the original data set or reinforcement learning, nobody knows exactly what is going on inside. You have to find the causes and then, with luck, you can observe the effect.
Customer Benefit Has to Be Drawn Out, Not Guessed
Customers tell you what delights them and what doesn’t, but the useful information lies in a cause-and-effect analysis. The Net Promoter Score gives you the signal. The analysis has to explain why the score turned out the way it did.
That analysis is the basis for the next feature list. Over the years, two small companies that started with a handful of people have grown into larger businesses this way, one in color quality and one in customer communication. In both cases, QFD helped them build the most important thing next.
Working out the causal explanation takes effort, because raving about customer benefit is easier. But thinking it through forces you to see what matters, and sometimes your own mistakes too. What looked easy often isn’t.
Frequently Asked Questions
Are hallucinations in large language models a bug that can be fixed?
No. They stem from the design of neural networks, a concept that has existed for about 80 years: strong at recognizing patterns, weak at logic and explanation. Such a system cannot readily reveal how it arrived at an answer. That is why the term “explainable AI” is misleading, and anyone who needs causal traceability must obtain it separately.
Where is Quality Function Deployment (QFD) used outside of software development?
QFD originated in Japan and made its way to Germany through individual pioneers, where it was primarily used in the automotive industry. Volkswagen and Škoda made extensive use of the method, for example, to identify that a car needed a place for the driver’s handbag. BMW uses similar methods under a different name. Much of this remains unpublished as a trade secret.
Can the probabilistic logic of an LLM be combined with a cause-and-effect analysis?
Theoretically, yes, because both approaches use the same mathematical methods. A reasoning model provides clues as to which causal path to follow. If the causality check were properly integrated, hallucinations would be eliminated. Only then would certification (for example, in automotive manufacturing) be conceivable, something that is not possible with existing systems.
Is a high test density worth it as a competitive advantage?
Test density hardly pays off in the market. A well-tested car cannot be sold for a higher price than a poorly tested one, even though testing remains expensive, even with AI support. Theoretically, test density would be a distinguishing feature, but in practice, no one uses it, often due to time pressures leading up to the next release. This shifts the focus from the quantity of tests to their relevance.
Does prioritizing based on customer benefit also apply to security tests?
No. Security and privacy are non-negotiable and are not weighed against customer benefit. The customer never explicitly states this expectation, but is entitled to expect it: Anyone who gets into a car assumes that it will drive and brake with reliability. Benefit prioritization applies to functional testing, not this foundational level.
How many requirements can still be meaningfully compared in a QFD matrix?
Even 20 user stories versus 25 test stories push the classic method to its limits. Real-world projects with 1,000 user stories and 5,000 test stories no longer fit into such a matrix. Large, sparsely populated matrices have only been practically resolvable since around 2014. This computational capability is inherent in LLMs, not in QFD, but it could be repurposed.
What is the purpose of transfer functions if the cause cannot be measured directly?
A transfer function maps how an effect arises from a cause, thereby making the underlying principle usable. Exoplanets cannot be measured directly; only their effects can be detected through the laws of gravity. A more everyday example is the conversion of digital information from a music file into acoustic data. In Six Sigma projects, transfer functions are used to minimize variation in production.
Is the Net Promoter Score sufficient to determine customer value?
No. The Net Promoter Score merely provides the signal; the useful information only emerges during the analysis, which clarifies why the rating turned out the way it did. This cause-and-effect analysis forms the basis for the next list of features. Two companies that started with just a handful of people, one in the field of color quality and the other in customer communication, have grown by following this approach.


