DEFOSPAM is an acronym for structured requirements analysis that bundles seven dimensions to check: Definitions, Features, Outcomes, Scenarios, Predictions, Ambiguity and Missing elements. The goal is to uncover gaps, contradictions and ambiguities in requirements documents early. AI can serve as an idea generator, but it doesn’t replace human review.
Key Takeaways
- The acronym DEFOSPAM stands for Definitions, Features, Outcomes, Scenarios, Predictions, Ambiguity and Missing elements, and it is used to question requirements systematically.
- Unresolved definitions such as “customer” or “order” lead developers and testers to understand the same document differently, so defects only surface late in the project.
- Requirements act as an oracle: if you can’t derive a prediction from a requirement about what the system should do with a given input, you can neither test nor build it.
- AI works well as an idea generator for scenario-outcome pairs the team overlooked, but it is no reliable substitute for human review, because it hallucinates and misreads questions.
- Using AI as a coach rather than a mentor means not trusting its answers to be complete, but treating its questions and suggestions as prompts for your own thinking, which you then evaluate yourself.
Requirements Analysis Means Asking the Right Questions
Good requirements analysis doesn’t stop at reading a document. You hold it up against a fixed set of questions and question it systematically. That is the idea behind DEFOSPAM, an acronym Paul Gerrard came up with more than ten years ago and wrote up in a slim book with Susan Windsor.
The name happened by accident. Paul sketched out the method in an hour, looked at the initial letters and got stuck on a word nobody forgets. The book never sold, and hardly anyone used it. Still, the approach follows a logic that is useful for every tester and every developer.
The background is a classic one. Requirements often read smoothly and look complete, but they aren’t. Examples only ever describe software partially. You would need infinitely many examples to specify a system fully and infinitely many tests to evaluate it fully. Specification by example is therefore only a partial solution.
What the Letters in DEFOSPAM Stand For
DEFOSPAM is a series of review questions for examining requirements from different angles. Each letter marks its own direction of search. Only the “De” at the start stands jointly for Definitions, which is why eight letters give you seven directions.
- De for Definitions: clarify terms such as customer, product, order or defect before you go any further.
- F for Features: what can the system do, and how can that be grouped into functions?
- O for Outcomes: what does the system actually do with a given input?
- S for Scenarios: which situations lead to which outcomes?
- P for Predictions: what behavior can you derive from the requirement?
- A for Ambiguity: where is the language or the logic ambiguous?
- M for Missing elements: which scenarios, data or rules are missing?
S and O belong closely together. An outcome without its scenario makes little sense, and neither does a scenario without an outcome. What works are scenario-outcome pairs following the pattern: if this, then that.
Why Definitions Come First
Definitions are the starting point, because most misunderstandings come down to unclear terms. People read a requirements document, pick out ideas and never question the words.
Years of practice have shown that the conversations you get when you really dig into a definition are often the most valuable ones. What exactly is a customer? What is an order? What is a defect? Testing itself knows the same fuzziness. Everyone carries around a favorite definition that the rest of the industry doesn’t share.
Agreeing on a set of shared definitions sounds boring and time-consuming. It is. But afterward your message gets through, and you understand the other person’s message.
Why Requirements Act as Your Test Oracle
The prediction is where a requirement fulfills its real purpose: it serves as an oracle. If you enter A, B and C, the system should do D, E and F. If you can make that prediction and read it out of the document, the requirement works as a source of knowledge.
Both sides use the same oracle. The tester derives expected results from it. The developer derives what to build.
A perfect oracle doesn’t exist. It would have to be a huge document written in mathematical logic. In practice, you work with compromises and trade-offs.
This is where the tester’s mindset becomes productive. You can imagine a sensible outcome, but you can’t find it in the requirement. Say a document describes order processing with all its rules but says nothing about invalid values. The requirement only knows “rejected.” Being rejected is one thing, but under certain circumstances a value might be acceptable and under others not. If those rules aren’t in the document, the tester doesn’t know what to test and the developer doesn’t know what to build.
“By acting like a tester and looking critically at the requirements, you help everyone: the developers, the users and yourself.”
(Paul Gerrard)
Ambiguity Hides Between the Pages
Ambiguity shows up on several levels, and the most dangerous kind isn’t in the individual word. At the language level, the problem is still easy to spot. “Customer” can mean a new or a regular customer, a private or a business customer, a high-value or a low-value one. Without narrowing it down, the term stays vague.
Structural contradictions spread across the document are harder. Two very different scenarios lead to the same outcome, one on page 23, the other on page 38. Nobody noticed the connection because the passages are so far apart.
The reverse case is just as tricky. The same scenario produces different outcomes depending on where you look. On page 23 the system behaves one way, on page 43 another. Both can’t be right. Finding contradictions like these means keeping the whole document in your head at once, which is real work. The time can be well spent.
Missing Elements Are the Hardest to See
Missing elements are the hardest category, because it’s difficult to picture what isn’t there. A certain combination of inputs doesn’t appear anywhere. A data field is missing, such as the user’s age, which would actually matter in one situation.
Requirements can never cover every eventuality. Otherwise the document would grow without limit. So the point isn’t completeness but prioritization: the most important missing things stay the most important, and that is exactly where you need to focus.
This is where an outside nudge is valuable. Experience tells you that a certain situation needs handling. But no experience is comprehensive, complete and all-knowing. A broader range of examples covers more than the memory of a single team.
AI as a Coach in Requirements Analysis
AI can be a useful assistant when you work through this list of questions, without replacing the work. Paul is currently testing this in experiments and building a prototype to find out what’s possible.
The first observations are concrete. Give the AI a text, and it identifies named entities, meaning nouns, actions and activities, fairly reliably. It suggests usable definitions, maybe not perfect but good enough to start. It generates scenario-outcome pairs and matches them up.
A requirement that looks good yields surprisingly little material. It assumes a lot of implicit knowledge that a human would fill in but the AI doesn’t have. Ask for more, and a handful of pairs quickly grows. For an e-commerce application, the model knows there are stock limits, price calculations and discounts, and it suggests matching scenarios. Five examples turn into eight more without much effort, and most of them make sense.
You can also ask the system directly: about linguistic ambiguities, inconsistent scenario pairs, missing situations or data. It returns a list of what it finds.
Trustworthy, No. Helpful, Yes.
For reviewing requirements, AI is a less-than-perfect assistant you shouldn’t trust blindly. Its ideas aren’t guaranteed to be sound. You need to review them, keep an eye on what the model is doing and make use of every good suggestion.
Its strength is memory, not thinking. The model remembers everything, calculates reliably and, thanks to its training material, knows more cases than any single team member. It isn’t the best member of the team. It just has the better memory. It can hardly think, debate or truly challenge anything.
Its mistakes resemble human ones. It exaggerates, hallucinates, misunderstands one question and answers another. That calls for supervision.
The more productive role is coach, not mentor. A mentor would have to know everything about everything and pull you out of the mess. A coach asks good questions, makes suggestions and gets you to think more deeply. In that role, the tool gets you maybe 80 percent of what a team would achieve on its own with more effort. That moves you forward, but it doesn’t solve the whole problem.
The same goes for generated code. Few people trust it blindly, but it can deliver a large part of the work. If it follows your instructions and you know what it does, you save time. You still have to test it, or at least review it.
Frequently Asked Questions
Why Are Examples Alone Not Enough to Specify Software?
Examples always describe a system only partially. A complete specification would require an infinite number of examples, and a complete evaluation would require an infinite number of tests. Specification based on examples therefore remains a partial solution. It aids understanding but does not replace the systematic examination of definitions, rules, and missing scenarios.
Why should scenarios and expected results always be documented together?
A result without an associated scenario makes little sense, and a scenario without a result makes just as little sense. The information only becomes useful when presented as a pair following the pattern: if this, then that. Such scenario-result pairs are the core of a testable requirement because they can be used to derive both test cases and implementation decisions.
Why are vague terms like “customer” or “order” dangerous in requirements?
Because developers and testers read the same document differently without realizing it. People pick out ideas from the text and never question the words themselves. The error only becomes apparent late in the project. Agreeing on common definitions is tedious and boring, but once that’s done, your message actually gets across.
How can you tell if a requirement isn’t testable?
When no prediction can be derived from it: If I enter A, B, and C, the system does D, E, and F. If this deduction is missing, neither the tester knows what to test nor the developer knows what to build. A typical example: A document describes order processing but, in the case of invalid values, only specifies “rejected.”
What form of ambiguity is most often overlooked during a review?
Structural ambiguity, which is spread throughout the document. Linguistic vagueness in individual words is still noticeable. More serious is when two very different scenarios on page 23 and page 38 lead to the same result, or when the same scenario is handled differently on page 23 than on page 43. Both cannot be correct.
Does a requirements specification have to be complete?
No. Requirements can never cover every eventuality; otherwise, the document would grow to an unimaginable size. Instead of completeness, prioritization is key: the most important missing elements remain the most important, and that’s what you focus on. Since no one’s experience is all-knowing, a broader selection of examples is more helpful than the memory of a single team.
What exactly does a language model do when reviewing requirements?
It recognizes named entities such as nouns, actions, and activities quite reliably, suggests useful definitions, and forms scenario-outcome pairs. When prompted, the output grows rapidly: five examples easily yield eight more, most of which make sense. For an e-commerce application, for example, it generates inventory limits, price calculations, and discounts.
Can AI replace human review of requirements?
No. Its strength lies in memory, not in thinking: It remembers everything and knows more cases than a single team member, but it rarely debates or questions. It exaggerates, hallucinates, and sometimes answers a different question than the one asked. In the role of a coach who asks questions instead of providing answers, it achieves about 80 percent of what a team would work out on its own.


