In a data warehouse, quality assurance starts with understanding the requirements, long before the first test runs. A business analyst with a testing background spots edge cases early, pushes back on half-baked solutions from the business side and gives testers the domain knowledge they need to challenge the specification itself, not just the code.
Key Takeaways
- Business analysts who come from testing find edge cases and gaps in requirements sooner, because they are used to questioning specifications rather than taking them at face value.
- Testers should challenge the business analyst’s specification, because business analysts make mistakes too, and a tester without domain knowledge can only test the technical side.
- Synthetic test data sets, specified from the business perspective, cover new structures and data combinations more reliably than production data, which was never prepared for the changes under test.
- A temporary role swap, spending two to three months in another discipline’s job, builds mutual understanding and has a direct effect on the quality of everyday work.
Business Analyst vs. Tester: Who Decides What Is Right
Put simply, a tester checks whether something is right, and a business analyst defines what right should be. That is a crude simplification, but it captures the core of both roles.
People who move from testing into analysis bring a way of thinking that pure requirements work often lacks. Testers think in edge cases. They look hard at the spots where specifications break. That habit pays off when writing requirements, because it exposes gaps before anyone writes code.
Philipp Huber made that move himself. He went from software tester to business analyst on a data warehouse project at a large banking group, and he sees testing and analysis as two sides of the same coin. In agile teams, where work goes to whoever is available and T-shaped skills matter, it helps to know both sides: the methods, the tools and the product, seen from testing and from analysis.
Why Testers Should Question the Specification
Testers should not treat the business analyst’s specification as gospel. Yet that is exactly what often happens. A data mapping is specified a certain way, so the mapping is taken as given. When the output looks wrong, testers hunt for the bug in the code or in their own test case, never in the specification.
Business analysts get things wrong, too. That is why testers need to challenge the specification instead of testing against it blindly. A tester who skips this step is left with a purely technical test, and that is not enough.
To challenge anything, testers need domain knowledge. In a banking data warehouse, many fields look like a data graveyard: they exist, but nobody quite knows what they are for. Without understanding what the fields mean and how the entities relate to each other, nobody can judge whether a result is correct.
“It matters a lot to me that testers question the specification. But for that, they need to know the background. If it’s a data graveyard to me and I have no idea what the fields are for, I can’t challenge it.”
(Philipp Huber)
Challenging Requirements Early Keeps Baggage Out
The biggest lever on quality sits before the first line of code. Business units often don’t show up with a requirement. They show up with a half-baked solution they have already worked out, and it is frequently not what they actually need.
This is where the business analyst earns their keep. Instead of adopting the proposed solution, they ask: What exactly do you need? Forget what the solution should look like and tell me what you want to achieve. In a setup with two data warehouses, each with a different focus, this often shows that elaborate data flows across both systems are unnecessary and that much of it can be done more simply.
Skip this step and you end up stacking solutions on solutions. The requirement already is the finished solution, it gets built as is, and the next one is bolted on top. Over time the architecture turns into a tangle that follows no clear principle.
In a data warehouse, that hurts more than elsewhere. Whatever gets carried along once is almost impossible to remove later. Better to put in the effort at the start than to drag the baggage along for years.
One Concept, Twenty Views: Harmonizing Is Analysis Work
Different business units see the same concept differently, but the warehouse needs a single view. Customer classification is a typical case. One project alone can pull in twenty different ways of classifying a customer, from retail customers to large corporate clients.
Twenty parallel customer type attributes in the warehouse help nobody. The business analyst’s job is to sit down with the business units and harmonize these views as far as possible before they become a permanent structure in the warehouse.
How Requirements Work Plays Out in Practice
Large organizations often have no single, central requirements process. How requirements get captured depends on the department, sometimes better, sometimes worse. In practice, Excel lists of the required fields dominate, backed internally by homegrown tracking tools.
A group-wide business data model can provide orientation, but only if departments actually follow it. To some, it describes their business. To others, it is an add-on someone in the warehouse team dreamed up. In a large company, the result stays inconsistent.
Moving from classic waterfall projects to agile work changes the mindset, but it does not clear up that inconsistency on its own. Anyone who confuses agility with a lack of structure has missed the point. Agile is not a license for chaos.
Test Data for New Structures Is the Real Problem
New fields and new entities come without usable test data, because the structures are still being built. Production data is of little help for testing such changes, since it was never prepared for them.
The manual workaround, adding or editing individual records, costs effort every time and is hard to maintain. Try to rerun the same test cases six months later on a different, more production-like data set, and they no longer work cleanly.
On top of that, there is a visibility problem. A table with ten million records tells you nothing about whether it contains every data combination you need. With new data, that is especially hard to find out.
Business-Driven Test Cases That Run Anywhere
The most effective approach is isolated test cases, each covering a specific data combination, that can be deployed at any time. Instead of digging through a forest of millions of production records, you describe exactly the scenarios you want to check.
In one project, the team built this with Excel macros and Quality Center. The tooling was rough, but it worked. The workflow looked like this:
| Step | What happens |
|---|---|
| Enter the data combination | Data combinations are captured in an Excel sheet; related entities can be copied as a test set |
| Define the expected result | The same sheet records how a consuming application expects the data |
| Load the data | An Excel macro writes the entries straight into the staging area, from where they are loaded through the warehouse into the core |
| Compare automatically | Quality Center test cases check the expected result against the actual output at the end of the pipeline |
The approach automates two things at once: generating test data from manually specified test cases and checking the output. At its core, the question is whether the transformation is right, that is, whether data arriving in one structure shows up correctly in a different structure in the consuming system.
Excel has obvious limits as a foundation. Different versions of the same file float around on file shares and in emails, which is no long-term solution. The workflow itself holds up, though, and works just as well with better tools.
Acceptance Beats Elegance in Tooling
A test data approach is worthless if the business units won’t use it. That was exactly the strength of the Excel solution: the business side could create test data directly, because Excel is the tool they already use every day.
For someone who lives in Excel, formulas and copying test sets are second nature. Learning an unfamiliar specialist tool would be a hurdle. Excel is their comfort zone, and that drives acceptance.
The best tool is useless if people don’t understand it and don’t accept it. This solution was born out of necessity: the existing tools couldn’t keep up with so many new structures, and there was no time for a lengthy tool evaluation.
Combine Production Data and Synthetic Test Data
No single approach covers everything. What works in practice is a combination: specific test cases for the most important functionality, plus regression testing on production-like data.
Much of the work runs as regression on anonymized production data. A baseline of production-like data is compared against the previous release to catch unexpected differences. The catch is in the details: whenever a difference shows up, you have to know whether it was expected.
That is where testing with production data alone falls short. The data isn’t prepared for the planned changes. So you also need targeted data sets, including synthetic ones, that cover the core functionality, while regression makes sure everything fits together in the end.
AI as a Source of Ideas for Test Scenarios
Generative AI can suggest test scenarios, but it can’t hand them over ready to use. Describe your structures to a language model, ask for possible test cases, and you get useful ideas that work as a starting point.
A business analyst can pass concrete scenarios on to testers this way, without spelling out every edge case down to the expected result. That detailed work stays with the tester. The input sets the direction; the testing side does the elaboration.
Put Yourself in the Other Person’s Shoes
The most effective way to broaden your view of quality is a temporary role swap. Spend two to three months doing someone else’s job, and you will see things differently afterwards.
Nobody becomes a perfect tester or business analyst in three months. But you see what matters. It works both ways, and you can take it further: a business analyst who spends a few months with the business unit that depends on the warehouse comes back with a much better sense of why things are the way they are.
That understanding feeds straight into the quality of the work. People who stay curious and step outside their own box take more from unfamiliar tasks than someone who only does their own role as well as they can. That change of perspective sits at the heart of agile work.
Frequently Asked Questions
Why Is a Purely Technical Test in the Data Warehouse Not Enough?
Without domain expertise, a tester can only verify whether the code meets the specification. Whether the specification itself is correct remains an open question. In a banking data warehouse, many fields look like a data graveyard: they exist, but their purpose is unclear. Only those who understand the meaning of the fields and the relationships between the entities can truly scrutinize the mappings.
What are the benefits of understanding both testing and requirements analysis in agile teams?
Tasks are distributed as needed in agile teams, which is why T-shaping pays off. Those who understand methodology, tools, and the product from both perspectives can identify edge cases even while formulating requirements. Testing and analysis are two sides of the same coin: the tester verifies whether something is correct, while the business analyst defines what should be correct.
What happens when the business department delivers a ready-made solution instead of a requirement?
Then solutions pile up on top of solutions. The solution they bring is implemented, the next one is grafted on top of it, and eventually the architecture no longer follows any clear principle. It’s better to ask what the business department wants to achieve, rather than what the solution should look like. This often reveals that complex data flows across two data warehouses aren’t necessary at all.
How do you handle it when business units define the same concept differently?
The views must be standardized before they are incorporated into the data warehouse as a permanent structure. When classifying customers, a single project can quickly accumulate twenty variants, ranging from retail customers to large business customers. Twenty parallel customer-type attributes are of no use to anyone. Clarifying this with the business departments is an analytical task, not a technical issue.
Where does test data come from for structures that don’t yet exist?
From synthetic datasets specified from the business perspective. Production data is never prepared for planned changes, and a table with ten million records does not reveal whether all necessary combinations are present in it. Effective test cases are those that map individual data combinations, can be deployed at any time, and include the expected result of the consuming system.
Is regression testing on anonymized production data sufficient?
No. Comparing a baseline of production-like data against the previous release will find unexpected discrepancies, but for every discrepancy, the question remains whether it was expected. Because production data does not reflect the planned changes, targeted test cases, including synthetic ones, are also needed for core functionality. Regression testing then verifies that everything fits together in the end.
Can generative AI provide test cases for data warehouse testing?
Not on a one-to-one basis. If you describe the structures to a language model and ask for possible test cases, you’ll get useful suggestions that serve as a starting point. The detailed work (formulating every edge case with an expected result) is still up to the tester. The input sets the direction; the test team handles the elaboration.
What helps testers and business analysts understand each other better?
A temporary role swap: spending two to three months doing the other person’s job. No one will become the perfect tester or business analyst this way, but you’ll see what matters most to the other side. This works both ways and can be taken a step further: A business analyst could also spend a few months working in the business unit that relies on the data warehouse.


