Data testing for data processes means checking machine-generated output against an expected result derived directly from the business specification. The case distinctions in that specification form equivalence partitions, which produce test coverage automatically. Test data is the input the process runs on. Anonymized production data typically covers around 70 percent of the cases; the remaining 30 percent is added synthetically on purpose.
Key Takeaways
- Production data covers only about 70 percent of case combinations in testing, because rare edge cases simply never show up there and so go undetected.
- Business departments can only spot data errors in processing reliably when the results are shown visually and can be followed without any SQL knowledge.
- Expected results generated automatically from the business specification replace hand-built individual test cases and deliver higher test coverage for the same effort.
- The rule set for the data comparison can double as the detailed business specification, because it describes the mapping logic precisely and in terms non-technical people understand.
- AI is a useful assistant here, for example to generate SQL queries and show them visually, but the business department stays responsible for checking the content.
Data Errors Surface Too Late When the Business Side Can’t Reach the Data
In many companies, errors in data processing are found late because the people who understand the data and the people who can access it are not the same. IT has no trouble getting at the data technically, but lacks the business knowledge to see when the content is wrong. The business department has that knowledge, but no access. Data testing that starts early closes exactly this gap.
In practice this turns into tedious back-and-forth. The business department comes up with individual test cases, gets data delivered for them and checks them by hand. That covers only a fraction of the possible cases and takes a lot of time.
The later an error is found, the more effort it causes. That is why it pays to start testing data as early as possible, ideally while the processing logic is still being developed.
Data Quality vs. Testing Data Processes
Data quality and testing data processes are two different jobs. With plain data quality, you receive data without knowing how it was produced, so all you can do is check it for plausibility.
An example: in a list of living people, a birth date of 1810 is unlikely. You can’t go further than that plausibility check, because you know neither the input data nor the mapping rule.
When you test a data process, there is a business specification that says how input data is to be transformed into output data. That lets you build an expected result and hold it against the actual result. The question then becomes: was the output produced according to the business specification?
Generate Expected Results Systematically Instead of Building Single Test Cases
The more effective approach is to generate expected results systematically and compare them in full with the actual results, rather than defining individual test cases by hand. The basis is the set of business rules the department has to define anyway to describe the data mapping.
From those rules you can generate the expected result automatically and even run it against the complete production data set. Every row and every column gets checked.
The effort stays manageable. Anyone trying to reach the same coverage with manual test cases needs at least as long and still ends up with poorer test quality, simply because they can’t cover that many cases.
Why Production Data Is Not Enough for Data Testing
If you run automatic expected-result generation against production data, in practice it often covers only about 70 percent of the relevant cases. The missing 30 percent has to be added to get broad coverage.
The reason lies in the business logic itself. Case distinctions naturally form equivalence partitions. If there is no matching data for one of them, say the case “age over 90”, you have no idea whether the test object handles it correctly.
Gaps like that only show up when a real case comes along later and goes wrong. If you rely on production data alone, you simply never test those combinations. So the anonymized 70 percent from production is deliberately enriched with the missing combinations.
For ongoing testing, you then build a test set that covers every case combination with one representative each and runs quickly. The full data set is checked separately, more as quality assurance of the production data.
Replicating the Mapping Is Not a Double Implementation
A common objection to data comparisons is that you are building the system a second time and dragging in the same sources of error. That doesn’t apply here, because only the business mapping rule is replicated, not the system.
Most mappings are large but not complex. They consist of case distinctions: if this, then that, otherwise something else. Logic like that can be taken straight from the specification, without worrying about performance or other constraints.
The difference in effort is significant. Where generating the expected result takes a day, implementing the system keeps a bigger team busy for one to two weeks. You are not doing more work; above all, you are leaving a lot of work out.
“We don’t rebuild the whole system. We take the specification and use it to recreate the mapping rule directly. It’s not about exactly when what has to run, we only test the mapping.”
(Joshua Claßen)
Really complex calculations call for a different route. For a complex simulation, you get verified expected results from an independent source and compare them within a defined tolerance.
Data From Many Source Systems Has to Be Merged Before Testing
In heterogeneous system landscapes there is rarely just one input system. Several source systems feed in data that has to be harmonized and passed to a central interface.
One example is an anti-money-laundering check on transaction data. The input data from the various source systems is brought together and delivered to that interface. This merging step can be tested with the same principles as a single mapping.
Why Data Comparisons Need Tolerances
A data comparison is not always an exact one-to-one match. In some cases small deviations are acceptable from a business point of view and have to be tolerated through defined thresholds.
In a bank’s trading system, for example, cash flows are generated algorithmically. Because of numerical effects, the same result can differ by a cent or two depending on the order of the calculations. Deviations like that are harmless as long as they don’t build up beyond acceptable thresholds.
Tolerances aren’t limited to numbers. They can make sense for text fields too, for example when upper and lower case shouldn’t matter in the comparison.
Business Rules That Double as the Specification
The rule set used to generate the expected result can also serve as the detailed business specification. Test basis and business requirement then live in one artifact instead of in separate documents that drift apart.
A rule like that needs no technical background, just common sense. Here is an example of a naming rule:
- If first name and last name are both filled in, both are output separately.
- If only the first name is filled in, only the first name is output.
- If only the last name is filled in, only the last name is output.
- In all other cases, an error is output.
This pays off twice. When the business department sees the results of the defined logic early, it notices straight away if a case distinction is missing. An implementation can match the specification perfectly and still produce wrong results, because the specification itself has gaps.
Visualization Gets Business and IT on the Same Page
Data comparisons need to be visualized so the business department understands them and both sides work on the same object. A document written in technical terms keeps business and IT apart; a visible definition of the data comparison brings them together.
In data testing the same patterns keep coming back. Testing always means comparing an expected result with an actual one. Most data is tabular, but it can also be compared as XML or JSON. Because these patterns recur constantly, low-code components implement them faster and more simply than having every department program them from scratch.
That redundancy is exactly what you see at large banking clients with many departments. Each one builds its own CSV comparison tool, each one reinvents the wheel. Nobody writes their own word processor either.
Even people who know SQL benefit from a good visualization, because it is quicker to take in than raw code. If you want 99 of 100 columns, SQL makes you list all 99; in a graphical interface you just click the one column away.
AI Makes Data Testing Faster, Not Obsolete
AI doesn’t replace the work of data testing, it speeds it up. An AI system can make suggestions, such as setting case distinctions in intervals or drafting a test, but the result has to stay verifiable.
This is where visualization matters. If an AI only generates code, you have to be able to read code to verify its suggestion. If the result is shown visually, even someone without deep coding skills can judge quickly whether it fits and fix the rest themselves.
Sensitive company data sets a hard limit. That data doesn’t go to an external service on the internet. The trend is therefore toward more efficient language models that run locally and can turn a natural-language requirement into SQL and from there into a visual no-code query.
AI also works as a co-pilot for operating the tool itself. It can suggest how to use a tool or answer a question about the best way to solve a specific task. A fully autonomous “here is my system, test it”, on the other hand, is not coming any time soon.
Frequently Asked Questions
Why are errors in data processing often discovered so late?
Because in many companies, data access and domain expertise are disconnected. IT can access the data without any technical issues, but cannot identify content-related errors. The business department would have the expertise but lacks access. This creates a back-and-forth: The business department devises individual test cases, has data delivered, and checks it manually. The later the error is detected, the more expensive it becomes.
Is it enough for the business department to devise individual test cases for data processes?
No. Manually defined individual test cases cover only a fraction of the possible scenarios and are very time-consuming. It is more effective to automatically generate a target set based on the business rules that the business department must define for data mapping anyway, and then compare it comprehensively against the actual data. Every row and every column can be checked this way.
Is production data sufficient as test data for data processes?
No. In practice, it often covers only about 70 percent of the relevant case scenarios. The reason lies in the business logic: case distinctions form equivalence partitions, and for rare classes such as “age greater than 90,” there are simply no data records. The anonymized production data is therefore specifically enriched to account for the missing 30 percent.
Isn’t replicating the target logic a duplicate implementation of the system?
No, as long as only the business mapping rule is replicated and not the system itself. Most mappings are extensive but not complex: if this, then that; otherwise, something else. Performance and constraints are left out of the equation. Where generating the target result takes one day, implementing the system with a larger team takes one to two weeks.
Do the target and actual values have to match exactly when comparing data?
Not in every case. Minor deviations may be acceptable from a business perspective and are tolerated within defined thresholds. In trading systems, for example, algorithmically generated cash flows can result in differences of one or two cents, depending on the order of the calculations. The situation only becomes critical when such deviations snowball. Text fields can also have tolerances, such as for uppercase and lowercase letters.
What are the benefits of having the test rules serve as the business specification at the same time?
The test basis and business requirements are then contained in a single artifact rather than in separate documents that may differ from one another. The second benefit is more significant: if the business department sees the results of the defined logic early on, it can immediately identify any missing case distinctions. An implementation can be correct according to the specification and still produce incorrect results because the specification is incomplete.
Is a visual representation of data comparisons worthwhile even for people who are proficient in SQL?
Yes. A visualization is easier to grasp than plain code. If you want to query 99 out of 100 columns, you have to list all 99 in SQL; in a graphical interface, you simply uncheck the one you don’t need. Furthermore, the business unit and IT work on the same object: a document written in technical terms separates the two sides, while a visual definition connects them.
Can AI take over the testing of data processes?
No, it speeds up the work but does not replace it. AI can suggest case distinctions at intervals or design a test, but the responsibility for verifying the content remains with the business department. It is important that the suggestion remains verifiable: Pure code requires reading skills, whereas a visual representation does not. Furthermore, sensitive company data must not be sent to external services on the Internet.


