Skip to main content

Search...

Form Testing: Model-Based Test Data via SMT Solver

Form testing for over 255 tax form variants across 16 states: an SMT solver derives valid test data from the model and exposes contradictory rules.

• • Updated: • 11 min read
Cover of the expert talk on 'Form Testing: Model-Based Test Data via SMT Solver' with Simon Bergler, Lilia Gargouri and Richard Seidl.

Model-based form testing is an approach in which the data and UI models of a low-code platform automatically produce test data sets and UI locators. An SMT solver calculates valid test data sets from the validation and calculation rules. Code generated from the models then fills in forms with hundreds of fields and thousands of rules fully automatically, instead of relying on manual testing or individual tests that are expensive to maintain.

Key Takeaways

  • More than 255 form variants across 16 German states, each with over 100 fields in the UI and some forms with up to 100,000 fields, make manual testing practically impossible.
  • A test data generator based on an SMT solver detects contradictory validation rules while generating data, which stops faulty models from being deployed in the first place.
  • Selectors and test code can be derived automatically from the UI model of the low-code platform, so a single line of code fills in an entire form and a second one checks it.
  • Testing that would take several hundred hours by hand runs automatically in twelve to fourteen hours per state.

Why Form Testing of Government Tax Forms Can’t Be Done by Hand Anymore

Government tax forms reach a level of combinatorial complexity that rules out manual testing. The project discussed here covers more than 255 form variants across 16 German states. Every state has its own variant, and every form has over 100 fields in the user interface.

For some tax types the numbers climb even higher. One tax form alone has around 2,400 fields and about 2,200 validation and calculation rules. One particular tax reaches up to 100,000 fields. At that size, no person can simply work through a form from top to bottom.

On top of that, everything keeps changing. The business logic changes every year and differs from state to state. Even a valid test data set, once you have found it, stops being valid as soon as a rule shifts. The effort grows in every direction at once.

“Even if I find a valid test data set that passes every validation, it keeps changing. The effort goes through the roof immediately, in every direction.”

(Lilia Gargouri)

The Maintenance Trap Hits Automated Tests Too

Simply starting to automate the user interface doesn’t solve the problem. It just moves it. With this many forms, even automating a single happy path leads straight into the maintenance trap. On top of implementing each test case, there are ongoing maintenance costs every time fields or rules change.

The second bottleneck comes even earlier: before a tester can check anything, they need a valid test data set. That step blocks manual and automated testing alike. Without valid data, no test gets off the ground.

For the project, the consequence was clear. Everything has to be fully tested within four-week sprints that end in a release, and consistent test coverage across the whole project gets harder with every new form variant.

Using the Models as the Source for Tests

The approach Lilia Gargouri and Simon Bergler developed starts with the models that already exist. The platform is built on a low-code platform in which all business logic is modeled first. Behind every form are two models, a data model and a UI model, each stored as a JSON file.

The data model holds the fields along with the validation and calculation rules. Case workers and business analysts capture the business logic in this model rather than implementing it in code. They define a string field for a name or a date field for a date, for example, and write rules in the platform’s own rule language.

A simple example: if a name is entered but no date of birth, a rule fires with an error message asking for the date of birth. This set of rules is exactly what makes the model the ideal source for testing. Anyone who tried to hand-code 2,200 rules instead would fail on implementation and maintenance.

How the SMT Solver Generates Valid Test Data from the Data Model

The team generates valid test data sets with high coverage automatically from the data model. To do this, the JSON file is converted into an arithmetic model. An open source SMT solver solves this system of equations and finds a test data set with as many fields filled in as possible.

The team extended the open source solver with project-specific validations, for instance for how the tax number is assigned. The result goes into an Excel file in which every column is a complete, valid test data set, ready for manual or automated testing.

This step has a side effect that catches bugs early. If rules in the model accidentally rule each other out, generation fails right away. The solver can’t find a solution, and the modeling error shows up before the JSON files are deployed.

When the business logic changes, the rules change, and the team simply regenerates the test data. The existing tests then run again right away.

The UI Model Supplies the Code for UI Test Automation

The second generic approach derives UI automation directly from the UI model. In this model, the modelers define which label a field has, how fields are grouped, whether a string field appears as a text field or a text area, and whether a choice is shown as a checkbox or a radio button.

The team takes everything it needs for UI test automation from this model. Filling in fields needs neither the validations nor the rules, because it doesn’t matter whether the entered data is valid. The selectors, or locators, for the test tool in use are built from the field definitions.

Each element is identified by the combination of field type and label, such as a text field labeled “Name.” A library provides the matching method for each field type to fill in or check the element. For a text field, it picks exactly the method that handles a text field.

Two Lines of Code for an Entire Form

For the tester, the effort shrinks to a few lines. One command essentially says “fill in the form.” The tool reads the Excel file with the valid test data sets and works through the form blindly from top to bottom.

A second line checks the entire form and confirms that it actually contains what is expected. Two lines of setup code cover a complete form, no matter how many fields it has.

This flexibility allows different test depths with the same code. The tester can run through the whole process: create a form, fill it in, complete it and archive it. Or they can stop at a specific point, use that run as setup and continue from there manually or with a targeted feature test.

How Much Time Model-Based Form Testing Saves

The test duration shows why manual testing is not an option here. Running all automated tests for a single state takes 12 to 14 hours in total. Multiplied by 16 states, that is the real testing effort.

Done by hand, the effort would be many times higher, an estimated several hundred hours, at least around 500. The automated tool tests quickly and consistently.

In a short time the team built 300 to 400 concrete tests, plus around 20 dedicated feature tests. Much of this gets reused, because the building blocks work across many forms.

The benefit goes beyond the testing role. When someone finds a bug, the developer first has to navigate to the affected area to look for the cause. Instead of spending two hours retracing that path by hand, the prepared navigation runs in eight to ten minutes.

Next Step: Checking PDF Content Automatically

After a form is submitted, the system creates a PDF, and that becomes the next thing to test. The end user is simulated along the happy path: the long form is filled in and submitted. Then comes a confirmation and finally a PDF to download, which contains sensitive data and must comply with legal requirements.

The approach is well suited to this check because the test data set used is known. The team knows exactly which fields it filled in and can compare the PDF content in minute detail. Testing PDF content is no longer web UI testing, though. It is a separate task that nobody wants to implement or maintain on their own.

That is why the team is working on a generic library for it, together with the subsidiary responsible for the test tool it uses. A solution like this is especially relevant for the insurance industry and for e-government.

Accessibility Becomes a Test Field of Its Own

Accessibility is turning into a major topic, in the public sector as well as in insurance. Tax forms that citizens use to file their income tax electronically have to be fully accessible.

The problem for testers: without accessibility expertise, you don’t know which rules to check or how to respond to a violation. Getting familiar with the standard accessibility library takes time the sprint doesn’t have.

The answer follows the same pattern as for testing forms. The team is building an accessibility library on top of the standard library with its checks and rules. A team of accessibility experts supplies the additional requirements, and together they defined user-friendly error codes. The tester essentially calls “check this screen” and gets readable run logs that mark in red where each error code applies.

Frequently Asked Questions

How many fields and rules can a single government tax form contain?

The sheer scale makes manual verification impossible. In the project described, there are over 255 form variants for 16 federal states, with each form containing more than 100 fields on the interface. One tax form alone has around 2,400 fields and about 2,200 validation and calculation rules; a single tax type can have up to 100,000 fields.

Why doesn’t UI test automation solve the problem of large form landscapes on its own?

Because it merely shifts the burden. Given this diversity of forms, even automating a single “happy path” leads to a maintenance trap: Implementing the test case incurs ongoing maintenance costs as soon as fields or rules change. The real bottleneck lies even before that, however, because without a valid test data set, neither manual nor automated testing can run.

How does an SMT solver help generate test data?

It calculates valid test data sets from the set of rules. To do this, the data model is converted into an arithmetic model; the solver solves this system of equations and searches for a solution with as many populated fields as possible. The result is exported to an Excel file, where each column contains a complete, valid data record that can be used for both manual and automated testing.

Can test data generation uncover errors in the business model?

Yes, and that is a significant side effect. If validation rules accidentally contradict each other, the solver cannot find a solution, and the generation process aborts. The modeling error thus becomes apparent before the model files are even deployed. Testing the model occurs as a byproduct of data generation.

What happens to the tests when the business logic changes every year?

The test data is regenerated, and the tests run again immediately afterward. This is precisely where the difference from the manual approach lies: A valid data set found once loses its validity as soon as a rule changes, and business logic changes annually and by state. The model-based approach makes this recalculation a routine step.

How much test code does a form with hundreds of fields require using this approach?

Two lines are all that’s needed for setup. One line populates the form by having the tool read the Excel file containing the valid data sets and process the fields from top to bottom. The second line checks the entire form against the expected values. The number of fields doesn’t matter because selectors and methods are derived from the UI model.

How much time is saved by automated form testing compared to manual testing?

All automated tests for a single state run in 12 to 14 hours. Manually, this would take an estimated several hundred hours, at least around 500. This also pays off when it comes to debugging: Instead of spending two hours manually navigating to the affected area, the prepared test run completes this process in eight to ten minutes.

Is it possible to perform accessibility testing without having accessibility expertise on the testing team?

Only with preparatory work. Without specialized knowledge, a tester won’t know which rules to check or how to respond to an error, and getting up to speed on the standards library takes sprint time. The team therefore encapsulates them in a separate library with user-friendly error codes that a team of accessibility experts helped define. The tester runs the test and reviews the results in the run log.

Share this page

Related Posts