Software testing as a discipline of its own means that specialized testers check programs systematically, separate from the development team. In this model, testers are paid for results rather than time: the number of test cases, the code coverage achieved and the defects found. What still matters most is that every component has clearly defined inputs and outputs that can be checked in a measurable way.
Key Takeaways
- Europe’s first independent software test lab opened in 1978 because Harry Marsh Sneed couldn’t find anyone willing to work purely as a tester.
- Harry Marsh Sneed refused on principle to pay testers by the hour: billing was based on test cases, program branches covered and defects found, which motivated the Hungarian testers directly.
- Agile development often fails in practice because components are cut too large to be truly finished and tested within one sprint.
- Static analysis alone can’t fully validate software, because the specification would have to sit at the same semantic level as the code itself.
- Project managers routinely underestimate testing effort: anyone who names a realistic figure risks being pushed out of the project, as the electronic health card shows.
Why an Independent Test Lab Was Something New in 1978
In 1978, Harry Marsh Sneed founded Europe’s first independent lab for software testing. It grew out of a simple observation: testing took more time than development. On a Siemens project for an interactive database query, a team of five needed three to four days to write the code and at least as long to test it. More than half of the effort went into testing. Why testing is so expensive became the question that shaped the following decades.
Harry set up the lab because he couldn’t find testers in Germany for a large project for the German Federal Railway. By the end, the project had around 200 modules and several hundred thousand lines of code. Something that size couldn’t be programmed the conventional way, like a small project. Testing needed an organization of its own.
The answer came from Budapest. Through a contact made at a testing conference in London, Harry found programmers there with a background in mathematics who were willing to do nothing but testing. Nobody in Germany wanted that job at the time. The profession of the dedicated tester barely existed.
Paying Testers for Results, Not for Time
A tester should be paid for what they deliver, not for the hours they sit out. That principle carried the whole setup of the lab. Neither Siemens nor the Hungarian institutes were billed by the hour. Billing was based on measurable results.
Three parameters defined testing performance:
- Number of test cases
- Branches covered in the program (coverage)
- Defects found
For the Hungarian institutes this was unfamiliar, because until then they had hired out their people to German companies by the hour. In the end they accepted the model. The testers themselves got a bonus for every defect they found, and it clearly motivated them.
In Harry’s view, motivation is still the core problem of software testing today, whether the testing is manual or tool-supported. If testers share in the results, you get better tests.
What Hasn’t Changed in Software Testing for Decades
Every test compares one thing against another. That principle hasn’t changed since the late 1970s, whether the test object is a batch program on a mainframe or a web service running in the cloud.
Every test needs a baseline: a description of what goes in and what should come out for that input. Without that defined expected result, there is nothing to test against.
Harry’s early tool, the test bench, was built on exactly this idea. It used an assertion language to define a test case: if a certain condition holds, say age greater than 90, then a certain output has to follow. The computer compared expected and actual values and listed the deviations. The testers then had to work out why each deviation occurred.
At the component level, testing is therefore a formal check of input and output parameters against the returned result. You don’t need domain knowledge for that. That changes in integration testing, once all modules are linked and the end result has to be checked against the business requirements.
Static Analysis: Finding Defects Without Running the Program
For years, Harry’s goal was to validate software without executing it. The idea: you feed in the source code, and the tool reports where it contains errors and contradicts the specification.
So the test bench had two parts. Dynamic analysis ran the program. Static analysis parsed the code, documented the architecture (who calls whom, who passes which data to whom) and looked for weak spots such as rule violations or poor structures.
Full static validation fails because of the effort. It requires the specification to sit at the same semantic level as the program. You would have to describe a separate test for every loop, every case decision and every if statement. That is so expensive that Harry settled for flagging major rule violations.
One recurring fight was over the GoTo. Until the late 1980s, many programmers still thought in assembler: check a state, jump, move data around in between. That way of thinking ran deep, and Harry spent a long time arguing against GoTo constructs.
Why Agile Projects Fail on Component Size
Agile sprints only work if the components delivered are small enough to be tested and usable within the sprint. In Harry’s experience, this is exactly where things break down. Teams try to finish components that are far too big in two to four weeks, and when time runs out, they close the sprint and hand over something that is nowhere near done.
The central question is what “done” actually means. How far does testing have to go before a sprint result is really finished? If you don’t answer that, you only appear to deliver.
“The architecture of the software has to fit the agile development theory.”
(Harry Marsh Sneed)
Harry has examined several failed agile projects, among them one for an oil company in Vienna and one for a hearing aid maker in Graz. Both ran well over budget and were never completed. His conclusion: they hadn’t really worked in an agile way. They had squeezed monster components into a single sprint.
His research at the University of Dresden looked at how big a web service can be. Working from the time needed to test it, he arrived at about 300 to 400 statements, depending on the language and complexity. A component module shouldn’t be bigger than that.
Measure Software Before You Test It
Before you test components, it pays to measure them. How big are they, and how complex? From those numbers you can decide the order in which components get tested and whether a component fits into a sprint at all.
This approach deliberately starts from the actual size of the building blocks. The common opposing view works the other way round: start at the top with the goals, derive the architecture from them, and let component size follow from the design. Harry counters that you spot problems earlier if you measure the programs directly.
The lesson from many projects is a sober one. Defining interfaces cleanly was always the hardest part, back then with software modules, today with REST and web services. The interface is everything.
Cut Testing Now, and the Project Fails Later
Squeezing the testing budget comes back to bite you. On the electronic health card project, Harry estimated several thousand person-days just to develop the tests, plus more days to run them. The project manager threw him off the project. The number was too high.
In the end the project failed because it was too buggy and hadn’t been tested enough, which is exactly what Harry had predicted. The mechanism is always the same: managers keep their eyes on the deadline, the deadline has to be met at any cost, and testing is the first thing to be cut.
Invest in proper testing early, and you pay once. Save on it, and you pay more later, often with the whole project.
Frequently Asked Questions
Why does testing often cost more than the actual programming?
Because the testing effort grows with the number of combinations, while the code can be written quickly. In a Siemens project for an interactive database query, a team of five people needed three to four days to program and at least as long to test. More than half of the total effort went into testing. This observation led to the establishment of Europe’s first independent test lab in 1978.
Is there a way to bill for testing work other than by the hour?
Yes. Harry Marsh Sneed did not bill his test lab’s work based on hourly rates, but rather on measurable results: the number of test cases, the program branches covered, and the bugs found. The Hungarian institutes, which had previously assigned their staff on an hourly basis, accepted this model. The testers themselves received a bonus for bugs found, which visibly motivated them.
What is the minimum required to be able to perform testing at all?
A baseline: a description of what goes in and what should come out for that input. Without a defined target specification, there is nothing against which to test. Every test compares one object against another, and this principle applies regardless of whether the object under test is a batch program on a mainframe or a web service in the cloud.
Do testers need domain-specific application knowledge?
Not at the component level. There, testing is a formal verification of the input and output parameters against the returned result, and the technical specification is sufficient for this. Domain knowledge only becomes necessary during integration testing, when all modules are linked together and the final result must be compared against the business requirements.
Can static analysis replace the execution of programs?
No. Complete static validation requires that the specification be on the same semantic level as the program: A separate test would have to be described for every loop, every conditional statement, and every if statement. This effort is practically impossible to justify. Sneed’s test bench therefore limited the static analysis to architectural documentation and noting major rule violations.
How large can a component or web service be?
About 300 to 400 statements, depending on the language and complexity. This figure comes from Sneed’s research at the University of Dresden and is based on the time required for testing of the component. A component module should not exceed this size, otherwise it cannot be fully tested within a single sprint.
Why do agile projects fail even though they work in sprints?
Because the components are too large. When time runs out, teams wrap up the sprint and deliver something that is far from finished. Sneed examined several failed projects, including one at an oil company in Vienna and one at a hearing aid company in Graz: both significantly exceeded their budgets. His finding was that massive components were crammed into individual sprints.
What happens when the testing effort is estimated realistically?
It’s often rejected as too high. In the case of the electronic health card project, Sneed cited several thousand person-days just for developing the tests, plus additional days for executing them. The project manager then threw him off the project. The project later failed due to too many errors. Managers keep an eye on the deadline, and testing is the first item where they cut corners.


