Skip to main content

Search...

Hybrid App Testing at ZDF: Manual First, Then Automation

Hybrid app testing at ZDF: 70 to 80% automation, and 90% of failed tests are real bugs. Manual and exploratory testing still catch what scripts miss.

• • Updated: • 10 min read
Cover of the expert talk on 'Hybrid App Testing at ZDF: Manual First, Then Automation' with Benedikt Broich, Anika Strake and Richard Seidl.

Hybrid app testing combines automated and manual tests to close gaps that neither approach covers on its own. Automation checks many devices and a lot of content systematically, while manual and exploratory testing catches visual defects, usability problems and device-specific edge cases. An automation share of around 70 to 80 percent doesn’t replace the manual part, it complements it.

Key Takeaways

  • Test automation without a manual share misses exactly the defects that only the human eye notices: overlapping elements, usability problems and quirks on specific devices.
  • A setup of 15 to 20 devices in test automation doesn’t cover all relevant OS variants, because manufacturers such as Samsung put their own layers on top of Android, which creates thousands of special cases.
  • Around 90 percent of failed automated test cases turn out to be real bugs in the software, not problems with the automation itself.
  • Automation breaks sooner than planned, because a single iOS update can knock out the tool in use, and manual testing covers exactly that gap.
  • Test automation decays when teams under deadline pressure put maintenance on hold: if a third of the tests are broken after a sprint and nobody fixes them, the suite loses its value.

Why Mobile App Testing Can’t Be Fully Automated

Hybrid app testing exists because full test automation of mobile apps runs into device diversity. At ZDF, the apps have to run on a wide range of devices, including old models, so that the content stays accessible to everyone. That range is exactly what makes pure automation expensive and fragile.

Anika Strake describes the core problem: different operating systems, different Android versions and manufacturer-specific interfaces, such as Samsung’s, create thousands of special cases. If you try to cover all of that with automation, you build up a maintenance burden that quickly gets out of proportion.

The result is a trade-off rather than an all-or-nothing rule. You have to decide what gets automated, on how many devices, and which gap remains. That remaining gap can’t be automated away. It can only be closed in a different way.

How a Tender for Pure Automation Turned into Hybrid App Testing

The original brief was: automate everything. ZDF’s tender said that only automated testing should be offered. Quality assurance had been neglected before, and automation seemed like the logical next level.

Appmatics won the tender and still took a different route, for practical reasons: the delivered software was poorly suited to automation. Tags and IDs, which make it easier to address elements automatically, were partly missing. For an app under continuous development, that drives maintenance effort up even further.

As a result, the kick-off phase dragged on. But the client wanted regular tests and regular results. So the team started manually and added automation later. What began as a stopgap became the guiding principle.

Manual and Automated App Testing Complement Each Other

The strength lies in combining both approaches. Automation reaches into many small corners of an app, and manual testing catches what machines are bad at judging.

Automated tests, for example, click through every program page in the media library and check whether content is there and links work. Doing that by hand would take forever. A person, on the other hand, spots usability problems, overlapping elements or hidden buttons much faster, and finds defects during exploratory testing that no script would have anticipated.

The basic setup includes 15 to 20 devices in test automation, and manual testing adds more devices on top. If a defect only occurs on a specific OS version, narrowing it down manually is often faster than digging through all the automated tests.

How to Decide Whether to Automate a Test Case

Criticality and frequency decide. Key features and happy paths belong in automation, because they must never break.

Manual testing itself gives you a second signal. If a test feels tedious and repetitive, that’s a first sign it should be automated. That’s exactly why it pays to start manually: you get a feel for the app and for the places where problems keep showing up.

Frequent defects in a particular area are the third criterion. If similar problems keep coming back, even outside the core features, that case moves into automation.

Benedikt Broich describes the typical sequence: in week one, a critical defect turns up during exploratory testing. In week two, it gets retested manually, and a test case already exists. If that test takes a lot of effort, or the defect comes back, the team adds automation as a safety net.

Automation Breaks, and the Setup Has to Cope with That

Mobile automation doesn’t stay stable on its own. It depends on many tools and vendors that have to work together. A single iOS update can be enough to make the automation tool stop working with the new version.

The manual side absorbs those outages. If an iOS version can’t be tested automatically at the moment, it simply gets covered manually. That makes the hybrid setup not a luxury but insurance against brittle tools.

The setup effort is high, but ongoing maintenance stays manageable. Problems on the automation side can usually be fixed within a day. Of the failed test cases, about 90 percent are real bugs in the software and around 10 percent are issues with the automation itself.

Results Are Verified Manually, Not Reported Blindly

Automated results go into the team’s own tool and then through a human review. A test analyst looks over the results and makes a first assessment: is the problem in the automation or in the app?

Then the defect is reproduced manually and narrowed down. Does it only happen on iOS or only on Android? Is it tied to specific OS versions? Is it a visual defect related to screen size or tablets?

For a media library, this verification matters a lot. If full-screen mode doesn’t scale properly or content is displayed incorrectly, those are classes of defects that automation alone struggles to find.

Capacity Freed Up by Automation Goes Back into Manual Testing

Automation covers around 70 to 80 percent of the cases. Even so, manual testing isn’t scaled back.

The capacity freed up by automation goes back into manual testing: more granular test cases, more exploratory testing, more devices. So the automation share grows because new cases are added, not because manual tests are cut.

That sets this approach apart from the usual back and forth over resources. Only at the very beginning do automated tests replace manual ones. Beyond a certain point, manual testing stays in place as permanent support.

How to Handle Test Data That Keeps Changing

With dynamic content such as a media library, generalization works better than fixed test data. Programs change all the time, content expires, new episodes are added.

The automation handles this with a data provider. The script scans the app, finds the available programs and runs the same test on each one until no new program turns up. That way a single test can be rolled out across all programs.

This generalization has a deliberate limit. The automation doesn’t check whether a title is worded correctly or contains a typo, because the effort would blow the budget. Content judgments like that are left to manual testing, which is part of the setup anyway.

Why Test Automation Breaks Down Under Deadline Pressure

The most common mistake is handing test automation to the development team as a side job. The expectation is that developers write their automated tests and everything takes care of itself. In reality, the opposite happens.

When the next deadline comes, testing gets pushed to the back. Suddenly a third of the tests are broken, and in the next sprint there’s no time to fix them. Because automation needs a lot of maintenance, especially when several teams develop in parallel and release candidates get merged, tests inevitably break.

The countermeasure is to reserve a fixed allocation, whether that’s a full position, a half position or a share of each sprint. Only then does quality assurance stay both flexible and functional.

“In the end, nobody is helped if the feature is done by the deadline but broken.”

(Benedikt Broich)

Defects Show Up at Every Level

No single test level finds most of the defects. The biggest problems are spread across manual, automated and exploratory tests.

A broken player where full-screen mode no longer works is more likely to be caught manually. Content that doesn’t load is found by automation, right where it happens. Strange defects on rare edge-case devices often only show up when someone follows a problem in an exploratory way.

That distribution is what confirms the hybrid approach. Only the combination of all three levels makes it possible to find so many defects in the first place.

Frequently Asked Questions

What bugs go undetected when apps are tested exclusively using automation?

Primarily visual and device-specific bugs. A human can detect overlapping elements, hidden buttons, and usability issues much faster than a script. Added to this are edge cases specific to individual devices, as well as error categories such as full-screen mode that doesn’t scale correctly. Exploratory testing also uncovers bugs that weren’t anticipated in any test case.

What makes testing Android apps so time-consuming?

Manufacturer-specific user interfaces. Vendors like Samsung overlay their own layers on top of Android, and on top of that, there are different Android versions and older devices that still need to be supported. This results in thousands of edge cases. A setup of 15 to 20 devices in test automation does not cover this range, and attempting to cover everything drives maintenance costs out of proportion.

Is test automation still worthwhile for an app that lacks tags and IDs?

Yes, but not as a starting point. Without defined tags and IDs, it’s difficult to target specific elements, and as development continues, the maintenance effort increases even further. In the ZDF project, Appmatics therefore started manually and introduced automation later, so that the client could receive regular test results despite the lengthy kick-off phase.

Should a test project start manually or with automation right away?

Starting manually has a practical advantage: you get a feel for the app and for the areas where problems regularly arise. If a test feels tedious and repetitive, that’s an indicator for automation. Key features and happy paths should be automated early on anyway, because they absolutely must not fail.

How reliable is a failed automated test case as an error report?

About 90 percent of failed test cases are actual bugs in the software; about 10 percent are issues with the automation itself. Nevertheless, reports aren’t filed blindly: A test analyst first assesses where the error lies. The issue is then manually reproduced and narrowed down, for example to a specific operating system, OS version, or screen size.

Does hybrid testing save on staff as the automation rate increases?

No. While automation accounts for about 70 to 80 percent of cases, manual testing does not decrease. The capacity freed up is reinvested in more granular test cases, more exploratory testing, and more devices. Only at the very beginning do automated tests replace manual ones; after that, manual testing remains an integral part of the process.

How does test automation handle content that constantly expires and is added?

Through generalization rather than fixed test data. A data provider scans the app, identifies the existing items, and runs the same test on each one until no new ones appear. This allows a single test to be rolled out across all content. Content-related evaluations (such as spelling errors in titles) remain the responsibility of manual testing because the effort required for automation would be too high.

Why shouldn’t test automation be handled on the side by the development team?

Because when under deadline pressure, it’s the first thing to get left behind. If a third of the tests are missing after a sprint, there’s no time for fixes in the next one, and the test suite loses its value. Automation requires a lot of maintenance, especially when multiple teams are developing in parallel and release candidates are being merged. It’s helpful to have a fixed allocation: a full-time position, a half-time position, or a portion of a sprint.

Share this page

Related Posts