Skip to main content

Search...

Mobile App Testing: Second Attempt with Robot Framework

Mobile app testing, second attempt: HUK Coburg scrapped four years of automation. Now manual testers write keyword-driven tests with Robot Framework.

• • Updated: • 13 min read
Cover of the expert talk on 'Mobile App Testing: Second Attempt with Robot Framework' with Felix Doppel, Alicia Heymann, Christoph Singer and Richard Seidl.

Mobile test automation works when it pursues two clear goals: non-technical testers can write test cases themselves, and device diversity is covered systematically. A keyword-driven approach with Robot Framework makes both possible. Just as important is building a stable infrastructure before the number of test cases starts to matter.

Key Takeaways

  • Test automation that creates more work than it saves, and that nobody can trust, is better shut down completely than kept running.
  • Before the second attempt, Felix Doppel and his team drew up a requirements catalog with weighted criteria and asked for a proof of concept with at least three test cases before hiring a service provider. A package of around 30 test cases followed as a stress test.
  • Keyword-driven testing with Robot Framework lets the manual testers at HUK-Coburg write and run test cases themselves, without any development skills.
  • More testers can’t solve Android device diversity. Only a mobile device cloud with physical devices, grouped into priority packages based on real user data, can.
  • Isolated test automation teams fail, because working test cases only come about when the domain knowledge of manual testers and the technical skills of automation engineers come together.

Why the First Attempt at Mobile Test Automation Failed

Test automation that creates more work than it saves has missed its purpose. That is exactly what happened with mobile test automation for HUK-Coburg’s telematics app. The team started automating in 2018, shortly after development began, without first thinking through the tool, the approach or the goals.

The result was green test runs that nobody trusted. While the automated suite passed from start to finish, manual regression testing found 20 to 25 critical bugs. So the tests were checking something, just not what mattered.

Both promises of automation fell away. It was supposed to take load off manual QA and provide confidence. Instead, it tied up extra capacity and gave no reliable signal. Felix Doppel, a tester at HUK-Coburg, puts part of this down to maturity: the team started too early, before it had worked out its roles and responsibilities.

Pulling the Plug Is Sometimes the More Honest Move

Keeping a failed automation effort alive just because a lot of money has gone into it is an expensive reflex. The team chose differently and shut the automation down completely in mid-2022. Four years of work were deliberately brought to an end.

The trigger was a sober cost-benefit analysis. In an agile project, what counts is not having automation to show off, but automation that actually delivers. Because the manual testers were strong, the team was simply better off without it.

That wasn’t easy. Management, predictably, asked about the money already spent. The discussions were tough, and it hurt to admit that the investment hadn’t paid off at that point.

After the hard cut came a deliberate break of about four to five months with no automation work at all. The team used that time to work through the failure instead of jumping straight into the next attempt.

Goals First, Then Tools: The Second Attempt

The second attempt started with a requirements catalog instead of a tool. That goes against the agile reflex of trying things out quickly, but given the scope and cost of test automation, it was necessary. This groundwork was exactly what the first attempt had lacked.

In internal workshops, the team asked each role directly: What does a product manager need, what does a test manager need, what do the manual testers need, what does development need? Some of the answers turned out to be competing goals that had to be reconciled and redefined. In one workshop, the team even had several groups sketch out rough solutions from scratch.

In the end, there were two clear goals. First, a non-technical user should be able to specify and run test cases themselves, so that manual QA could take an active part. Second, device coverage should grow, to protect quality.

The target for the level of automation was deliberately modest. Rather than aiming for as much as possible, the team set its sights on 60 to 70 percent, but stable and on around ten different devices.

Why Device Diversity Becomes a Risk in an Insurance App

In the telematics app, every functional bug has an insurance product attached to it, and that turns device diversity into a real risk. The app uses a sensor to record driving behavior and evaluates speed, acceleration, braking and steering. Drive safely, and you pay a lower premium.

If trip recording stops working, users get in touch right away. If recording fails for about three weeks, it escalates through customers and support all the way up to the department heads.

iOS is still manageable: one new model a year, few screen sizes, and usually the latest operating system. Android is a different story. There are hundreds of combinations of devices, models and operating system versions. You can’t test that many by hand, no matter how good your testers are.

Why Keyword-Driven Was a Better Fit Than Behavior-Driven

Which automation approach fits is decided by the team, not the textbook. The first attempt used Cucumber with Gherkin syntax, a behavior-driven approach. The second relies on Robot Framework and a keyword-driven approach.

The team didn’t know in advance that the keyword-driven route would work better for its people. That was something they could only find out by trying.

A common mistake: behavior-driven development is often introduced with great enthusiasm, and then every acceptance criterion gets squeezed into Gherkin without anyone asking who on the team can actually write it. Not every task can be expressed that way in a meaningful form.

Robot Framework had a practical advantage in this setting. Android and iOS share the same test flow and only differ at the lowest level. Above that, the business keywords keep everything the same, and test cases can be written in natural language and driven by test data, without anyone having to program methods.

How the Proof of Concept Went

A tool proves itself on your own test cases, not on paper. Instead of adapting to the technology as they had the first time, the team turned the order around and asked for real test cases first.

The external tender had two conditions: the requirements catalog had to be met, and there had to be a proof of concept with at least three submitted test cases. The team only gave the go-ahead after reviewing those test cases.

The contract went to the service provider imbus. Christoph Singer, a consultant at imbus, describes the starting point as unusually advanced, because the requirements catalog and test cases were already quite mature. Often this step has to be made up for at the client first, instead of simply grabbing the first tool a Google search turns up.

At the end of 2022, a larger package followed as a stress test: around 30 test cases within about two months. That showed whether the approach would hold up not just for three examples but at scale.

Quality Comes from the Whole Team, Not an Isolated Automation Team

From HUK-Coburg’s point of view, an isolated automation team that simply gets test cases thrown at it was one of the main reasons the first attempt failed. When automation engineers work in a silo and only receive finished test cases “over the fence,” there is no feedback on the domain side.

What makes the difference is combining strengths. An automation engineer is technically strong but often lacks deep enough domain knowledge. A manual tester understands the test case because they’ve run it a hundred times, but lacks the technical tools. Bringing those two sides together is what really makes it work.

“You have to bring the whole team along. If you isolate it, the test automation engineers feel isolated and they don’t support each other.”

(Felix Doppel)

For the manual testers, the keyword-driven approach wasn’t a big break, because their test cases were already structured along those lines. In a joint workshop, they wrote their first test cases, and working with the existing keywords, they quickly got a feel for what was already there and what was missing. When a keyword is missing, the testers pass the technical ball back to imbus.

Close communication mattered from day one. Interim results were presented regularly and checked for clarity, instead of unveiling a finished product at the end.

Where Mobile Test Automation Stands Today

About a third of the regression test suite is automated at the moment, some of it fully, some only partly. The regression suite has around 150 test cases. The team considers the 60 to 70 percent target achievable by the end of the year.

The sheer number of test cases deliberately isn’t what counts. Eighty automated but unstable (flaky) test cases would make a nice number for management and still be worthless. Stability beats quantity.

Fully automated means the manual test case can be dropped entirely. But a lot can only be partly automated. Some features stay out of scope: trip recording itself and pairing the telematics sensor can’t be tested automatically in any meaningful way.

Alongside the test cases, the surrounding infrastructure was built too: interface tests and the connection to a mobile device cloud. Maintenance is an ongoing cost as well. The 30 test cases from the end of 2022 already needed maintenance again in spring 2023.

Real Devices Instead of Simulators, Organized into Three Packages

The automation runs on physical devices in a mobile device cloud, not in a simulator. A simulator makes a lot of things easier, but side effects and reliable results only show up on real devices.

Device selection is prioritized into three packages:

PackageScopeExpectation
Priority 1Most-used devices (mainly Android)Tests must run here
Priority 2Devices that are still widely usedTests should run here
Priority 3Less commonly used devicesRun occasionally

The selection follows real usage data, not chance. In the background, the team monitors which devices users are on and reviews the data every month. That decides which devices land in which package.

An insurer’s security requirements come into play here too. Test devices are updated automatically, which is why there are no longer any physical test devices running Android 7, for example. The app supports Android 7 and up and iOS 14 and up, and older versions can be covered specifically through the cloud.

One practical advantage of the mobile device cloud is reporting. Every failed test case comes with a video recording that shows exactly where things went wrong. With the purely technical setup before, all you got was a stack trace, which ended up back with the developers.

Roll Out Slowly Instead of Losing Trust Early

The automated tests aren’t fully integrated into regression testing yet. They run in Jenkins, currently every night, but not yet automatically with every new build. That is a deliberate decision.

If the team switched the 50 to 60 or so automated test cases on for real right away, it would have to take them out of the manual test. It will only take that step once the surrounding infrastructure is fully in place, meaning the Jenkins and Jira integration is up and running.

The reason is the experience from the first attempt. If the test cases aren’t stable enough, the mood in the team can turn quickly again. That’s why the automation should only go into full use once everything around it is reliable.

How often to run the tests also takes judgment. Running them on every change drives up maintenance. Well-chosen times and releases with the right focus matter more than running constantly.

The automation is already paying off. The regular runs have found two or three critical bugs that would otherwise only have surfaced in regression testing. And finding bugs earlier comes down to money, once again.

Frequently Asked Questions

When does it make sense to completely discontinue an ongoing test automation process?

When it creates more work than it saves and no one trusts its passing results. For the HUK-Coburg telematics app, manual regression testing found 20 to 25 critical bugs, while the automation ran through the entire suite without issues. The team wrapped up the project in mid-2022 after four years, rather than continuing to chase after the money already invested, and took a break of four to five months.

What should be clarified before selecting a test automation tool?

The goals, specifically from the perspective of all involved roles. In internal workshops, product management, test management, manual testers, and development were surveyed separately; conflicting goals were resolved and redefined. In the end, two requirements emerged: Non-technical users should be able to specify and execute test cases themselves, and the variety of devices should increase. The target was 60 to 70 percent coverage across about ten devices.

When is a keyword-driven approach better than behavior-driven testing with Gherkin?

That depends on the team, not on the textbook. Behavior-driven development is often introduced with great enthusiasm, and then all acceptance criteria are crammed into Gherkin without asking who can actually write them. The keyword-driven approach with Robot Framework allows for test cases written in natural language and driven by test data, without anyone having to program methods. Android and iOS share the same test flow in this approach.

How can you evaluate a test automation service provider before hiring them?

Through real test cases, not presentations. The request for proposal set two conditions: the requirements specification had to be met, and there had to be a proof of concept with at least three submitted test cases. The contract was awarded only after these were reviewed. As a stress test, a package of approximately 30 test cases followed over about two months to see if the approach would hold up at scale.

Why can’t the wide variety of Android devices be covered with more manual testers?

Because on Android, there are hundreds of combinations of devices, models, and operating system versions. No team can manually perform testing on this many devices, no matter how skilled the testers are. iOS is much simpler: one new model per year, a few display sizes, and usually the latest operating system. The solution lies in a Mobile Device Cloud with prioritized device packages.

What criteria are used to select devices for automated app testing?

Real-world usage data. The team monitors which devices users are actually using in the background and evaluates this data monthly. This results in three packages: Priority 1 consists of the most frequently used devices on which the tests must run; Priority 2 includes devices that are still frequently used and on which the tests should run; and Priority 3 consists of rarely used devices that are included sporadically.

Are simulators sufficient for test automation for mobile apps?

No. Simulators make many things easier, but side effects and reliable results only become apparent on physical devices. A second advantage of the Mobile Device Cloud is reporting: For every failed test case, there is a video recording that makes it possible to trace the location of the error. With the purely technical solution used previously, developers were left with only a stack trace.

Should automated tests run with every new build?

Not necessarily. Running them with every change drives up maintenance costs; running them at strategic times and during releases with a specific focus yields better results. In the project described, the tests ran nightly via Jenkins, but not yet with every build. The team wanted to switch them on fully only once the integration with Jenkins and Jira was in place, because unstable test cases can quickly erode trust.

Share this page