System testing commercial kitchen appliances means checking heating, motor control and bus systems fully automatically on real devices, backed by a dedicated test infrastructure. A Python framework controls hardware such as relay modules and power supplies. Around 2,000 test cases run every night, and on weekends a full run takes up to 30 hours.
Key Takeaways
- RATIONAL’s combi steamers contain around 700,000 lines of code, spread across several electronic components that talk to each other over bus systems.
- Ten test automation specialists at RATIONAL work alongside roughly 70 embedded developers, which makes scaling the automation infrastructure the biggest strategic challenge.
- Nightly runs across several branches cover around 2,000 test cases, and the bottleneck is not the framework but the physical time the real devices need to heat up.
- Where heating cycles would take too long, the team feeds simulated temperature values into the system to test shutdown logic without heating the device every time.
- RATIONAL is gradually moving test design into the business departments: the test team provides infrastructure and framework, while the departments are expected to write and maintain their own tests.
Commercial Kitchen Appliances Are Software Systems in a Steel Case
A combi steamer for a commercial kitchen is a networked software system, and testing it takes a test infrastructure to match. RATIONAL’s appliances run on around 700,000 lines of code, spread across several electronic components that communicate over bus systems and run different firmware.
The manufacturer develops the main software in-house, while the electronic components come from outside suppliers. Both worlds have to work together, and that interface is exactly where it gets decided whether an appliance works reliably.
The scale has nothing in common with a home kitchen. The largest models cook more than 90 chickens in one go, stand taller than a person and deliver up to 70 kilowatts of heating power. Software that controls that much energy needs a test setup that reproduces that power for real.
Why the Test Infrastructure Needs Real Hardware, Not Just Simulation
When you test an appliance with 70 kilowatts of heating power, software alone won’t do. At RATIONAL, the tests are backed by dedicated hardware in the team’s own racks: relay modules, remote power supplies, logic analyzers and the ability to switch electricity, gas and water on and off on demand.
Everything is reachable over the network, down to the serial console on the device. USB runs over the network as well, which Andreas Berger himself calls a challenge. This hardware costs a lot of money and material, and the team is clear that it is the expensive part of the setup.
In day-to-day testing, the limiting factor is not the software but the hardware under test. When a device has to heat up so the heating can be checked, the test spends most of its time waiting for physics. A server with eight cores and 8 GB of memory handles 18 test devices running at once without breaking a sweat. The heat still takes as long as it takes.
A Clear Focus: System and Regression Testing, No Unit Tests
The test automation team set its focus early. It does system testing, regression testing and low-level testing down to the bus systems. Unit and module tests stay with development, acceptance testing with the requester.
That boundary is a decision, not an accident. If you try to test everything, you end up testing nothing properly. Focusing on system behavior suits a product where the interplay of many components makes the difference.
A typical test case selects an operating mode, such as cooking with convection, and then checks at the controller level whether the right controllers respond, whether the motor turns, whether the heating elements switch on and whether the set target temperature is reached. On top of that come system integration tests that include the connectivity solution.
What Connectivity Actually Does for the Appliances
Connectivity here doesn’t mean the appliances talk to each other. It means updates and cooking programs can be rolled out remotely. Large chains need to push their cooking programs or a software update to every appliance in every location.
The manufacturer provides the service for that. A small restaurant can use the same connectivity solution with its own account. For testing, this means the rollout has to be checked just as thoroughly as the cooking itself.
Why Python Was the Right Choice for the Test Framework
The team chose Python because it didn’t want to compile anything and needed a compact language with plenty of libraries. Andreas Berger comes from the C and C++ world and knows the difference firsthand.
“If I try to set up a network connection or a serial connection there, it takes me forever. If I do the same thing in Python, it’s two lines of code.”
(Andreas Berger)
Getting rid of compile times was part of the point. Today the team maintains several Python packages of its own, uploads them automatically to its own registries and runs its own package infrastructure around them. Looking back, Andreas considers Python the right call.
Alongside Python, GitLab and a test design tool form the backbone. GitLab runs the pipelines and the quality assurance for the team’s own framework. The test design tool handles variant management: one abstract test case can be applied to 14 different device types through variable test steps, without multiplying the test case.
How Testing Runs Around the Clock
The goal is simple: every branch that needs testing runs fully automatically overnight, from start to evaluated report. Two branches get tested each night, each taking about four to five hours, because a device can’t be used by two runs at once and a queue sits behind them.
Each branch runs about 2,000 test cases. Results should be in by six in the morning. On weekends, a full test runs for around 30 hours.
During the day, the devices are busy with other work. Developers trigger on-demand runs when they have changed something and want to check it right away. Daytime is also when new test design happens, the team provides debugging support for complex bugs, checks its own packages after refactorings and maintains the hardware.
Here is how the schedule breaks down:
| Time window | Task | Scope |
|---|---|---|
| Night | Automated run, two branches | About 2,000 test cases per branch, 4 to 5 hours per branch |
| Weekend | Full test | Around 30 hours |
| Daytime | On-demand runs, test design, debugging support, hardware maintenance | As needed |
Risk-Based Testing Means Prioritizing by Customer Impact
The team finds bugs regularly and prioritizes them by impact. Anything that hurts the customer or would trigger lots of service calls is critical. A bug ten levels deep in a service menu can usually wait.
Stabilization branches are calmer, the main delivery branch is the critical one. And there is a classic problem every company knows: something gets changed and testing doesn’t hear about it in advance. It happens.
Bugs don’t just sit in the tool. The team goes to the developers, talks it through and works out together where the problem lies and how to fix it. Some failures turn out to be infrastructure issues, such as a network outage, and those have to be ruled out first.
Testing Safety Before the Device Gets Damaged
Protecting people is covered first by the inherent safety of the individual components and their certification. Those protective shutdowns kick in late. They are there to prevent injury and damage to the building, even if the device is already destroyed by then.
Testing starts earlier. It checks the shutdowns the software triggers before the device gets damaged. If a component is about to fail, the heating should switch off in time.
These tests come with their own traps. An empty steam generator is ruined within a short time if you heat it. So during the test of the heating lockout, it must not actually be empty, but the system has to believe it is. You have to think of every one of these conditions.
This is where simulation helps. Once earlier tests have covered the heating itself, the team feeds temperature values directly into the system and sets them one degree above or below the threshold to trigger the heating shutdown. The device stays intact, and the test still means something.
The Real Challenge Is Scaling, Not Technology
The biggest challenge is keeping pace with development. Embedded development has around 70 software developers at two locations, against ten people in test automation. The team wants to keep up without matching development headcount.
That has led to a change in strategy. Instead of delivering complete test frameworks and finished test design, the focus is shifting toward infrastructure. The team provides the hardware and software for testing, and the business departments are increasingly expected to take over test design.
The reasoning is easy to follow: whoever changes the requirements can update the tests along with them and cut back on their manual testing in return. That puts the workload where the domain knowledge is.
Then there is geography. One lab is in Landsberg, another in Wittenheim, France, where a subsidiary handles the tilting pan systems. Two colleagues there belong to the team. Regular visits in both directions hold the collaboration together and keep each site from drifting off into its own group.
Test Impact Analysis Is the Next Step
The next big lever is test impact analysis. The team is working on reordering test cases automatically based on what has changed in the code. Instead of running the same block every night regardless, the relevant tests should move to the front.
Asked about the highlight of the coming years, Andreas names two things: getting test impact analysis up and running, and having kept pace as the company scaled. The two goals belong together, because smart test selection is exactly what keeps a small test team effective next to a growing development organization.
Frequently Asked Questions
What hardware is required for automated system testing of devices with high power consumption?
In addition to the device under test, this requires a dedicated test infrastructure in racks: relay modules, remote power supplies, logic analyzers, and the ability to selectively turn power, gas, and water on and off. At RATIONAL, everything is accessible via the network, right down to the serial console on the device. This hardware is the expensive part of the system, not the software.
What limits the runtime of nightly test runs on real hardware?
The physics of the device under test, not the computing power. If a device needs to heat up first to test its heating system, the test run spends most of its time waiting for this process to complete. A server with eight cores and 8 GB of memory can easily handle 18 test devices running in parallel. Nevertheless, two test branches per night each take four to five hours.
Which test levels should a test automation team handle for embedded systems?
Not all of them. The team at RATIONAL focuses on system testing, regression testing, and low-level testing down to the bus systems. Unit and module tests remain the responsibility of the development team, while acceptance testing is handled by the client. This distinction is made deliberately: If you try to perform testing on everything, you end up not testing anything properly.
Is a scripting language a better choice than C or C++ for a test framework?
Yes, for test development. The deciding factor was that nothing needs to be compiled and many libraries are readily available. A network or serial connection, which requires longer code in C or C++, takes just two lines in Python. The trade-off is having to maintain the team’s own package infrastructure: several self-maintained packages that are automatically uploaded to the team’s own registries.
How do you keep test cases maintainable across many device variants?
By using high-level test cases with variable test steps instead of creating copies for each variant. At RATIONAL, a single test case can be applied to 14 different device types without having to duplicate it. A test design tool handles variant management, while GitLab manages the pipelines and quality assurance for the in-house framework.
What criteria are used to prioritize bugs found in system testing?
Based on the impact on the customer. What matters most is what would cause the user problems or trigger a large number of service calls. A bug in the tenth sublevel of a service menu is usually less urgent. The branch also matters: stabilization branches are less critical, while the main release branch is the more critical one.
How do you test safety shutdowns without destroying the test subject?
By specifically simulating individual measured values in the running system. If the heater has already been tested in preliminary tests, the team inputs temperature values and sets them one degree above or below the threshold to trigger the shutdown. Example: An empty steam generator would quickly break down if heated, so the system only needs to treat it as empty.
What does a small test team do when development is growing faster than the team itself?
It shifts the test design to where the domain expertise resides. At RATIONAL, ten people in test automation work alongside approximately 70 embedded developers at two locations. Instead of delivering finished tests, the team provides the hardware, framework, and infrastructure; the business units write and maintain their own tests and, in return, reduce the need for manual testing.


