Skip to main content

Search...

Hardware-in-the-Loop Testing: Kernels on Real CPUs

A kernel has to run on every CPU generation, so hardware-in-the-loop testing puts it on real devices. Developers and CI reserve clean machines.

• • 8 min read
Cover of the expert talk on 'Hardware-in-the-Loop Testing: Kernels on Real CPUs' with Markus Napierkowski and Richard Seidl.

Hardware-in-the-loop testing is the automated testing of system software, such as a Linux kernel, on the physical hardware it has to run on. A platform-specific image gets its final acceptance only there. An open-source framework reserves test machines in a clean initial state, boots the image and hands the devices to developers and the CI pipeline through a REST API.

Key Takeaways

  • Hardware-in-the-loop testing puts a kernel on every CPU generation it has to support. Virtualized tests leave exactly that step open.
  • The framework boots every reserved device from a USB drive with the matching image: that is how a HIL test starts from a clean initial state.
  • Instead of pushing a change and waiting for the pipeline, a developer reserves the hardware-in-the-loop machine directly and returns it to CI when done.
  • Each test machine hangs off a Raspberry Pi for USB simulation, HDMI capture and the serial console, with about five lines of code on the software side.
  • HIL testing runs inside GitLab CI before or after every pull request and reports each test case as a JUnit XML artifact.

A Kernel Only Proves Itself on Real Hardware

Hardware-in-the-loop testing, HIL testing for short, means the software under test runs on the actual target device while a test environment controls that device from the outside, powers it on, supplies the boot image, and reads back what it does. Markus Napierkowski and his team work at the Linux kernel level and have built microkernels in the past. A kernel has to run on every CPU generation, so it has to be tested on every one of those hardware generations as well. Virtualization covers a lot of ground, but it cannot stand in for the real platform.

The same holds one level up, with embedded PCs. Most of the work can be verified in a virtualized setup. The final acceptance test cannot, because an image tailored to one specific platform only proves it boots when it boots on that platform. From day one, the goal was the same as with unit tests: fast feedback, nobody clicking through a device by hand.

The hardware in the loop lab that colleagues of Markus set up more than ten years ago was the starting point. Use cases and solutions have changed several times since. The current version is a framework built over several years with a partner organization, closed source for a long time, and on GitHub for two months now.

Why Is a Clean Initial State the Sticking Point in Hardware-in-the-Loop Testing?

Because test automation at the software level is usually in place already and still proves little. Almost every team manages to upload a script and a payload to the device and run it. What is missing is setup and teardown for the hardware itself. Without both, nobody knows for sure what state the device is in after a run, and the next red test says little about the software.

A test machine gets reserved and is then in a defined initial state. The test specifies the boot image, plus any other devices around the machine that need initializing, plus a network of several machines if the case calls for it. Then the hardware powers on, and the actual test runs as if someone had logged in over SSH and started it by hand.

Markus is open about the limit. The framework keeps clean initialization simple because it is optimized for the case where booting from a USB drive is good enough. Flashing firmware onto a device is not supported yet; anyone who needs it has to build that part themselves. In his experience, the USB case still covers a large share of the team’s own use cases and many of the ones they see at customers.

How Do Developers and the CI Pipeline Share the Same Test Hardware?

Through reservations backed by a secret. The tests run in the pipeline before or after every pull request, whichever fits the project’s strategy, and the same infrastructure is open to developers directly. Anyone who wants to investigate a misbehavior takes the same machine for the same test without interrupting the CI. Once the machine is free again, the pipeline takes it back.

“As a developer, I can now reserve a piece of hardware and use it as if it were right next to my desk.”

(Markus Napierkowski)

The trigger was a working rhythm many developers know. Push a change, wait for the pipeline, get a coffee, come back, still waiting. Or the message in the team chat that the machine is now reserved and would everyone please not deploy to it. Which happens anyway. A reservation, by contrast, belongs to one person until that person hands the machine back. When several units of the same hardware exist, the system distributes requests dynamically between CI and people.

The secret can be shared, and that pays off in one situation. A test fails, and two people look at the web UI together to see what the screen output of the individual machines shows right now and why.

One Raspberry Pi per Test Machine, Five Lines of Nix per New Device

Adding a new piece of hardware costs about five lines of code on the software side plus one deployment. Markus and his team run on Nix and NixOS and keep the entire test infrastructure in one large Infrastructure-as-Code repository. A new device there is a duplicated piece of configuration, rolled out once.

On the hardware side, every test machine is connected to a Raspberry Pi. The Pi simulates USB devices such as the boot medium, reads the machine’s HDMI output, and connects to its serial console.

The Pi has to be prepared once. After that, the infrastructure description hides the rest of the complexity; anyone who has understood the overall system once extends it without fiddling with individual devices.

How Do You Write a HIL Test Without Knowing the API?

By letting the tooling build the environment before the test script starts. Underneath everything sits a REST API; the web UI uses it, and so does the command-line tooling. Whoever reserves a machine gets a secret back that acts as the context for all reserved machines. The same path networks the machines, with a VLAN switch as part of the setup.

A test describes in JSON which machines it needs, which network they belong to, and which boot medium they get. The test script itself begins at the point where the machines are reachable over SSH. In the best case, the person writing the test never needs to know how the API underneath talks.

One major use case was GitLab CI. Each test case of a run comes out as a JUnit XML artifact, so the CI displays results in a structured way. Testing on hardware matters for performance benchmarks too. Virtualized measurements are often not stable enough, and sometimes the number sits in a product requirement. A router that has to be up within ten seconds needs measurement data that proves it, and those values can be recorded from the runs and forwarded to an infrastructure that collects and processes them.

An Internal Tool, Deliberately on GitHub

Markus still calls the framework an internal tool. It covers the team’s own cases and the partner’s; whether a few tweaks would make it useful to others is an open question. There are first conversations, but no proof and no outside investment. Beyond some cleanup after the release, there is no roadmap. The candidate that emerged from conversations at TACON is an interface for getting firmware onto devices instead of only booting from a USB drive. For the current use cases, Markus considers the state good enough.

The framework went open source for two reasons. There was no commercial value behind the proprietary version that would have justified keeping it closed. And the team’s entire infrastructure is built on open-source software; Markus describes it as a matter of honor to give back what the team built, even if it turns out to need too many changes to be useful beyond their own company.

If you test embedded systems with changing hardware configurations and reservations still go through the team chat, the project is worth a look. Markus is currently collecting feedback on which use cases the framework is still missing.

Frequently Asked Questions

What is hardware-in-the-loop testing?

Hardware-in-the-loop testing runs the software under test on the real target device while a test environment controls that device from the outside, powering it on, supplying the boot image, and reading its output. For kernels and embedded images it is the only way to show that a build actually works on a specific platform.

How do hardware-in-the-loop tests run in a CI pipeline?

As a regular pipeline step, for example before or after every pull request. The tooling reserves a machine, boots it with the specified image, runs the test, and releases the machine afterwards. The results of the individual test cases land in the CI as a JUnit XML artifact, which GitLab displays directly in its interface.

Is a virtualized test enough, or does a HIL test need real hardware?

Virtualization is enough for most of the development work, but not for the acceptance test. An image tailored to one platform only proves that it boots and runs once it does so on that platform. The same applies to showing that a kernel works on every supported CPU generation.

Can hardware-in-the-loop testing be used for performance measurements?

Yes, especially when a number is written into the requirement. A router that has to boot within ten seconds needs measurement data from the real device, because virtualized measurements often fluctuate. The values from the test runs can be recorded and forwarded to an infrastructure that collects and processes them.

Share this page