Using a reference implementation as a test oracle means that the tester builds the same functionality a second time, independently of the developer, in a simple high-level language such as Python or MATLAB. A back-to-back test then compares both results. This replaces the manual calculation of expected values, turns the review into a code review and scales even to thousands of signals.
Key Takeaways
- With a reference implementation as test oracle, the tester rebuilds the software independently of the developer, and a back-to-back test compares the two versions.
- Testers must not have access to the original source code while writing the reference implementation, so they can’t unknowingly copy the developer’s coding errors.
- Reviewing a reference implementation is a structured code review with a checklist, which takes far less effort than recalculating expected values by hand.
- Automatically generated stimuli for structural code coverage can be run against the reference implementation to cover the last few percentage points without creating a self-fulfilling prophecy.
- When a requirement changes, only one spot in the reference implementation needs a code change instead of dozens of expected values in the test cases.
What Is a Pseudo Test Oracle Based on a Reference Implementation?
A test oracle is whatever tells you the expected result of a test. In a pseudo test oracle based on a reference implementation, that source is a second, simplified version of the function under test, built by the tester independently of the developer. The delivered software and the reference implementation run on the same input data, and their outputs are checked against each other.
The developer writes the production code that later goes into the vehicle. In parallel, the tester implements the same logic in a simple high-level language such as Python or a MATLAB script. If both implementations produce the same outputs, the behavior counts as compliant with the specification.
Stefanie Leitner works for an automotive supplier within the Volkswagen Group that develops embedded software for autonomous driving and vehicle safety. The team has used this approach for more than ten years to test complex safety-critical functions such as lane keeping assistants and emergency braking assistants.
Why Manual Expected Values Break Down with High Complexity
Calculating expected values by hand doesn’t scale once the number of signals to check runs into the thousands. That is the problem the reference implementation solves.
The very first software components, which verify incoming bus signals, don’t check 10 or 20 signals but 2,500 to 3,000. Format, validity and value range have to be right for every one of them. Nobody can keep up with calculating those expected values by hand. If a small detail changes, testing stalls and can’t give early feedback.
The second driver is the algorithms themselves. Trajectory planning in autonomous driving, for example deciding whether the lane is clear and an evasive maneuver is possible, involves many consecutive calculation steps. For every test step, the tester would have to work out the expected value separately. Both problems led to the same answer: an automated oracle that calculates the expected values instead of someone entering them by hand.
How the Reference Implementation Is Built
The reference implementation covers the full logic of the function, not just a stub. At the end, a back-to-back test compares the original software with the replica, signal by signal.
Unlike production code, the reference implementation doesn’t have to go through the full development process. Coding guidelines don’t matter, and neither does optimizing for resources, because the code never goes into the vehicle. Python libraries provide many functions out of the box, which speeds up the work.
For the approach to hold up, a few fixed rules apply:
- Independence: The developer and the tester have to be two different people. Nobody builds their own model and then tests it against itself.
- No access to the original code: Testers don’t see the production source code while they write their implementation. That keeps them from borrowing solutions or reproducing the same coding error.
- Traceability: As in the original code, the requirements are linked in the reference implementation. That makes completeness checkable and helps keep it consistent with the specification.
The size of a reference implementation depends on the unit, typically between 100 and 1,000 lines. At unit level, it is deliberately kept small.
The Review Becomes a Code Review Instead of a Math Marathon
The biggest everyday gain is in the review. Where reviewers used to recalculate every manually determined expected value, they now do a code review with a checklist.
ISO 26262 requires an inspection of the test cases from ASIL B upward. With the classic approach, that meant working through every hand-calculated value a second time: laborious and error-prone. With a reference implementation, the reviewer checks the code instead, uses the linked requirements to see whether the tester interpreted them correctly and so makes sure the tests are complete and consistent.
How much this matters became clear in one project where a project manager shied away from the effort and went back to the classic approach with scripts. By the time of the review at the latest, everyone was complaining. The lesson was clear: never again.
Developers and Testers Start in Parallel
As soon as the requirements are approved, developers and testers start at the same time. One implements the production software and the developer tests, the other the reference implementation and the test case specification.
Before that, the requirements go through a double review, once from the developer’s and once from the tester’s point of view. That catches a lot of ambiguity early. Ideally, both sides finish at the same time and can move straight into test execution.
If the comparison fails later, the question is which side the defect is on. The rule: the tester checks their own reference implementation first. Only when they are sure it behaves as specified does a defect ticket go to the developer, who then debugs the software.
Where the Real Value Lies: Test Depth and Maintainability
The approach compares every output signal at every point in time, not just the few signals within the scope of a single requirement. That adds a lot of test depth.
It is used mainly at unit and integration test level. Because there is a value for every signal at every test step, edge cases and side effects show up that you would miss if you looked only at the specified output signals.
It pays off when things change, too. If a requirement changes, a single line in the reference code is often enough, for example >= instead of >. After the back-to-back test, the expected values are up to date, and nobody has to adjust 60 expected values in the test cases by hand. Variants, coding and calibration data can be programmed in directly: swap the data set, and you’re done.
Another effect showed up with structural coverage. From ASIL B, decision coverage is required, from ASIL C, MC/DC. The last five percent to full coverage are tricky for testers.
“If you generate the expected value, it’s a self-fulfilling prophecy. But we have our reference implementation. We take the stimuli, run them through the reference implementation, and that tells us for the last five percent too whether the software behaves the way it should.”
(Stefanie Leitner)
The tool only generates the stimuli from the code. The expected values come from the reference implementation. Otherwise the code would simply confirm itself.
Early Doubts in the Team: Is the Bug in the Code or in the Model?
The biggest hurdle isn’t technical. It is a question of trust. At first, developers doubt whether a reported defect is really in their software or in the reference model.
Only the safeguards built in beforehand can resolve that doubt: independent implementation, review, traceability. Everyone makes mistakes, and now and then something slips through. In practice, though, the comparison holds up, and when in doubt, both look at the reference implementation together. Often it turns out that someone understood the requirement differently or forgot something in the code.
For developers, the reference implementation even helps with debugging. They can see how the tester solved the requirement and find out more quickly which of the two didn’t implement it as intended.
What Skills Testers Need for This Approach
Testers need to be able to write code, but they don’t need specialist knowledge. Programming basics such as loops are enough to start with, and the rest can be learned.
Many testers in embedded development have a computer science background anyway, so the basics are there. Python and MATLAB scripting haven’t caused problems so far and could be taught in internal training. Instead of strict coding standards, there is a template and a few guidelines, mostly for traceability. That way, one tester quickly finds their way around another’s work if someone is out.
When a Reference Implementation Isn’t Worth the Effort
Not every function deserves a reference implementation. For very simple units with only a few mathematical calculations, the effort often outweighs the benefit.
That is why there are criteria for when the approach makes sense and when it doesn’t. These criteria are to be refined further. The team is also looking at whether AI can speed up writing the reference implementation, which would shift the threshold for borderline cases.
Frequently Asked Questions
At what point does it no longer make sense to calculate expected values manually?
As soon as the number of signals to be tested reaches the thousands, calculating expected values manually is no longer feasible. The first components for verifying incoming bus signals check 2,500 to 3,000 signals for format, validity, and value range. Even multi-stage algorithms, such as trajectory calculation in autonomous driving, require each expected value to be determined individually. This hinders early feedback.
What’s the point of implementing the same function twice?
The second implementation provides the test oracle. Developers and testers build the same logic independently of one another; the tester uses a simple high-level language such as Python or a MATLAB script. Both run on the same input data, and their outputs are compared in a back-to-back test. An automotive supplier within the Volkswagen Group has been testing lane-keeping and emergency braking assistants this way for over ten years.
Can the developer write the reference implementation themselves?
No. Developers and testers must be two different people; otherwise, someone would be testing their own model against itself. Additionally, testers do not see the production source code at the time of implementation, so they cannot copy solution approaches or unknowingly reproduce a coding error made by the developer. The requirements are linked in both implementations.
How do you verify test cases whose target values are no longer calculated manually?
Manual verification is replaced by a structured code review using a checklist. ISO 26262 requires an inspection of the test cases starting at ASIL B. Traditionally, this meant recalculating all manually determined values a second time. When reviewing the reference implementation, however, the reviewer instead follows the linked requirements and assesses whether the tester has interpreted them correctly.
Who resolves the error if the software and the simulation produce different values?
The tester first checks their own reference implementation. Only when they are certain that it behaves as specified is an error ticket sent to the developer. At first, developers often doubt whether the error really lies with them. When in doubt, both parties look at the reference implementation together, and it often turns out that there was a difference in how the requirement was understood.
How time-consuming are requirement changes with this test approach?
Usually, a single code change is sufficient, for example “greater than or equal to” instead of “greater than.” The back-to-back test then generates the updated target values; no one has to manually adjust 60 expected values in the test cases. Variants, encoding, and calibration data can be programmed in: Swap the data set, done.
Can structure-based coverage be achieved automatically without the code validating itself?
Yes, if the stimuli and expected values come from different sources. The tool generates the stimuli from the code on its own, while the reference implementation provides the expected values. Otherwise, it would create a self-fulfilling prophecy. The last five percent needed to reach the required coverage is particularly difficult: Decision Coverage starting at ASIL B, MC/DC starting at ASIL C.
Are there functions for which a reference implementation isn’t worth the effort?
Yes. For very simple units with only a few mathematical calculations, the effort often outweighs the benefit. That’s why criteria are needed to determine when this approach makes sense. The full benefit is realized when every output signal is compared for each test step, thereby revealing edge cases and side effects.


