Fair AI describes the requirement that AI systems demonstrably do not discriminate and meet values that society accepts. Because fairness measures can contradict each other mathematically, AI fairness can’t be solved on a purely technical level. An assurance case framework structures the requirements instead: from one main claim to testable subclaims down to evidence such as test results or software features.
Key Takeaways
- Fairness measures in AI contradict each other mathematically: optimizing one inevitably makes another worse, which is why computer scientists shouldn’t make that call on their own.
- The assurance case approach from safety engineering carries over to fairness: a main claim is broken down step by step into testable subclaims and evidence that can be checked.
- In a real industry project with a medical rotation planner, the framework showed that about a quarter of the required fairness measures weren’t even in the development plan.
- Asking stakeholders brings up fairness requirements that go far beyond non-discrimination, including transparency, autonomy and practical points such as short waiting times between departments.
- A structured argument in the form of an assurance case protects developers legally, because if discrimination is alleged they can show that they worked to the best of their knowledge and according to the state of the art.
Fair AI Starts with Deciding What Fair Should Mean
AI fairness is not something developers can define on their own. Fairness is a requirement that comes from society and has to be translated into testable criteria, and that translation step doesn’t belong in the hands of the programmers.
The reason is how widely the technology is used. Software has always been embedded in social processes, but AI is now used in so many areas that requirements suddenly appear that nobody had thought about before. The people who build AI have five to seven years of computer science education behind them, but no training in ethics or social questions and no domain knowledge, in medicine for example.
Marc Hauer explains the core problem with the example of hiring. In the past, nobody looked over the shoulders of HR staff for every single decision, and there was no need to: a person makes maybe ten or twenty such decisions a month. An AI applies the same logic at a completely different scale. That scaling effect is exactly what forces you to clarify the concept of fairness up front.
Why There Is No Objective AI Fairness
Fairness can’t be resolved mathematically in a single way, because different ideas of fairness can contradict each other. If you optimize one measure, you run into another.
A simple example makes this clear. With child benefits, a distribution counts as fair if everyone gets the same amount of resources. With welfare, it counts as fair if everyone reaches a minimum level after the distribution. Both feel intuitively just, but you can’t satisfy both principles at the same time.
The computer scientist Kleinberg showed this tension for three clearly defined fairness measures: they contradict each other. That is exactly why the decision about what is fair in a specific case must not rest with computer scientists.
There is another trap. If you are simply told to build a fair system, you can run every available fairness measure and will almost certainly find one that produces a good value. That is not a sound statement about fairness.
How an Assurance Case Breaks Fairness Down into Testable Requirements
The approach comes from safety engineering and is called the assurance case framework. It turns an abstract claim into a structured argument that can be backed up with evidence and challenged.
It starts with a main claim. In safety engineering that is “My system is safe”, in the fairness context “Our system is fair”. Below it sits a layer of argument that asks: which subclaims have to hold for the main claim to be true?
These subclaims are broken down further until they are small enough to be backed up concretely. Depending on the case, the evidence can be:
- an existing feature of the software,
- test results and traceable test processes,
- references to scientific publications,
- existing business processes, for example a way to contact the responsible planner.
Openly stated assumptions are part of the structure too. One assumption might be that you work only with fairness measures and look at nothing else. If an assumption like that is written down openly, a reviewer can challenge it. That is the real value: the argument can be checked, criticized and sharpened from the outside.
Asking Stakeholders Instead of Optimizing a Metric
Fairness becomes concrete once you ask the people affected and the users what they would actually accept as fair. The answer usually goes far beyond non-discrimination.
Tobias Krafft tested the framework outside the safety world, in an industry project for a medical rotation planner. Medical students rotate through different hospitals and departments for fixed periods, and an AI component was supposed to create the plans for these rotations.
Asking the stakeholders produced several requirements that, for them, are part of fairness:
- non-discrimination
- transparency
- autonomy and control over their own plan
- short waiting times between departments
- having to move as rarely as possible
Each of these requirements became a subclaim under the main claim “Our system is fair”. For transparency and control, that meant in concrete terms: students can see their plan, see KPIs such as average waiting times, can challenge their plan and voice criticism, and can contact the planner who makes the final decision.
The Assurance Case Reveals Gaps Before the System Is Finished
The biggest practical benefit showed up during development: the assurance case makes visible which requirements haven’t been planned for at all.
In the rotation project, stakeholder acceptance was the top goal, so fairness was assured during development, not afterwards. At the time the results were published, about half of the evidence couldn’t be provided yet, because the software wasn’t that far along.
The second finding stood out more: about a quarter of the required evidence wasn’t even in the development plan. The assurance case thus identified new requirements and features, which were then added to the plan. One example was the option to compare two alternatively generated plans side by side and see directly where one is better than the other.
AI Makes Existing Discrimination Visible
The debate about fair AI often overlooks that human decisions were rarely better before. People had their biases and their tunnel vision, but nobody looked systematically.
When AI is used, people check much more closely whether a result is fair or discriminatory. That scrutiny brings problems to light that simply went unnoticed in a purely human process. The key difference: with AI, you can name the adjustable parameters, understand them and challenge the result.
Media coverage paints a one-sided picture. It reports on cases of discrimination, denied welfare, money demanded back. Where AI helps rarely makes the news. That imbalance can cause economic losses as soon as attention pounces on an incident.
A Documented Argument Protects Developers
An assurance case worked out according to the state of the art serves as the basis for your argument when an allegation of discrimination comes up. It moves responsibility off the developers’ shoulders and to where it belongs.
“So far the rule is: you are responsible for your product. But as a computer scientist, I learned how to develop. I have no training in ethics, no domain knowledge in medicine, and I have no idea what consequences this might have.”
(Marc Hauer)
With the documented assurance case, including all evidence and assumptions, you can say after an incident: we worked to the best of our knowledge and in good faith, this and this is what we did, and no better approach is known today. That lowers the barrier to using AI in areas where fairness matters.
Not every use case needs this effort. If an AI on an assembly line checks whether a screw meets the standard, non-functional requirements like fairness are irrelevant. That is exactly where teams should be able to work quickly and in an agile way.
Regulation by Risk Instead of One Rule for Everything
The European approach regulates AI based on risk rather than treating everything the same. That shifts the effort to where a decision becomes critical.
The AI Act defines risk classes. Closer scrutiny applies where an AI has a say in decisions about human rights or fundamental questions. Where the benefit outweighs any critical consequences, there is room to work fast.
AI never operates in a legal vacuum anyway. If an AI supports a doctor, the existing rules for medical practice still apply. In that sense, a risk assessment doesn’t necessarily lead to overregulation. It applies existing legal frameworks to a new tool.
Templates and Open Source as a Way to Scale
An assurance case can’t be copied one to one from one project to the next, but useful templates can be prepared for specific fields. That is the lever for making the method usable beyond individual projects.
So far there are two scientific publications and a guidebook that explains the approach for a broader audience. The German standardization roadmap for AI includes the method as a proposal. A working group on AI fairness in financial services is developing a template for fairness aspects in credit scoring systems.
The biggest practical hurdle is getting the right stakeholders around one table. Within a tight project scope that is hard, and an open call to the community brings in many competing ideas.
As a long-term goal, Tobias outlines a common task framework modeled on image recognition: the community gets a sufficiently specified use case, builds assurance cases for it, shares and improves them, and after a year or two picks the best one. The safety field has shown that the method works, and an open source model would be a way to share exactly that experience widely.
Frequently Asked Questions
Can fairness in AI systems be achieved purely through technical means?
No. Different measures of fairness are mathematically contradictory: optimizing one inevitably worsens another. Computer scientist Kleinberg has demonstrated this for three clearly defined measures. This conflict is also evident in everyday life: In the case of child benefits, a distribution is considered fair if everyone receives the same amount; in the case of welfare, it’s fair if everyone receives at least a minimum amount. You can’t have both at the same time.
Why is fairness a bigger issue in AI than in traditional software?
The reason is scale. Software has always been embedded in social processes, but a human resources manager might make ten or twenty hiring decisions a month, whereas an AI applies the same logic to a completely different scale. Added to this is the gap in training: Developers have a degree in computer science, but no training in ethics and no domain-specific knowledge, such as in medicine.
How do you substantiate the claim that a system is fair?
The main claim is broken down into subclaims until they are small enough to be supported by concrete evidence. Evidence includes existing software functionalities, test results along with a traceable test process, references to scientific publications, or existing business processes, such as a way to contact the responsible planner. Assumptions are openly disclosed so that a reviewer can challenge them.
What is the benefit of asking stakeholders and users about their understanding of fairness?
Requirements come to light that go far beyond non-discrimination. In an industry project for a medical rotation planner, participants cited transparency, autonomy, and control over their own schedule, short wait times between departments, and as few transfers as possible. Each of these requirements became a separate subclaim under the main claim.
At what point in the development process is it worthwhile to ensure fairness?
During development, not afterward. In the rotation project, about half of the evidence could not yet be provided because the software was not far enough along. More important was the second finding: About a quarter of the necessary evidence was not included in the development plan at all. For example, the feature to compare two alternatively generated plans in parallel was added.
Why isn’t it enough to simply test a system for good fairness scores?
Because a suitable metric can always be found. Anyone tasked solely with building a fair system can run calculations using all available fairness metrics and will almost certainly discover one that yields a good value. That is not a reliable statement about fairness. What is needed is a predefined, well-reasoned, and verifiable line of reasoning.
Does every AI system need fairness safeguards?
No. If an AI on an assembly line is checking whether a screw meets the standard, non-functional requirements such as fairness are irrelevant; the goal there is to work quickly and agilely. The European AI Act regulates on a risk-based approach using risk classes. Closer scrutiny is applied where an AI plays a role in decisions involving human rights or matters of principle.
Does a documented fairness rationale help in the event of a discrimination allegation?
Yes. With a state-of-the-art assurance case, complete with all supporting evidence and transparently documented assumptions, it is possible to demonstrate in the event of an incident what steps were taken and that no better approach is known. This shifts responsibility away from the developers and lowers the barrier to using AI even in areas where fairness is a concern.


