Skip to main content

Search...

Architecture Fitness Functions: Making Decisions Testable

Architecture fitness functions turn design decisions into testable KPIs, so hard-to-reverse Type 1 decisions defend themselves against violations.

• • Updated: • 12 min read
Cover of the expert talk on 'Architecture Fitness Functions: Making Decisions Testable' with Ralf Enderle and Richard Seidl.

Architecture fitness functions are automated checks that translate architecture criteria into a measurable, testable form. Software architecture itself is the blueprint of a system: which cornerstones apply, which constraints must be respected and which decisions have to be made. Those decisions fall into two categories: irreversible Type 1 decisions with a big impact and revisable Type 2 decisions. A living architecture defends itself, for example through measurable KPIs and automated tests.

Key Takeaways

  • Architecture documents that only sit in a wiki don’t defend themselves: ADRs and fitness functions protect the architecture from violations only once they can be checked automatically.
  • Type 1 decisions, such as choosing between a monolith and microservices, can hardly be reversed in practice, so they need a clear rationale and have to be defended for the long term.
  • Type 2 decisions, such as the choice of logging framework, may and should change when team feedback or metrics speak against them, because their impact doesn’t put the architecture at risk.
  • Architecture boards should discuss Type 1 decisions only, because Type 2 topics get stuck there, blocking evolution the team could simply allow in a direct conversation.
  • AI can help make soft architecture rules technically checkable: it recognizes from the context when a catch block is logged but the log exposes sensitive data such as credit card numbers.

Architecture Is Out of Date the Moment It Is Written Down

A software architecture describes how a system should be built: what matters, which constraints apply, what to watch out for. It is the blueprint, a few levels above the code. And that very plan often lags behind reality. Architecture fitness functions are one way to keep it honest.

Thinking an architecture through takes days, weeks, sometimes months. Meanwhile product, software and requirements keep moving. In a strict waterfall this shows less, because the world there stays relatively stable. As soon as a team works in an agile way and requirements change, a long-crafted architecture risks missing what is actually needed.

That is no argument against architecture work, but one for the right dose. Thinking too long can mean holding on to requirements that no longer exist in that form. The job is to balance stability against changeability.

Why Early Architecture Decisions Shape Testing

Performance and security need to be tested early, because they depend on the architecture. Start too late and you end up shaking fundamental structures, which gets expensive.

That creates a balancing act. On one side are decisions with a big impact on the system and its quality characteristics, which can hardly be turned around later. On the other side is the wish to stay fluid enough to react to change. Serving both at once is the real difficulty.

The way out is to tell them apart: which decision is a cornerstone you have to defend? And which one is necessary but can be revised without drama?

Telling Type 1 and Type 2 Decisions Apart

Jeff Bezos’ analogy gives you a simple grid. A Type 2 decision is a door you can walk back through. A Type 1 decision is a door that falls shut behind you: the wall is high and there is no way back.

For every architecture decision, it pays to ask up front which category it belongs to. That determines how much energy it deserves when you think it through and how much energy it will cost later to defend it.

A classic Type 1 example is the choice between a microservice architecture and a monolith. In theory you get from a monolith to microservices by simply cutting it apart. In practice it isn’t that easy. The reverse route, merging microservices into one unit by copy and paste, also sounds doable and turns out to be hard.

There is some relief in this: if there are good reasons for a monolith, that is fine. Microservices are not mandatory. They bring their own drawbacks, such as lots of small plumbing between the services. What counts is that the choice is well reasoned and carried through the application’s lifecycle.

The logging framework is the counterexample. At some point a choice has to be made, but if it turns out after eight weeks that it doesn’t fit, you swap it out. The system keeps running, nothing fundamental changes. You are allowed to make decisions like that and make them differently later.

Type 1Type 2
Reversibilitypractically noneany time
ExampleMicroservices vs. monolith, hyperscaler lock-inLogging framework
Effort to think it throughhighmoderate
Where to decideBoard, with management backingwithin the team
Documentationwatertightkeep it short

Building Generically Has Its Limits

Replaceability costs effort, and at some point that effort tips into overengineering. With logging, you’d have to get a lot wrong before you can no longer swap it. Other topics can only be kept generic enough to leave the choice open forever with considerable effort.

If you make everything as generic as possible, you build a massively overengineered system, for reasons nobody needs. The more open a part is kept, the fewer concrete guidelines there are for implementing it. Then developers lack examples.

Some supposed Type 1 decisions can also be broken down into Type 2 once a suitable pattern exists. If there is a clean pattern on top, the concrete implementation hardly matters.

The Cloud Question Depends on Strategy, Not Just the Architect

How much room a decision leaves is not up to the architect alone. The choice of hyperscaler shows this. If it is open, the architect has plenty of options. If the company prescribes one of the big hyperscalers, the direction is set.

If that directive comes with the addendum “keep the system generic anyway”, the classification shifts. Committing to one hyperscaler is, in principle, a Type 1 decision. Building cloud-native, however, ties you less tightly to a provider and wins back some room to move.

If, on the other hand, you integrate provider-specific services deep into the system, it is clearly Type 1. You won’t get away from that hyperscaler again, at least not with reasonable effort. Migrations are possible, but the effort becomes enormous.

An Architecture That Defends Itself

Architecture documentation only works when it shows up beyond the sheet of paper. arc42 templates, C4 and ADRs are the right tools for capturing ideas without getting lost in detail. An ADR on its own, filed in a wiki nobody looks at, won’t achieve anything.

The goal is documentation that pushes back when someone violates it. It should switch on a red light: the idea you have here goes against what was agreed.

How much effort that takes depends on the type. Type 1 decisions need clean, reliable documentation, and everyone on the project should know it. Type 2 decisions can be documented more briefly, and you should assume that not everyone on the team knows them.

“If I manage to set up my ADRs so they can push back when they are broken, or at least switch on a red light, then someone comes along and it tells them: think about it again.”

(Ralf Enderle)

How Architecture Fitness Functions Make Decisions Testable

An architecture defends itself as soon as its criteria are translated into a checkable form. “The architecture should perform well” sounds nice, but it is worthless until it becomes KPIs you can test against.

The lever is linking those KPIs to the decisions behind them. You choose microservices, for instance, to scale per service and hold a certain level of performance. Exactly those goals become metrics. If a test fails, a light goes on and reports that something is off.

Fitness functions, a concept by Neal Ford, translate hard architecture criteria into a technically checkable form. On paper the criteria feel rigid; at the moment of testing they become soft enough to grasp technically.

Where AI Adds Context to Architecture Checks

AI complements technical checks where pure rule checks miss the intent behind a rule. An example: every catch block should show up in the log so you can react to production errors. Regular expressions can verify that a log statement follows a catch block. What gets logged there, the regular expression can’t tell you.

If you also describe why this logging matters, an AI can read the surrounding context. In the context of a credit card payment, it can recognize that whatever is logged from the error might contain the credit card number. A regular expression would never catch that.

This check doesn’t give you 100 percent certainty. AI systems don’t wake up in the same mood every day. But there is no reason to do without the extra safeguard it provides.

Architecture Decisions Belong Where They Have an Impact

An overarching board that reviews the architecture every few weeks and contributes ideas runs the risk of debating far away from the reality of the software. Here, too, the type distinction helps. Type 1 decisions can belong in such a body, because they need management backing.

Discussing Type 2 decisions in large meetings is wasted time. It means negotiating topics without impact and blocking evolution that would otherwise happen effortlessly, instead of dealing with the hard things. These decisions belong in the team.

If the team notices in a retro that a Type 2 decision isn’t holding up after a few weeks, it changes it. That is exactly what it is built for, with little impact. Why defend a decision that doesn’t hurt to change? Clinging to it means missing out on evolution you could have had for free.

What Testers Should Ask of Architects

Testers shouldn’t only ask what the architecture looks like, but how it translates into checkable factors. The central request to the architect is: think about how your architecture shows up in KPIs, in soft or hard factors that can be checked.

That question creates a shared goal: making sure the architecture chosen for good reasons is actually lived and not violated. Fitness functions are one way to get there, AI-supported context checks another.

Many teams already have an architecture process in place. It can be adapted. A retro is a good place to raise the question of whether the current approach defends the right decisions and leaves the right ones open.

Frequently Asked Questions

Why does a carefully designed architecture often no longer meet requirements?

Because the design process takes days, weeks, or months, and the product, software, and requirements continue to evolve during that time. In a strict waterfall model, this is hardly noticeable because the environment remains more stable. In agile teams, however, an architecture that takes a long time to develop can easily miss the mark when it comes to actual needs. This isn’t an argument against architectural work, but rather for finding the right balance.

Why should performance and security be tested early in the project?

Both of these quality characteristics are directly tied to the architecture. If you wait until late in the process to test them, you’ll have to shake up fundamental structures, and that gets expensive. This is precisely why there’s a balancing act between high-impact decisions that are difficult to reverse later and the desire to remain flexible enough to adapt to changes.

How can you tell if an architectural decision can still be revised later on?

The image of the door from Jeff Bezos’ analogy is helpful: You can walk back through a Type 2 door, but a Type 1 door closes behind you. This classification should come at the start of every decision, because it determines how much effort is worth putting into thinking it through and how much it will cost later to defend the decision.

Are microservices a better choice than a monolith?

No, microservices aren’t mandatory. If there are good reasons for a monolith, that’s fine. Microservices come with their own drawbacks, such as a lot of small details between services. What matters is that the choice is well-founded and sustained throughout the application’s entire lifecycle, because it’s one of those decisions that are practically irreversible.

Is it worth building system components to be as generic and replaceable as possible?

Only up to a point. Replaceability comes at a cost in terms of effort, and at a certain point, that effort tips over into overengineering, often for reasons that serve no one’s needs. The more open a design choice is kept, the fewer concrete guidelines developers receive, and the fewer examples are available. With logging, minimal planning is sufficient; other areas can hardly be left open indefinitely.

Is committing to a cloud provider always irreversible?

At its core, it’s a Type 1 decision, but its impact can be mitigated. Those who build cloud-native solutions tie themselves less tightly to a single provider and regain flexibility. If provider-specific services are deeply integrated into the system, it clearly remains a Type 1 decision: While a switch is technically possible, the effort required becomes enormous. Often, the company sets the direction anyway.

Is it enough to document architectural decisions as ADRs in a wiki?

No. An ADR in a wiki that no one looks at achieves nothing. Documentation is only effective when it extends beyond the page and triggers a warning as soon as someone violates it. ARC42, C4, and ADRs are suitable tools for this, but they must be made automatable for verification.

What does AI do when verifying architecture rules that regular expressions cannot?

AI captures the context behind a rule. A regular expression can check whether a log statement follows a catch block, but it cannot determine what is being logged there. If you also describe the intent of the rule, AI can recognize, in the context of a credit card payment, that the log entry might reveal the card number. This does not provide 100% certainty.

Share this page