Microservices testing at REWE spreads automated tests across three levels: unit tests for business logic, API tests for each service and a small number of end-to-end tests for the most important user journey. Flaky tests are quarantined and either stabilized or removed. Production monitoring and alerting take over much of what large test environments used to do, because fast rollbacks keep the risk of each release small.
Key Takeaways
- Microservices solve an organizational problem, not a technical one: turning Conway’s Law into an architectural principle scales teams, not just code.
- A test suite made up almost entirely of end-to-end tests that runs for hours is not a safety net. It is a maintenance burden that slows down the whole development process.
- Putting flaky tests into quarantine and asking the responsible team to fix them is an effective way to shrink a bloated test suite step by step.
- Monitoring and alerting in production can partly replace end-to-end tests on test stages, because in e-commerce a quick rollback costs less than elaborate test environments for rare production defects.
Why REWE Split Its Monolith into Microservices
Microservices testing at REWE started with an organizational problem rather than a technical one. The company’s online shop began life as a monolith, wired to a long list of legacy systems and grown chaotically over years of deadline pressure. The goal was to break the code into modules that could be maintained, and to let many teams work on it at the same time.
In a monolith, that is exactly the hard part. Michael Kutz describes the “merge hell” that follows when one person refactors a file and three others have changed the same code in different ways. You end up with a hundred modified files and no idea which version should win.
Microservices address this by turning Conway’s Law into an architectural principle. Each team owns a clearly bounded area and can deploy on its own schedule. The price is a new set of problems: coordination, managing the APIs between services and a more complex code base.
“Technically, microservices don’t really make sense. It’s ‘just’ an organizational problem that you lift to a technical level. And the ‘just’ is where I see the mistake in that sentence, because in IT everything is an organizational problem.”
(Michael Kutz)
Long End-to-End Test Runs Slow Everything Down
If you protect a system almost entirely with end-to-end tests, you pay for it with long run times and results that tell you very little. The old monolith’s test suite took three to four hours to run, and still a good two hours after it had been optimized.
There is an understandable reason teams end up there. When a system changes constantly, developers are reluctant to write detailed unit tests, because those details keep changing. So they focus on what stays stable: the end-to-end test that checks the requirement directly.
The team wanted something else. Taking a cue from Extreme Programming, they wanted to know within about ten minutes whether a change had broken anything. A red light several hours later doesn’t come close.
Failure analysis was the most expensive part. A failed end-to-end test gives you a red signal and nothing else, no hint about the cause. Often the culprit was a slow database at the wrong moment, something the code that had just been deployed had nothing to do with.
Three Red Runs Before a Failure Counts
One honest, if unflattering, rule from practice: a test failure was only treated as a real defect once the test had gone red three times in a row. Rerunning the suite was simply cheaper than investigating every single failure.
Nobody should copy this, but it shows the core problem with flaky end-to-end tests. Once analysis costs more than a rerun, the test stops working as an early warning system. At that point you are testing the stability of your infrastructure more than your code.
Carving Microservices Out of the Monolith Step by Step
The way out of the monolith followed the strangler fig pattern. The existing code was treated as a given and left alone, because any change could have destabilized a structure that had long since set like concrete. The assumption that this code was free of defects is never true, but it was the working basis.
Features were replaced one at a time by code in microservices. A bypass in the monolith routed calls to the new service, which took over the function. At heart, the first services were databases with APIs: very little business logic, mostly data storage that wrapped the monolith’s complex database structure.
After that, the business logic moved over to these services piece by piece. In the end, the monolith did nothing but render HTML. One deliberate architectural decision was the switch to asynchronous APIs, even though eventual consistency and end-to-end tests don’t get along well.
Microservices Testing: How to Build a Real Test Pyramid
Every new microservice got a cleanly designed API from day one, along with its own tests. Using Spring Boot’s test tooling, the team started only the service under test and tested its API thoroughly without looking at the internal implementation. Unit tests were added wherever there was business logic.
Over time, the tests took the shape of a pyramid. What was missing for a long while was a good top. The old full test suite stayed on as the end-to-end layer, and everyone was responsible for it. When everyone is responsible for something, in the end nobody usually is.
Cutting back the end-to-end tests followed a clear logic. If a feature was already covered at the API level or by unit tests, there was no need for ten end-to-end variations of the same case. One test per risk is enough when that risk is already covered lower down in the pyramid.
Quarantine Instead of Maintenance: Sorting Out Flaky Tests
Once many teams share a test suite, nobody can maintain it centrally anymore. Four to six teams grew to twelve to fourteen. What started out as a kind of guild working on the suite together turned into a handful of people looking after it.
The fix was a simple mechanism. Any test that failed on its first run went into quarantine, and the team most likely responsible was asked to stabilize it, because flaky tests were not welcome.
Often the answer was that the team didn’t even know the test existed and had no use for it. That is how the suite shrank on its own: the stable tests stayed, the rest dropped out. Eventually the suite was rewritten from scratch.
The new suite deliberately covered only the money path. Instead of many separate Given-When-Then cases, it became one continuous test that walks through the happy path and breaks exactly where something is wrong. That makes it much easier to pin down the defect.
Shift Left, but Look Right: Smaller Tests and Production Monitoring
The most effective partner for smaller tests is good observability in production. The team made two moves. First, shift left: tests became smaller and moved to lower levels to get feedback earlier.
Then came “shift left, but look right”. The team expanded logging and observability, and above all monitoring and alerting. The rule: whenever a customer sees an error message, a red light should go on somewhere. Pure network issues are the exception, for example when someone’s train goes through a tunnel.
Releasing often lowers the risk of each deployment. With many small changes, any single release can only break a little. The weeks of acceptance testing the monolith used to need were no longer necessary.
That changes what end-to-end tests are for. For an e-commerce business, it can be enough to deploy to production when in doubt, wait for an alert and roll back if needed. Recreating every rare failure on a test stage beforehand costs far more than the damage it could prevent.
Why Rollback Is a Requirement, Not a Nice-to-Have
None of this works without a clean deployment mechanism. While moving to microservices, you redeploy parts of the system all the time, and you can’t afford downtime. That calls for techniques like blue-green deployment or canary releasing.
Some lessons still hurt. Database migrations need a lot more attention and care in this setup. The payoff is worth it, though: you can deploy something to production and check there whether it works.
Teams that can’t make rollback work often build huge test environments instead, at great expense of energy. That is frequently the more expensive route, and it only covers up the real gap.
Cleaner Tests Motivate Less Than You Might Think
Faster feedback motivates developers. Cleaning up tests rarely does. People are pleased when they cut the test suite’s run time in half, because everything moves noticeably faster. Replacing one big end-to-end test with six unit tests, on the other hand, is hardly anyone’s idea of fun.
The reason is how many developers see their job. They want to write code that solves problems, not code that causes them. A failing test first feels like a problem, not like help. Champions of good test code are rarer than most people think.
The practical takeaway: if you want better test quality, tie it to a benefit people can feel, such as shorter run times and faster feedback. Cleaning up for its own sake, with nothing visible to show for it, is a hard sell.
Cleanup Needs the Same Planning as Feature Work
Cleanup sprints often fail because nobody plans them. The time is always too much and too little at once: you’ll never clear up the really big mess, and most of the time you don’t want to anyway.
Michael compares it to a kid’s bedroom. Tell someone to tidy up, and the first question is: where do I start? The job is too big and has no structure.
The flawed assumption is that you already have a plan for the cleanup. You don’t, because your head is in feature mode. Paying down technical debt takes the same planning effort as building features. Without that plan, the time you were given simply evaporates.
Why Test Data Remains the Hardest End-to-End Problem
The biggest open issue is not test technique but getting the right data onto the stages. Hundreds of stores with different product ranges, different logistics setups and a huge number of configurations all have to be represented.
Warehouses that deliver directly to consumers sometimes work differently from other sites, and the IT process has to reflect those variations. Providing an environment that can supply every variant at any time takes a lot of effort.
This is where dedicated testers come back into the picture. For a long time REWE had no testers at all, only development teams that did their own testing. Today there are a few very good testers who test end to end, partly on the mobile devices warehouse staff use. Not every developer has those devices on their desk.
Exploratory Testing Belongs in the Development Teams
One concrete recommendation is to build exploratory testing into the teams themselves. A development team should now and then set aside half an hour in a sprint to put its own product through its paces.
A good charter often comes straight out of the code. While reading an API, a developer notices that the authorization might not work the way it should. Inconsistencies like that are quick to uncover through exploration.
Simple for the Business, Complex in the Code: The Legacy Trap
Some requirements are trivial from a business point of view and surprisingly complicated technically. For the business, the change is easy. In the code, it isn’t. Michael is clear about this: things that are simple for the business shouldn’t really be technically complex.
The cause is history. Legacy systems that have been rebuilt three times and passed through four teams are hard to change. Platform engineers and the business side alike often need help understanding where the actual testing problem lies.
A well-cut microservice and data architecture pays off most under outside pressure. The VAT cut could be implemented surprisingly quickly because the data architecture allowed it. Regulatory changes also go faster when the organization and the code fit together well.
Frequently Asked Questions
Do microservices solve a technical problem?
No. The driving force is usually organizational: Many teams need to work on a codebase in parallel without blocking each other. In a monolithic architecture, this ends in merge hell when four people refactor the same file in different ways. Microservices turn Conway’s Law into an architectural principle: clearly defined responsibilities, independent deployment. The trade-off is coordination, API management, and more complex code organization.
What does it mean when a test is only taken seriously after three failed runs?
It shows that the test suite has lost its value as an early warning system. In the practice described, it was more cost-effective to run the suite again than to investigate the cause after every failure. Often, the problem was a slow database at the wrong time. In such cases, testing tends to focus on the stability of the infrastructure rather than the code.
How do you gradually break down a monolith that has grown over time?
Using the Strangler Fig pattern: The existing code remains untouched because any change would destabilize the cemented structure. A bypass leads from the monolith to the new service, which takes over a function. The first services were essentially databases with APIs that encapsulated the complex database structure. Only then did the business logic migrate, until all that remained in the monolith was the rendering of HTML.
How many end-to-end tests does a service really need?
One test per risk, provided the risk is already covered further down the pyramid. If a feature is validated at the API level or in a unit test, ten end-to-end variations of the same case are unnecessary. The newly written suite deliberately tested only the money path, as one continuous end-to-end test that fails exactly at the point of the error.
How do you scale back a bloated, fragmented test suite?
Through quarantine instead of maintenance. A test that fails on the first attempt is placed in quarantine, and the team likely responsible is asked to stabilize it. Often, the team wasn’t even aware of the test and no longer needed it. This is how the suite shrinks organically. The background to this was a growth from four to six teams to twelve to fourteen teams, which made centralized maintenance impossible.
Can monitoring and alerting replace end-to-end testing?
Partially, yes. The rule is: As soon as a customer sees an error message, a red light goes on somewhere, except for pure network issues such as a train going through a tunnel. For an e-commerce business, it may be sufficient to deploy to production, wait for the alert, and roll back if necessary. Reenacting every rare error in advance on a test stage is disproportionate to the potential damage.
What prerequisites are needed to detect errors only in production?
A clean deployment mechanism without downtime, that is, methods such as blue-green deployment or canary releasing. Database migrations, however, require significantly more attention and care. Those who cannot make rollbacks a reality often build massive test environments instead, which consume a great deal of resources: this is frequently the more expensive route, which merely masks the actual problem.
Why do cleanup sprints for test code fail so often?
Because no one plans them. Reducing technical debt requires the same planning effort as feature development, but in “feature mode,” this plan doesn’t exist, and the time set aside goes to waste. Added to this is the issue of motivation: Everyone is happy when test runtime is cut in half, but hardly anyone is pleased to replace a large end-to-end test with six unit tests. Therefore, link test quality to tangible benefits such as faster feedback.


