A cloud migration spans three interdependent levels: infrastructure, software architecture and organization. On the technical side, it usually starts with lift and shift, moving an application to the cloud unchanged, followed by step-by-step refactoring from monolith to microservices along functional domain boundaries. Without automated CI/CD pipelines, resilience testing and organizational change, the migration fails.
Key Takeaways
- Greenfield rebuilds fail again and again: teams that build a new system in parallel to the monolith usually abandon the project after about two years.
- A working CI/CD pipeline has to be in place before any migration work, not after, because a system with many microservices and databases simply can’t be maintained without automation.
- Cloud migration changes job profiles in concrete ways: DBAs lose tasks when Oracle is replaced by a managed service, and developers who only deliver JAR files have to widen their area of responsibility.
- Classic test suites lack resilience tests, because container failures, higher latency and brief outages hardly ever happened on-premise but are part of normal operation in the cloud.
Why Companies Move to the Cloud
Three motives drive most cloud migrations, whether they stop at lift and shift or go all the way from monolith to microservices: hoped-for cost savings, flexibility and the pressure to follow everyone else. The first two are tangible. The third is a hype effect that builds its own momentum.
On cost, many expect their own infrastructure and the staff that runs it to go away. On paper, the numbers look good. In practice, the picture often turns out differently, because new complexity appears somewhere else.
Flexibility is the stronger argument. The big hyperscalers run data centers all over the planet. If you want to serve customers in Australia, you can’t sensibly do it from Switzerland or Germany. Hardly any single company has that kind of geographic reach on its own.
A VM in the Cloud Is Just a Computer Somewhere Else
Technically, a virtual machine in the cloud is at first exactly that: a computer in a different location. This sober view helps break a common illusion, the idea that the cloud solves the old problems automatically.
The move brings problems to light that were always there locally. Security, access control, update cycles. Suddenly it matters how often a VM gets rebooted or what day-two operations look like for a database system or a Kubernetes cluster.
These questions existed on-premise just the same. Linux systems had to be updated every two weeks before, too. The cloud only makes these duties visible, because they can no longer hide in your own data center. That, together with the cultural issues, is what makes a migration hard.
Three Construction Sites: Infrastructure, Architecture, Organization
Every cloud migration breaks down into three areas, and only one of them is purely technical.
- Infrastructure: The path from local to cloud infrastructure. Cloud providers offer an API that lets you set this up fairly quickly. The managed services behind it, however, bring their own complexity, which makes the whole effort harder.
- Architecture: Monoliths that have grown over many years are often badly built. What many people overlook: a monolith can have an architecture too, for example through clean Java packages. The task is to split such monoliths up sensibly.
- Organization: Over the years, organizational units have formed around on-premise databases that no longer fit the new world. This is about ownership and agile ways of working.
Where the main problems lie differs from company to company. Banks and insurers carry especially long-lived systems and equally long-lived organizations. You start at the weakest point.
How a Cloud Migration Works Step by Step
The sensible way in is lift and shift: you take the existing application, containerize it if needed and start it in a cloud VM. This first step is mainly about gaining experience.
Further stages build on it. In re-hosting, the monolith moves to the cloud, which already requires a CI/CD pipeline in place. In re-platforming, the whole thing goes into a Kubernetes cluster, and logging and monitoring run on the cloud provider’s services.
Only after that comes refactoring. This is the real architecture question, where you improve the structure if you want to. There are good reasons not to take this last step at all.
From Monolith to Microservices: Greenfield Fails, Step-by-Step Cutting Works
Trying to rebuild a monolith from scratch with a separate greenfield team goes wrong in practice. The pattern is familiar: three teams maintain the monolith, a fourth is supposed to reimplement it with clean boundaries. After about two years, projects like that are usually abandoned.
The path that works goes through the existing monolith. You carve out the functional domains step by step, spread over several iterations, not in a single sprint.
Once the interactions between domains are minimal, the interfaces are clean and the database schema is separated, a domain can be cut out and packed into its own container. The cut is functional, not technical. That is the basic principle of microservices.
One example from banking is onboarding. A customer comes in, their data is captured, and first products are offered based on the investment amount and risk appetite. This onboarding process is a self-contained domain you can start with.
The Database Is the Hardest Construction Site
Many companies run a large Oracle platform, with DBAs in front of it who often block architecture changes. The issue is technical and organizational at once.
Technically, you might be able to port an Oracle database to a big hyperscaler, but usually that’s not what you want. This is exactly one of the biggest tasks in any migration.
A data migration from one database system to another has to make sure the data is still the same afterwards. That starts with character sets and field lengths and goes much further. When you move from an old system to PostgreSQL, things will definitely happen that you have to catch before going live.
A Microservice System Can’t Be Maintained Without a CI/CD Pipeline
Before you migrate anything, set up a continuous delivery pipeline. It is the first step toward the cloud, not the last.
A system of 40 or 50 microservices with replicas and databases underneath can no longer be maintained without tests and a CI/CD pipeline. Automation has to be built in from the start.
“When you build a microservice like this, you always set up the CI/CD pipeline first and only then start writing the code.”
(Christopher Schmidt)
The pipeline includes unit tests, component tests and properly tested microservices. This is exactly where the errors that a data migration is guaranteed to produce will show up. And the pipeline itself needs a defined process. If you haven’t worked out a clear approach beforehand, you build something that keeps breaking.
Migration Changes People, Not Just Machines
A cloud migration can’t be run as a stealth project. If a single team decides that local IT is too slow and quietly moves to the cloud, it won’t last long. Because of the organizational consequences, management has to back such projects.
The classic silo structure doesn’t fit the new world well. An infrastructure team, an ops team, developers who neither know base images nor write deployment manifests because ops handles that. This separation needs to go, and that only works by adjusting objectives and incentives. That takes time.
Specific roles change. Today, DBAs are responsible for query performance and optimize queries directly in the database system. After a switch to a managed service such as Amazon RDS, they are no longer needed to the same extent. Developers whose self-image ends at “my artifact is a JAR, nothing else concerns me” have to move, too.
At the same time, nobody can know everything in a complex environment. Core expertise has to stay with specific people in the teams. The idea that everyone has full access to all knowledge doesn’t work.
Resilience Testing Belongs in Every Cloud Pipeline
The cloud adds a test category that classic functional testing doesn’t cover: resilience against latency and failures. You probably had functional, unit and component tests before. That isn’t enough anymore.
In a cluster, containers are sometimes unavailable, for example because a node has rebooted or died. An API goes down briefly, latencies rise. A monolith turns into separate processes spread across different machines and sometimes different data centers. That creates latencies and failure situations that didn’t exist before.
You feel this most in a hybrid cloud setup. If part of the application runs in the cloud while service or database systems stay on-premise, the connection goes through VPNs or SSL. Latency then jumps from around ten to a hundred milliseconds. That is often enough to make thread pools overflow.
Resilience tests like these belong in the CI/CD pipeline. Leave them out, and you’ll wonder later why parts of the new software keep crashing. Something similar may have happened in your own data center, but you ignored it, or it happened less often.
Maintenance Window or Fast Response
Managed services let you bundle updates into maintenance windows, say on Saturday or Sunday. That sounds convenient, but it has a price you need to weigh deliberately.
If you schedule the window for the weekend, a vulnerability can stay online until then. Fixed vulnerabilities are listed in the release notes, and anyone who wants to break in checks exactly those well-known Java or Linux issues to see whether a system is still exposed.
The recommendation is clear: work without a maintenance window and deploy fast. An application that can cope with short outages and higher latency doesn’t need the planned window as much anyway. That brings the question back to the resilience of the architecture.
Frequently Asked Questions
Do Companies Really Save Money by Migrating to the Cloud?
Not necessarily. Three motives drive most migrations: the hope of cost savings, flexibility, and the pressure to follow the crowd. The cost analysis looks good on paper because it eliminates the need for in-house infrastructure and the associated staff. In practice, however, new complexities arise elsewhere. The more tangible argument is the geographic reach of hyperscalers: Customers in Australia can hardly be served effectively from Germany.
Does moving to the cloud automatically solve existing operational problems?
No. A virtual machine in the cloud is, at first, just a computer located somewhere else. Security, access control, and update cycles remain the same; they simply become visible. Suddenly, what matters is how often a VM is booted and what the day-two operations look like for databases or Kubernetes clusters. Linux systems had to be updated every two weeks before, too; it’s just that this could be hidden away in our own data center.
How can you tell if a domain can be carved out of a monolith?
Three conditions must be met: minimal interaction between domains, clean interfaces, and a separate database schema. Only then is it worthwhile to package the domain into its own container. The cut is functional, not technical. A good introductory example from the banking sector is onboarding: capturing customer data and offering initial products based on investment amount and risk tolerance.
Does the CI/CD pipeline have to be in place before the actual migration?
Yes. The continuous delivery pipeline is the first step, not the last. A system consisting of 40 or 50 microservices with replicas and underlying databases can no longer be maintained without testing and automation. The pipeline includes unit tests, component testing, and thoroughly tested services. The pipeline itself also requires a defined process; otherwise, the entire structure will constantly break down.
Why is the database the toughest part of a migration?
Because the issue is both technical and organizational. Organizational units have grown up around large Oracle platforms over the years, and DBAs often block architectural changes. Technically, an Oracle database could be ported, but in most cases, that’s not what you want. A data migration must ensure that the data remains the same afterward: character sets and field lengths are just the beginning. When switching to PostgreSQL, something is bound to stand out.
Why are functional and unit tests no longer sufficient for cloud applications?
They don’t account for resilience against latency and outages. In a cluster, containers may be temporarily unavailable because a node is rebooting or has failed, and APIs may experience brief outages. In a hybrid scenario with VPN or SSL, latency jumps from about ten to one hundred milliseconds, and that alone is enough to cause thread pools to overflow. Such tests belong in the pipeline.
Can a single team launch a cloud migration on its own?
It won’t last long as a “submarine project.” Due to the organizational implications, senior management must support the initiative. The classic silo structure (an infrastructure team, an ops team, and developers who neither understand base images nor write deployment manifests) doesn’t fit into the new world. It can only be broken down through adjusted objectives and incentives, and that takes time.
Are weekend maintenance windows a good solution for updates?
They come at a cost: a known vulnerability remains online until the window opens. Fixed security vulnerabilities are listed in the release notes, and attackers specifically target these known Java or Linux standard issues. It makes more sense to work without maintenance windows and deploy quickly. An application that can handle brief outages and higher latencies hardly needs the scheduled window anyway.


