Skip to main content

Search...

Legacy Modernization: Why Missing Experts Are the Bottleneck

Legacy modernization usually stalls on missing subject matter experts. How RAG makes legacy code queryable and why working in slices lowers the risk.

• • Updated: • 13 min read
Cover of the expert talk on 'Legacy Modernization: Why Missing Experts Are the Bottleneck' with Erik Doernenburg and Richard Seidl.

Legacy modernization is the work of moving old software systems, such as COBOL mainframes or early Java monoliths, onto modern architectures. The biggest bottleneck is usually missing expert knowledge about those old systems. Retrieval Augmented Generation (RAG) addresses this by turning the legacy code into a knowledge graph that developers can query directly with AI.

Key Takeaways

  • In one real legacy modernization project, the main reason for running three quarters behind schedule was a shortage of subject matter experts, not a lack of technology.
  • Retrieval Augmented Generation eases the SME bottleneck: the legacy code goes into a knowledge graph, and developers can ask the codebase targeted questions in a chat without retraining the model.
  • Microservices named after nouns (Customer Service, Product Service) tend to couple more tightly than services named after verbs, because a business process then usually touches all of them at once.
  • Too many tokens in a prompt lower answer quality, because the model gives less weight to information in the middle of the context window than to what comes at the beginning and the end.

When Software Counts as Legacy

Any legacy modernization effort starts with a definition, and age alone doesn’t provide one. Software becomes legacy through its technology and through how hard it has become to change. COBOL mainframe systems clearly qualify. So does software written around the turn of the millennium in early versions of Java or .NET.

Erik Doernenburg, who works at the consultancy ThoughtWorks, sees an especially large number of these systems in insurance and banking. Much of what clients want to modernize has been around for a long time and now needs to be changed, improved, or moved to the public cloud.

Two things keep teams from touching such systems. The first is missing tests. Nobody dares to make changes because nobody fully understands how the system works anymore. The people making changes today are rarely the ones who wrote the code. Knowledge has been handed down over the years, often from one service provider to the next.

The second is a missing grasp of the business domain. Insurance is a prime example: after acquisitions and mergers, nobody knows what the policies sold 15 years ago actually looked like. On top of that, technical concerns and business logic are often tangled together in the code. Anyone trying to understand the business logic keeps tripping over infrastructure code for the database or the user interface.

Subject Matter Experts Are the Real Bottleneck

In large modernization projects, the main problem is not the technology. It is access to people who can explain the old system. An analysis at a major German client made that plain. The project was three quarters behind schedule, costs were high, and the cause was a lack of subject matter experts.

These experts explain to the teams writing the new software how the old software works. That is exactly where the bottleneck forms. Hiring or moving people around internally doesn’t fix it. The experts have often left the company, or they are so scarce that you can’t get time with them when you need it.

“Subject matter experts, people who can explain to the teams writing new software how the old software works: that was the bottleneck.”

(Erik Doernenburg)

Why Modern Microservices Turn into Legacy Too

Microservices were supposed to make legacy as we know it a thing of the past. Small, deployable units could be swapped out for newer technology after five or ten years without touching the whole system. Those firewalls between parts let you recombine pieces and throw them away one at a time.

For many companies, that works. Well-cut services live quietly at the edges, run stably, and don’t get in the way of code that changes more often. Other companies end up with a distributed monolith, where the services are so tightly coupled that a change in one drags three others along. Sooner or later, systems like that become legacy again.

Erik offers a simple rule of thumb for spotting this early. It is all in the service names.

NamingLikely Consequence
Service is named after a noun (Customer Service, Product Service)Tight coupling is more likely, because business processes run across several services
Service is named after a verb (e.g., Order Capture)Isolated services that cover a complete business process are more likely

If changing a business process means reaching across the customer service, the product service, and more, the coupling is too high. Services that wrap a complete process can be replaced later without disturbing the others. Microservices were never meant to mirror an entity-relationship model. The idea was to break business functionality into pieces that are as autonomous as possible.

Legacy Modernization in Slices Instead of One Big Bang

The most important strategy in modernization is to avoid doing it in one big step and to work in slices instead. A mainframe modernization that runs for two or three years, delivers nothing along the way, and switches everything over on a single cutover day is a high-risk bet that hardly anyone is willing to make anymore.

Instead, teams pick out individual pieces of functionality that can be looked at and moved on their own. That takes enough people in the company who understand the business functionality.

The longer a team works together in a similar setup, the better it gets at forecasting. It keeps the slice size more consistent and can estimate how much work is left. If a slice takes about three to four weeks and everyone knows how many are still to come, you can plan, even when one of them runs long. That builds trust with stakeholders on the business side.

It also avoids what is known as watermelon reporting: green on the outside, red on the inside. Management hears that everything is on track until, three weeks before the end of a two-year program, half a year suddenly turns out to be missing.

Tests as a Safety Net for the Rebuild

You can’t modernize legacy safely without tests. The more test coverage a legacy system has, the better. In some cases, teams write tests first just to have any safety net at all.

End-to-end tests are valuable in modernization because they show after every change whether the system still works. Another strategy is to run the old and the new system side by side for a while, three to six months is realistic. That shows whether the new system produces the same output for every input from daily operations.

Month-end and year-end closing deserve a close look. Erik remembers a banking client where things had been cleaned up and deleted. At the end of the month there was trouble: a process had been sending messages that nobody could trace anymore, but that mattered to the receiving side.

The same discipline applies to the new code. If you move fast at the start and pile up technical debt, you lose speed later. The teams rely on test-driven development, write the code with unit tests, and add end-to-end journey tests that run from start to finish. Edge cases belong in the fine-grained tests, and integration tests cover the middle layer for performance reasons. That is the classic test pyramid.

How LLMs Help You Understand Legacy Code

Large language models help most not when writing new code, but when making sense of old code. In forward engineering, meaning code generation, the productivity gains are modest. A Microsoft study reported a 55 percent speed increase with Copilot, but it measured the pure act of programming, and the task was a web server in JavaScript, a problem with hundreds or thousands of examples online.

With specialized business logic, the picture changes. There is far less public source code for the business processes of an insurance company, so there is less of it in the models, and the productivity gain shrinks accordingly. Even so, hardly any developer who has used such a tool wants to give it up, and compared with personnel costs, the license fees barely matter.

The real payoff for legacy lies elsewhere. Erik sees it in a pattern called Retrieval Augmented Generation.

Retrieval Augmented Generation: Pulling Knowledge Out of the Legacy System

Retrieval Augmented Generation tackles the hallucination problem by having the model find the answer in text it is given instead of generating it from its own training. Rather than asking “Where in the codebase does X happen?”, the prompt says, in effect: answer the question using the following text. And that text contains the company’s own documents.

Technically, this runs on embeddings. For each document, a model calculates a vector of numbers, and documents with similar content end up pointing in a similar direction. Those vectors go into a vector database. Before every request, the system calculates the vector for the user’s prompt, looks up the matching documents, and adds them to the prompt. That is the augmentation. Nobody has to train their own model, and the company’s documents don’t have to be inside the model.

For legacy code, the teams turned this into an internal toolkit. It is not a finished product but a set of scripts, instructions, and a chat interface. At its heart is a knowledge graph fed with source code, existing documentation, and the output of classic reverse engineering tools such as dependency analysis.

The toolkit works in two modes:

  • Generating briefing documents. A prompt such as “Describe the capabilities of an admin user” produces a few pages explaining how an admin user differs from other users in the code. At one client at least, subject matter experts spot-checked these documents and found no major errors.
  • Chat dialog. Developers ask the tool the same way they would ask a human expert. From the prompt, the system finds matching nodes in the knowledge graph and follows the edges, for example to functions that call other functions. That way it collects exactly the knowledge that belongs in the prompt.

Small, Precise Prompts Beat a Huge Context

More context is not automatically better. Dumping in as much unstructured information as possible is not the best approach. Often it works better to put less, but more carefully chosen, information into the augmentation.

The reason is the “lost in the middle” problem. When you pass in a very large number of tokens, the model seems to weigh information at the beginning and the end more heavily than what sits in the middle. Some people have started compressing their prompts with a smaller model for that reason.

With source code, there is a more elegant way. Instead of compressing with a sledgehammer, you can use existing reverse engineering knowledge about dependencies to enrich the prompt in a targeted way. One especially telling signal comes from version history: which files get checked in together. If three classes are always committed together, they belong together in terms of content, often more clearly than any static call graph analysis would show. Signals like that make prompt augmentation much better, because small, sharp prompts carry exactly the knowledge the answer needs.

Cost points the same way. Large input context windows, up to half a million tokens at Google in some cases, start costing real money once you use them often. That looks smaller next to personnel costs, but a precise prompt is usually still the better choice.

Frequently Asked Questions

Is a Java system from the early 2000s already considered legacy?

Often, yes. Legacy is not defined by a specific age, but rather by the technology used and how difficult it is to modify the system. COBOL mainframes clearly fall into this category, but so do applications built with early versions of Java or .NET. This is often compounded by the fact that business logic and infrastructure code for the database or user interface are intermixed in the source code.

Why do large modernization projects tend to fail because of people rather than technology?

Because there is a lack of subject matter experts who can explain how the legacy system works. At a major German client, a project was three quarters behind schedule, and the analysis attributed this to a lack of subject matter experts. This bottleneck cannot be resolved through new hires or internal reassignments: the people in question are often no longer with the company.

How can you tell early on that microservices are too tightly coupled?

A clue lies in the service names. Services named after nouns, such as “Customer Service” or “Product Service,” more often lead to tight coupling because business processes then span multiple services. Verb-based names like “Order Capture” indicate services that encapsulate an entire process and can later be replaced without affecting others.

Is a mainframe migration worth it as a single, large-scale transition?

No. A program that runs for two or three years, delivers nothing in the meantime, and switches everything over all at once on Day X is a high-risk gamble. It’s better to migrate individual functionalities in isolation. If each phase takes three to four weeks and the number of remaining phases is known, it’s possible to plan effectively and avoid “watermelon reporting.”

Why do teams run the old and new systems in parallel for a while?

To verify that the new system produces the same results for all inputs from daily operations. Three to six months of parallel operation is realistic for this purpose, and month-end and year-end closings deserve special attention. In a banking project, after cleaning up at the end of the month, it became apparent that a process was sending messages whose origin no one could trace anymore.

How much productivity do AI assistants really bring to writing new code?

Less than individual figures suggest. A Microsoft study reported a 55 percent increase in speed with Copilot, measured based on the pure act of programming and using a JavaScript web server, for which there are hundreds to thousands of examples online. For specialized domain logic, such as insurance processes, the gain is smaller because the models contain significantly less source code for these use cases.

Does a model need to be trained specifically to ask questions about one’s own codebase?

No. In Retrieval Augmented Generation, the model is not supposed to generate the answer on its own, but rather find it in the provided text. Vectors calculated from documents are stored in a vector database; the system searches for the relevant documents based on the prompt and returns them. The company’s documents do not need to be included in the model for this.

Does a larger context window improve answer quality?

Not automatically. If a very large number of tokens are provided, the model weights information at the beginning and end more heavily than that in the middle. A targeted selection makes more sense: For source code, the version history is helpful, because files that are always checked in together belong together in terms of content, often more clearly than a static call-graph analysis reveals. Large context windows also incur immediate costs.

Share this page