Skip to main content

Search...

AI Code Remediation: Two Thirds of the Fixes Work

AI code remediation fixes about two thirds of static analysis findings well. The rest is faulty, and models almost never admit they don't know.

• • Updated: • 11 min read
Cover of the expert talk on 'AI Code Remediation: Two Thirds of the Fixes Work' with Benjamin Hummel and Richard Seidl.

AI code remediation means letting a large language model automatically fix quality problems found by static code analysis. The model receives the affected code and the description of the problem and generates a proposed fix. Current models deliver usable results in about two thirds of cases, while in the remaining third the output is wrong or not even valid code.

Key Takeaways

  • AI-generated fixes for static analysis findings work well enough in about two thirds of cases to be adopted as they are or with small adjustments.
  • The remaining third produces unusable or simply wrong code, and reviewing it takes as much effort as fixing the problem by hand from the start.
  • Large language models almost never decline to answer: in a benchmark of one hundred findings, the models said “I don’t know” only twice and produced faulty code instead.
  • Closed-source code fares much worse in LLM-based security analysis than publicly available code, because it is simply not part of the training data.

Static Analysis Produces More Findings than Anyone Can Fix by Hand

On large systems, static analysis turns up so many issues that fixing them manually hits a hard volume limit. That is the starting point for AI code remediation. A typical rule set contains hundreds to more than a thousand rules. The result: a few thousand, a few tens of thousands, sometimes millions of findings in a single codebase.

These findings differ widely in weight. Some are trivial, like a missing comment on a method. Others are serious, such as a potential SQL injection or a possible null pointer dereference. Severity is all over the place.

In practice, the flood often leads to avoidance. Benjamin Hummel sees two common patterns: teams switch rules off, or they stop looking at the results altogether. Switching rules off is at least the more constructive of the two. His advice is to go after the serious findings first on purpose and not let the sheer mass overwhelm you.

Why Language Models Help Fix Findings at All

Language models move remediation from the pure syntax level to a level that understands code and natural language at the same time. That was hardly possible with classic static analysis.

Static analysis is good at parsing code, building syntax trees and tracing data flows. But it stays at the level of syntax and language semantics. The natural language in comments, names and descriptions was out of its reach.

A missing method comment shows how big the leap is. Older documentation generators took the method name, inserted spaces and produced mostly meaningless text. Feed the whole method body into a language model today and you get a useful summary with meaningful details. Code understanding has come a surprisingly long way here.

So the real value lies less in generating new code than in improving existing code. Instead of starting from scratch, you get a proposal that you only need to review and refine.

How AI Code Remediation Works in Practice

The naive approach turns out to be surprisingly workable. You give the model the affected piece of code, the problem description from the static analysis tool and, if available, the explanation of why it is a problem and what typical fixes look like. The model returns the corrected code.

This works across a wide range of problem classes, from adding comments and removing unused imports to more complex refactorings.

Long methods show how far it goes. A model suggests extracting two sensibly named methods, and the result is internally consistent. That is no small thing, because even doing it by hand, deciding where to split a method takes real thought.

Two Thirds of the Fixes Work, the Rest Is Noise

How reliable is AI static analysis when it comes to fixing code quality problems? A systematic benchmark shows that AI-assisted remediation works really well in about two thirds of cases. The remaining third is noise.

For the evaluation, around 100 rule violations were taken from a larger project and run through several models. Models seem to change every month, so comparing more than one was necessary. The assessment itself was manual work: for each model, every one of the 100 cases was checked by hand to see whether the fix made sense, whether the code was still valid and whether it still did what it did before.

The good two thirds: these proposals can be adopted directly or reused with small adjustments, such as a rename. That saves a noticeable amount of work.

The bad third is trickier. It produces unusable code, sometimes code that is not even valid. Reviewing such a proposal and checking it for subtle errors often takes as much effort as fixing it yourself, or more.

Models Want to Please Instead of Admitting They Are Stuck

One core problem: language models are heavily trained to present a solution, even when they clearly don’t understand the problem. Even if the prompt is open and explicitly allows the model to decline with “I don’t know,” that almost never happens.

Across all experiments, there were only two cases in which a model admitted it did not know the solution. Otherwise the models tried to force out some kind of answer.

That behavior causes more trouble than an honest refusal would. If the model said “I don’t know,” you would immediately know to fix the spot by hand without thinking twice.

“The models are very strongly trained to present a solution, even when you get the impression they can’t do it at all. They want to please.”

(Benjamin Hummel)

Mainstream Languages Work, Niche Languages and Closed Source Fall Short

The quality of the proposals depends directly on how much training data exists for a given language. Mainstream languages such as Java or JavaScript work well because they are heavily represented in the training data.

Niche areas look different. For ABAP-based SAP development, for example, there is much less open-source training data than for Python. Expect weaker results there. This matches how code generators behave in general: they produce noticeably worse output for less common languages.

The effect is even stronger for in-house code that has never been public. An ongoing thesis on security analysis with language models compared two data sets: a widely used benchmark and in-house code that was guaranteed not to be in any training data. Results on the in-house code were significantly worse.

This gap matters in practice. A large share of business software is developed as closed source and never shows up in training data. The methods are weakest exactly where they would have to work every day.

Many AI Review Tools Just Repackage Classic Linters

Many tools that promise automated code reviews on platforms like GitHub run a classic open-source linter in the background. A language model then rewrites the linter’s findings in natural language.

That creates the impression that the AI is doing the actual analysis. In reality, it mostly repackages the linter’s results. “AI” on the label has become a selling point.

The quality of genuine AI findings is still mixed. Tools report problems that aren’t problems, and many real problems go unnoticed. To be fair, AI shares this lack of completeness with every static analysis. None of them is complete.

Fix Findings Where You Are Already Working on the Code

For old, grown systems, AI is not the obvious answer. The proven strategy still applies: fix problems exactly where work is happening anyway.

If you touch a tax calculation module because of a change in the law, it makes sense to clean up the findings in that very module. Cleaning up parts nobody touches is rarely worth it. The exception is high-risk issues such as SQL injection, which you look for specifically. Step by step, this approach gradually leads to a better system.

Should legacy systems be cleaned up proactively with AI? The error rate argues against it. Even if 99.9 percent of the fixes are correct, clearing 100,000 findings will introduce 100 new bugs. Nobody wants that.

For the worst parts of a system, deeply nested modules whose architecture is almost impossible to follow, AI does not help at the moment. The only option is the serious decision to rebuild those parts and secure them with tests.

Generated Code Isn’t New, but AI Undermines the Trust in It

AI-generated code raises an old question with new urgency: how readable does code have to be if no human looks at it anymore? Generated code has been around for a long time, such as parsers built from grammars or application code generated from domain-specific languages. Different readability standards apply to that kind of code, because when something changes, you adapt the source model and regenerate.

The decisive difference with AI is trust. Classic generators and compilers are deterministic: same input, same output. They are developed and tested over many years, which is why hardly anyone suspects the compiler first when looking for a bug. Language models, on the other hand, hallucinate, make mistakes and produce different outputs because of their probabilistic nature.

That creates a concrete risk. If AI builds a system that you as a human no longer understand, and you also leave every change to the AI, you are left empty-handed at the point where it can’t go any further. That is exactly where testers become more important, because the question of whether the system still does the right thing remains open.

Frequently Asked Questions

Why Does Addressing Findings from Static Analysis Fail Due to the Volume?

A typical rule set comprises hundreds to over a thousand rules. In a single codebase, this results in several thousand, several tens of thousands, or in some cases millions of findings. There are two common reactions to this: teams either disable rules or stop looking at the results altogether. Disabling rules is the more constructive approach. It makes sense to address the most serious findings first.

What can language models do with code that traditional static analysis cannot?

They capture the natural language embedded in comments, names, and descriptions. Static analysis parses code, constructs syntax trees, and tracks data flows, but remains confined to syntax and linguistic semantics. Take method comments, for example: Earlier documentation generators simply parsed the method name and produced meaningless text. A language model generates a useful summary with meaningful details based on the method’s content.

Is it worth salvaging a flawed AI correction suggestion?

Usually not. In about one-third of cases, the result is unusable code, sometimes code that isn’t even valid. The effort required to review such a suggestion and check it for subtle errors is often just as great as, or greater than, simply fixing the section yourself. The usable two-thirds can be adopted directly or with minor renaming.

Do language models admit when they can’t solve a problem?

Practically never. Across all experiments, there were only two cases in which a model admitted it didn’t know the solution, even though the prompt explicitly allowed for a “I don’t know” response. Instead, the models try to force out some kind of solution. An honest admission of failure would be more useful, because then you’d know right away: time to do it manually.

Do AI-assisted fixes work equally well in every programming language?

No, the quality depends on the amount of training data available for each language. Java and JavaScript work well. For niche areas like SAP development based on ABAP, there is significantly less open-source training data than there is for Python, for example, so the results are correspondingly weaker. The quality drops even more sharply with proprietary closed-source code that doesn’t appear in any training data.

Do tools that promise automated AI code reviews really perform their own analysis?

Often not. Many tools on platforms like GitHub run a classic open-source linter in the background, whose findings are then rephrased linguistically by a language model. In other words, the AI primarily repackages third-party findings. Genuine AI findings remain mixed: problems are reported that aren’t actually problems, and many real ones go undetected.

Does it make sense to use AI to clean up all the findings in a legacy system across the board?

No, the probability of errors argues against it. Even with a 99.9 percent accuracy rate for fixes, clearing up 100,000 findings would mathematically result in 100 new errors. A more viable approach is to fix problems where work is already being done, such as in the tax module, which is being revised due to a change in the law. Exceptions are high-risk issues like SQL injection, which are examined specifically.

Share this page