Generative AI testing means using language models on purpose for specific jobs: reviewing test cases, finding missing edge cases, writing boilerplate code and discussing code in a review dialog. It pays off most when the requirements going in are precise. Unit tests generated blindly from existing code can reach high coverage and still hide logic errors instead of exposing them.
Key Takeaways
- AI-generated unit tests that are derived from existing code don’t uncover bugs, they cement them, because the model treats the faulty code as the truth.
- Generative AI helps most as a way into test automation: it takes over boilerplate and glue code and lowers the barrier for testers without deep coding experience.
- Code completion tools such as Codium know the repository and suggest code built on your own page objects and helper classes, not on someone else’s training data.
- Documentation generated only to feed an AI search, and never read by a person, piles up a growing layer of unchecked text with no value.
- The biggest hope for GenAI in software development isn’t speed, it’s paying down technical debt through better decisions in the code.
Generative AI Testing Is Here to Stay
For Matthias Zax, generative AI testing is no passing trend. It has found its place in software testing and will shape the next few years, and companies that pass on it will fall behind their competitors. Matthias is a test automation engineer at Raiffeisenbank International and works as an engineering coach for several teams.
The range of uses is wide: test case design, feature files in BDD style, help with the automation code itself. There are hardly any hard limits left. There are limitations, though, and you need to know them before you rely on the tools in production work.
The key distinction is between sensible use and blind trust. A language model produces output, but a human still has to validate it. Matthias puts it this way:
“If I ask about risks, I get risks. That doesn’t mean they are risks.”
(Matthias Zax)
How Gen AI Helps Testers Get into Test Automation
The biggest payoff comes where testers lack coding skills. Many of them come from manual or exploratory testing and are suddenly expected to automate. But test automation is software development, and that jump is hard.
The tooling around it doesn’t make the start any easier. CI/CD pipelines, version control with Git, collaboration through pull requests: there is a lot of overhead before a single test case runs.
This is where a key task of GenAI at the automation level comes in. The model writes the glue code and boilerplate, the scaffolding that turns an idea into a script that runs locally. Matthias describes it as a sparring partner who is always available, whom you never have to call, and who still gives you feedback right away.
Check Test Cases Before You Automate Them
A well-specified test case can be reviewed by a language model before anyone automates it. The guiding question: can this test case be automated at all, or are test data or environments missing?
AI also works as a consistency check for the test data itself. In practice, specifications often describe only the happy path, the version where the sun is shining and the birds are singing. Edge cases and negative tests are frequently missing.
That is where the real value lies, as long as you ask specific questions. Pasting in a user story and asking for feedback gets you generic answers. Ask instead: are all edge cases covered? Can the test data be improved? Which acceptance criteria are still open? For a whole feature with several acceptance criteria, that precision pays off.
Confidential Data Doesn’t Belong in a Public Model
In the financial sector, copying user data into a public AI is not an option. If you work in a regulated environment, you need a setup where the prompts never leave the company.
At Raiffeisenbank International, the language model runs in-house. The model itself is licensed, but every prompt stays inside the company. That makes it possible to work with real data, at least to a degree, without it getting out.
For coding, there are code completion tools that can be hosted internally. Confidentiality is then no longer the question, only cost is. Matthias expects the investment to pay for itself.
Code That Compiles Isn’t Necessarily Correct
Code that compiles is no proof that the code works. It can run and still not do what you meant it to do. If you aren’t a confident programmer, that gap is easy to miss, and a clean build starts to look like a finished result.
The models have become much better. The early copilots were poor. Matthias even suspects he was slower with them than without, because he kept having to dismiss or rework wrong suggestions.
Today the tools sometimes propose solutions that even an experienced developer didn’t know. That turns AI into a learning tool on the side, because it always explains the code it suggests. A new library, an unfamiliar approach: you read the reasoning and pick something up.
In practice, you’ll still refactor most of the time. Generated code is a starting point, not a finished product.
Why Generating Unit Tests with AI Is the Wrong Approach
Feeding functions into a language model and letting it generate 100 percent code coverage is about the worst thing you can do with it. The model will reach the coverage. But if there is a bug in the code, the generated test locks that bug in instead of finding it.
The problem is the direction. A test derived from existing code doesn’t check whether the code is right. It only cements what is already there, faulty or not.
Generating tests from the user story is somewhat better, but an error in the story still flows straight into the tests. Matthias prefers to go the other way round: write the tests himself and let the AI generate the business code.
That leads to a vision that takes the idea behind test-driven development one step further:
“My idea would be that I write my tests in Playwright and the application builds itself behind the scenes. I can only change my application through the tests.”
(Matthias Zax)
That would lift testing to a new level, because far more would be tested than today and much of the application would be generated. Look and feel, responsiveness and the specific behavior of an application make this hard to picture. But it isn’t impossible.
Code Completion Now Knows Your Context
Modern code completion tools base their suggestions on your own repository, not on code from somewhere on the web. The AI knows your page objects, your keywords, your helper classes.
That changes how you work. The suggestions come from your own project, so you don’t have to rewrite functions to fit. Matthias uses a chat interface less and the completion tools in the IDE more, privately for example Codium for open source work. He sees it as the next stage of the code completion people already knew from tools like IntelliJ.
It also removes an anti-pattern from repositories: excessive commenting. With the older copilots, you had to write long comments to get the right code out of them. The good coding practices of the past still hold. Code should be written so that it needs few comments, with comments only where they are really needed. Otherwise it becomes unreadable.
Generated Documentation Nobody Reads Is an Anti-Pattern
Generating documentation that nobody reads doesn’t solve a problem, it creates a new one. The typical case: someone pastes ten user stories into a model and has it write a fifteen-page test strategy, hallucinated gaps included.
What you get is a soup of generated text that nobody reads but that is now searchable. When companies then put their wikis and repositories into a language model’s context to search them, the AI ends up searching its own unread output. A closed loop with nothing inside.
There are sensible uses. README files are easy to generate, because the AI can work out from the Git repository what belongs in them. You then refactor the draft. Matthias does this himself.
Regulated industries are no excuse for a 200-page test manual either. The regulator doesn’t require 200 pages. It requires you to document how you test. Compact templates with the key information are enough: the architecture as a diagram, clearly marked what is tested, what isn’t and what lies outside. What nobody reads helps nobody.
Getting a Better Handle on Technical Debt
For Matthias, the biggest promise of AI lies in dealing with technical debt. Applications are meant to last, but often they can no longer be developed further because nobody knows how they were built, and upgrades become impossible.
Legacy often starts at a specific moment. Code becomes legacy as soon as the developers who wrote it leave the company. From then on, nobody can work on it.
Static code analysis has helped make problems visible in recent years. Scheduling dedicated refactoring sprints in response is better than nothing, but it isn’t a healthy state. The technical debt keeps growing anyway.
The hope is that AI makes engineering better rather than just faster. The wrong direction would be to shrink teams because AI tools supposedly make them faster. The productive direction: less cloned code, no classes with 2,000 lines, decisions documented where they matter, architecture that gets questioned, and a modern structure for modern applications, whether microservices or a cleanly modular monolith.
Understanding Regular Expressions without Memorizing Them
One everyday use deserves a mention of its own: regular expressions. You paste in a regex and ask what it does, or you describe what you need and let the AI write it.
It’s these small things that take the friction out of the day. You no longer have to worry about syntax you rarely need and never quite remember. Translating between code and plain language, in both directions, is the part that makes daily work noticeably easier.
Frequently Asked Questions
Is generative AI also worthwhile for testers who have primarily relied on manual testing up to now?
Yes, that’s where the greatest leverage lies. Test automation is software development, and making the leap to it is difficult: CI/CD pipelines, Git, and pull requests create a lot of overhead before a single test case runs. The language model handles glue code and boilerplate (the scaffolding for a runnable script) and acts like a constantly available sparring partner with immediate feedback.
How do you ask a language model about a user story so that the answers are useful?
Be specific rather than general. Simply copying a user story into the model and asking for feedback yields generic answers. Targeted questions are helpful: Are all edge cases covered? Can the test data be improved? Which acceptance criteria are still outstanding? Specifications often list only the “good cases,” while negative testing and edge cases are missing. For features with multiple acceptance criteria, this level of precision pays off.
Can you trust the risks an AI identifies for a test case?
No. If you ask for risks, you’ll get risks, but that doesn’t mean they’re actual risks. A language model provides output; validation remains a human task. AI is useful as a preliminary check, for example, to determine whether a test case can be automated at all or whether test data and environments are missing.
How can generative AI be used when confidentiality is involved?
User data does not belong in a public model. In regulated environments, a solution is needed where the prompts do not leave the company. At Raiffeisenbank International, the language model is purchased but hosted internally, so that real data can be used, at least to some extent. Code completion tools can also be hosted internally. That leaves only the question of cost.
Does high code coverage achieved through automatically generated unit tests improve quality?
No, this is one of the worst use cases imaginable. A test derived from existing code does not verify whether the code is correct. If there’s a bug in the software, the generated test actually covers up that very bug instead of finding it. Tests derived from the user story are somewhat better, but an error in the story carries directly over into the tests.
Why is AI-generated code particularly risky for inexperienced developers?
Compilable code is no proof of functioning code. It can run and still not do what was intended. Those who aren’t confident in their programming skills overlook this gap and mistake a successful compilation for a finished product. In practice, refactoring is usually still necessary: The generated code is a starting point, not a final product.
Should you have an AI write documentation?
Only where someone will actually read it. Dumping ten user stories into a model and generating a fifteen-page test strategy from them, hallucinated gaps included, produces a layer of unverified text that later even feeds your own AI search. Readme files work well because you can deduce from the repository what should go in them. You then refactor the draft.
Can generative AI help reduce technical debt?
That’s the biggest hope, but only if it improves the quality of engineering work rather than just making it faster. It would be a mistake to downsize teams just because they’re faster with AI tools. What’s productive is less duplicate code, no classes with 2,000 lines, documented decisions where they matter, and a critically evaluated architecture. Legacy code often emerges the moment the original developers leave the company.


