Software quality today rests on three factors that work together: a test process that actually functions, a solid grasp of non-functional software quality characteristics such as performance, usability and security, and a deliberate use of AI tools. Where the process foundation is missing, it catches up with every team sooner or later, and with AI faster than ever.
Key Takeaways
- Sloppy processes and missing groundwork slow down AI rollouts, because gaps in the setup hit back faster with AI than they ever did with test automation.
- Non-functional quality characteristics such as performance, usability and accessibility have gone from nice-to-have to a must, and their weight keeps growing.
- Testers will need a broader technical skill set, but they must not lose their instinct for bugs, because gut feeling uncovers a large share of defects.
- Quality has to go into the product early, not at the end: teams that bring in QA too late pay considerably more, both in money and in rework.
AI Only Speeds Up What Already Works
AI amplifies the weaknesses in a test process instead of fixing them, and that has a direct effect on software quality. Put AI on top of a patchy process and your problems come back faster, not slower. What still matters are the basics: a test process the team really lives, and software quality characteristics that go well beyond functionality.
Matthias Gross knows this mechanism from test automation, and he sees it repeating itself now. With automation it was already true: if the underlying process is garbage, it will blow up sooner or later. AI does exactly the same, only faster. Messy setups and gaps in the approach catch up with the team so quickly that the AI never gets used the way it was intended.
That leads to an uncomfortable truth for many projects. AI is often the thing people put up front so they don’t have to deal with the actual question. The real issues sit in the test strategy and the test process: How do I really test? What am I trying to achieve? Which problem am I solving? Using an AI tool to dodge these questions doesn’t improve anything.
Then again, the current hype may steer attention right back to them. If AI only works on a solid foundation, the groundwork becomes a must once more.
Why Groundwork for Software Quality Matters More Now
A shared understanding and a process the team actually follows are what make AI in testing work in the first place. Without that foundation, every additional layer gets harder, whether it’s automation or AI.
Some of the established testing methods are decades old, and the framework behind the common certifications has been around for roughly 20 years. That maturity is an advantage. These are proven, clearly defined concepts with a common vocabulary.
AI, by contrast, is new and changes constantly. Even the terms are still fuzzy. One person talks about AI agents, the next about prompt chains, and often each means something different. Before a team can use AI sensibly, it has to agree on what these terms actually mean.
This is where the old groundwork pays off. A company that can say “this is our process and this is how we work” has the base that AI needs. Without it, getting started with automation or AI is much harder.
Vibe Coding Produces Software That Still Needs Testing
With AI, software gets built at a pace that overwhelms the usual QA approach. Testers are structurally behind, because they need to understand a topic before they can verify it.
On the other side are people who simply get going with AI. A hackathon produces something that works, without anyone having thought about quality or security beforehand. The problem: whatever comes out of that is hard to rein in later.
“Please, please don’t turn a hackathon into a product. Think it through first. Build in quality assurance up front, build in security up front, because you won’t be able to catch up on it later.”
(Wolfgang Sperling)
Over the coming months and years, a lot of software will be thrown together quickly, and its quality will stay questionable. Testing has to deal with that. Fifteen-page Excel test lists won’t get you far. Those lists still exist and sometimes even work, but they can’t keep up with this new speed.
Non-Functional Software Quality Characteristics Move to the Front
The focus in testing is shifting from pure functionality to non-functional criteria. Functionality is still where it starts, but it’s no longer what makes software good.
Christian Mercier compares then and now. About 30 years ago it was enough for a piece of software to do its job at all. Today the bar is higher. Performance used to be an afterthought; now everyone expects fast responses. For a long time usability only had to be usable; now people expect excellent usability. Accessibility was once a unique selling point; today it’s standard.
The relevant quality models reflect this shift too. Wolfgang points out that they give functionality a single sentence, while several criteria cover the non-functional side.
AI agents also move part of quality assurance into production. Anyone deploying agents has to ask what happens during live operation and how to guard against risks there. The platform vendors provide tools for risk mitigation. Which of them you use, and how you use and configure them, becomes a concrete testing task.
Debugging AI Systems Means Reading Traces
When you test AI systems, you analyze traces rather than classic test results. A trace records the individual steps inside the AI and can easily run to several pages.
Reading them is tiring, but doable. You have to recognize whether there’s a bug at all and what kind it is. A typical case: a function is called three times and never returns a response. Then the question is whether the prompt needs work.
The mechanisms for finding bugs are the same ones testers have trained all along. That’s a good starting point. On top of that comes closer contact with the code and a feel for what happens during prompting and what comes back as a response.
With non-deterministic systems, it’s no longer testing in the classic sense but validation. You work with criteria instead of fixed expected values, because there’s no guarantee the same result can be reproduced.
The Tester’s Skill Set Gets Broader and More Technical
Domain knowledge alone is no longer enough. The role increasingly calls for technical depth, and at the same time the business side and IT need to move toward each other.
Matthias describes two directions. The business side has to get closer to the technology and learn to test more than functionality, including aspects like usability. IT moves closer to the business and helps with test automation, performance testing and compatibility. The goal is a shared product, not code thrown over the fence to the business department with the instruction “now test it.”
One thing must not get lost amid all the technology: the instinct for bugs. This professional intuition drives the search. Matthias estimates that 40 to 50 percent of bugs aren’t found through a clean test case but through a gut feeling someone decides to follow.
That intuition is a kind of unconscious competence that grows over the years. It complements methodical persistence. Together they cover the basics and still find what no test case anticipated.
In the End, Quality Is a Feeling
Quality can be described as a feeling, and every quality characteristic feeds into it. Functional and non-functional criteria are the tools that make this feeling measurable.
Users notice good quality. Their data is safe, the software is fast, it’s pleasant to use, and it does what they want. Each characteristic adds to that sense of satisfaction in its own way.
This gives the tester a new role: the user’s advocate, bringing the user’s perspective into the team. That whole-system view covers many criteria at once, and it’s a skill worth cultivating on purpose.
Part of it is questioning the status quo. If you’ve been maintaining and running test case catalogs for years, check whether they still fit. Is the granularity right? Am I testing the right things? Have the requirements changed? This kind of critical reflection isn’t a side task. If it isn’t part of the job yet, it has to become part of it, because nobody else on the project looks at things from this angle.
Sharing Beats Going It Alone
When things move this fast, sharing knowledge is the most effective way to keep up. No AI replaces the network, the look beyond your own backyard and the short path to other people.
Regional communities show how quickly this can take off. A newly founded testing community in southwest Germany brought together people who hadn’t known each other, and after a short time they were far better connected, with a much shorter path to getting someone’s opinion. Anyone who offers something like this soon notices that people come in numbers and want to talk.
The exchange works because different backgrounds meet there: system integrators, pure testing providers, people who build IT solutions themselves. That mix leads to conversations about everyday challenges, and everyone takes something home.
Looking ahead, one thing counts above all: share concrete solutions, not just concepts at a high level of abstraction. Show what worked for you so others can use it in their daily work. There’s a lot of untapped knowledge everyone could learn from. Have the courage to show your solutions. People will be glad you did.
Frequently Asked Questions
Why do AI implementations in testing often fail for reasons completely unrelated to the technology itself?
Because AI exacerbates existing weaknesses rather than solving them. If the underlying test process is flawed, those flaws will catch up with the team faster than they used to with test automation. Often, the AI project even serves to sidestep the real questions: How am I actually testing, what do I want to achieve, and what problem am I solving?
Do AI tools make traditional testing methods obsolete?
No. The established methods are, in some cases, decades old; the framework behind common certifications has existed for about two decades; and it is precisely this maturity that provides clearly defined concepts and a common language. With AI, even the terms are vague: one person talks about AI agents, another about prompt chains, and they often mean different things.
What are the arguments against turning a hackathon prototype into a product?
Quality assurance and security are nearly impossible to retrofit afterward. If something is created in a hackathon that works without anyone having thought about quality or security beforehand, you won’t be able to address those issues later. Both need to be addressed before implementation, not at the end of it.
Is it enough today for software to work correctly from the standpoint of functional correctness?
No. In the 1990s, it was enough for a piece of software to simply fulfill its function. Nowadays, fast response times are expected, excellent usability is required, and accessibility has gone from being a unique selling point to a standard. Even common quality models devote only a single sentence to functionality, while several criteria focus on the non-functional area.
How does debugging in AI systems differ from traditional test analysis?
Instead of traditional test results, you analyze traces that document the individual steps within the AI and can quickly span several pages. A typical finding: A function is called three times and never returns a response. With non-deterministic systems, the focus is on validation based on criteria, not on fixed target values.
Does the value of intuitive error detection diminish when testers have to work in a more technical manner?
No, it remains the second pillar. According to Matthias Gross’s assessment, 40 to 50 percent of errors are not found through a well-designed test case, but rather through a gut feeling that someone follows up on. This unconscious expertise grows over the years and complements the methodological perseverance that covers the basics.
What are the concrete benefits of collaboration within testing communities?
It shortens the path to gathering opinions and brings together different perspectives: system integrators, dedicated testing providers, and people who implement IT solutions themselves. However, this only becomes effective when concrete solutions are shared rather than concepts at a high level of abstraction. Show what has worked for you so that others can implement it in their day-to-day work.


