Skip to main content

Search...

Security Testing for AI: Checking Outputs Instead of Just Inputs

AI security testing covers the whole life cycle, from training data to live queries. Why output filters often beat input filters, and what OWASP offers.

• • Updated: • 11 min read
Cover of the expert talk on 'Security Testing for AI: Checking Outputs Instead of Just Inputs' with Jan Jürjens and Richard Seidl.

AI security testing spans the whole life cycle of an AI model: from training, where manipulated input data can skew the model, to day-to-day use, where attackers pull protected data out through carefully crafted queries. The OWASP Top 10 for AI and the OWASP AI Security & Privacy Guide give teams a framework to start from.

Key Takeaways

  • AI models can be talked into revealing protected data through indirect queries, even when direct queries are blocked, because the model doesn’t recognize a question about a job title as a workaround.
  • Security risks in AI-based systems run through the entire life cycle: from manipulated training data and theft of the model to misused queries in production.
  • Output filters often work better than input filters, because they check directly whether the generated answer contains sensitive content such as salaries or personal data.
  • The OWASP Top 10 for AI-based software is a structured starting point for security requirements, but putting it into practice takes expertise, either in-house or from outside specialists.

What Makes AI Security Testing Different

With AI-based software, you don’t know in advance exactly how it will behave. That is the whole idea. A machine learning model learns on its own instead of being explicitly programmed, and at best it solves the task better than code ever could. This is where AI security testing gets harder than testing conventional software.

That property creates the first security problem. If behavior isn’t fixed beforehand, it is hard to establish whether the system holds up against attacks. Security is one part of what Jan Jürjens calls trustworthy AI.

Most companies don’t build AI themselves. They buy a model or plug it into their own infrastructure as a component. So the question becomes: how do you make an IT infrastructure that uses AI secure against attacks and trustworthy overall?

Attack Points Span the Entire Life Cycle

Security risks in AI run from training all the way to use. There isn’t one single weak spot, but a chain of phases, each with its own requirements.

During training, an attacker may slip in data that influences the model in ways nobody intended. If you train your own model, you have to secure this phase. Once manipulated training data is in, it is very hard to identify cleanly later.

Once the model exists, classic security requirements apply, just to a new kind of asset. Nobody should be able to copy or steal the model, and nobody should be able to tamper with it. An AI model is often a business asset and needs the same protection as other critical data.

Use is the third area. Here the concern is that queries to the model don’t undermine the security rules and that no protected data leaks out.

Why a Chatbot Can Reveal Personal Data

Filtering the direct question isn’t enough, because the same information can be requested by a detour. That is the core problem in securing queries.

Jan Jürjens explains the pattern with a salary example. If a model was trained on employee data, a direct question about a specific person’s salary should be blocked. Ask instead about the salary of a particular role, say the compliance officer, and the system may happily answer. Anyone who knows who holds that role now knows that person’s salary.

The lack of explainability makes this worse. You get an answer, but no explanation of how the model arrived at it.

“AI doesn’t usually work like a search engine. It aggregates the information from its knowledge store, so that in the end nobody can say what the answer is made of.”

(Jan Jürjens)

Because the answer comes from aggregated information, it is almost impossible to prove afterwards whether something that shouldn’t be there went into it. That is exactly why protection has to start during training.

AI Security Testing Works like Penetration Testing

The most effective defense against manipulated queries is intensive testing from an attacker’s point of view. You think about how someone would go about it and send those queries to the system yourself.

This is the same approach as penetration testing for conventional software. You take on the attacker’s role and try to get a reaction or a piece of information out of the system that it shouldn’t give away. If you find a gap like that, you then block that kind of query.

You can never be sure you have found everything. It is a race, but that holds just as true for security testing of conventional software.

Output Filters Often Work Better Than Input Filters

Checking the output often beats trying to catch every nested input. The reason is practical: right before the answer goes out, you can check directly whether, say, a salary appears in it.

Both places make sense. At the input, you block problematic requests. At the output, you check the result before it leaves the system. Indirectly phrased queries are hard to catch completely, while an output check works no matter how clever the question was.

Limiting requests makes sense anyway. Even with models trained on public data, you don’t want someone to extract the entire model through mass queries and rebuild it. Rate limiting should therefore be standard.

Reputation is another reason for filters. Nobody should be able to get an application to produce offensive output, because that reflects badly on the model vendor and on the company running it.

OWASP Provides the Checklists for AI, Too

There are already established resources for AI security, backed by the OWASP community. The organization started out as the Open Web Application Security Project and renamed itself the Open Worldwide Application Security Project, because its work has long gone beyond web applications.

For AI, there are specific resources, including an OWASP guide on security and privacy and an OWASP Top 10 for AI. The Top 10 format is familiar from web security and has been carried over to AI-based software.

These documents are easy to follow and a good place to start. The real hurdle isn’t reading them. It is putting them into practice.

Easy to Read, Hard to Implement

As a developer, you can follow the OWASP material, but assessing your infrastructure takes someone who knows how to judge it. That is the honest answer if your assignment is to add a chatbot quickly.

To decide whether your infrastructure really is secure, you need either trained people of your own or outside support. Here, too, AI is nothing fundamentally new. It is the same situation as in classic security.

If You Run AI, You Own Everything around It

As soon as you use AI professionally, you count as a deployer under the EU AI Act, and that comes with obligations. The regulation distinguishes between those who develop AI, those who deploy it and those who use it.

In most cases, you build AI into your architecture as a black box. You can’t look inside the model or its internal processes. But you are responsible for everything that happens around it.

That covers several questions at once:

  • What data goes in, and is it protected?
  • How does data come out, and is it protected at the interfaces?
  • What happens with the results, and is that even allowed?

The EU AI Act also limits what AI may be used for. Following those rules is part of the deployer’s duty. As far as you can, also make sure that your vendor complies with the regulations.

The Biggest Weak Spot Is Still the Person at the Keyboard

People often use AI carelessly, and that is an underestimated risk. Documents get uploaded and analyzed without anyone checking what data leaves the building.

Many users trust a chatbot’s answer blindly. That is dangerous whenever something depends on the answer being right, at home or at work. AI isn’t one hundred percent reliable, and hallucinations appear as soon as you dig a little deeper.

Jan Jürjens gives an everyday example. Asked about a trampoline park in Koblenz, a chatbot confidently named one that doesn’t exist there. The system had probably found the name for another city and simply moved it to Koblenz. For a museum, it even gave an address in Aachen, a street that doesn’t exist in Koblenz at all.

Ask whether the answer is correct, and the system sticks to it at first. Only after pushing several times does it admit the mistake.

AI Gives You Hypotheses, Not Facts

Use AI to generate hypotheses, and do the checking yourself whenever the answer matters. That split is the practical consequence of its limited reliability.

Some tasks suit it well, such as quickly summarizing a long document. Even there, a statement may show up that wasn’t in the document, but you can check that against the original. You should always have that kind of fallback.

Be careful with answers you can’t verify yourself. With security topics in particular, a wrong answer delivered with confidence carries a lot of weight. Making users aware of this matters as much as the technical filters.

Frequently Asked Questions

Why is it harder to assess whether AI-based software is secure?

Because its behavior isn’t predetermined. A machine-learning model learns on its own, rather than being hard-coded in the traditional way. That’s precisely where its value lies, but it’s also the primary security challenge: what can’t be predicted is also difficult to test for secure behavior against attacks. Added to this is the lack of explainability, since the model aggregates information and the origin of a response can no longer be traced.

Do I need to worry about AI security if I’m not training the model myself?

Yes. Most companies purchase a model or integrate it as a component, but they remain responsible for everything related to the model. While the training phase is no longer their own task, they are still responsible for protecting the model, as a corporate asset, against copying and manipulation, as well as for securing queries during operation.

Is it enough to block direct questions about sensitive data?

No, because the same information can still be accessed indirectly. If a model was trained on employee data, it might block a question about the salary of a specific person. However, it may readily answer a question about the compliance officer’s salary. Anyone who knows who holds that role also knows the salary. The model does not recognize the role title as an attempt to circumvent the system.

Can penetration testing be applied to AI systems?

Yes, and it’s the most effective approach against manipulated queries. The process is the same as with traditional software: You take on the role of the attacker, think about how someone might proceed, and make these queries yourself. If you find a vulnerability, you block that type of query. Completeness is never achievable here, but the same applies to traditional security testing as well.

Where is it better to intercept unwanted content: at the input or at the output?

Both points make sense, but the output is often more effective. Nested or indirectly phrased requests are nearly impossible to block completely. Before output, however, you can directly check whether sensitive content (such as salary information or personal data) appears in the response. This check works regardless of how cleverly the question was phrased.

Why should the number of requests to an AI model be limited?

To prevent anyone from draining the model through mass queries and then replicating it. This also applies to models based on public data, because the model itself represents the company’s value. A volume limit is therefore standard practice. A second reason for filters is reputation: An application should not be allowed to produce offensive output, as this reflects poorly on the manufacturer and operator.

Is reading the OWASP Top 10 for AI enough to secure a system?

No. The OWASP materials, including the AI Security & Privacy Guide and the Top 10 for AI-based software, are easy to understand and a useful starting point. The challenge lies in implementation. To assess whether an infrastructure is actually secure, you need trained in-house staff or external support. In this respect, AI is no different from traditional security.

How should one handle chatbot responses that cannot be verified independently?

With caution. AI provides hypotheses, not truths, and the verification step falls to the user as soon as something depends on the response. Hallucinations become apparent as soon as you dig deeper: One chatbot confidently named a trampoline park in Koblenz that doesn’t exist there, and gave an address in Aachen for a museum. Tasks that allow for verification, such as summaries that can be checked against the original, are well-suited for this.

Share this page