Skip to main content

Search...

LLM Testing: A Second Model Acts as the Judge

LLM testing can be automated if you accept non-determinism: rubrics and acceptance ranges replace yes/no asserts, and a calibrated LLM judges outputs.

• • Updated: • 11 min read
Cover of the expert talk on 'LLM Testing: A Second Model Acts as the Judge' with Anupam Krishnamurthy and Richard Seidl.

Automated LLM testing makes non-deterministic AI systems systematically testable. Because language models do not return the same output every time, acceptance ranges and scoring rubrics take the place of classic pass/fail assertions. In the LLM-as-a-judge approach, a second LLM rates the outputs of the system under test, calibrated against the judgments of human experts.

Key Takeaways

  • Non-determinism in LLMs is a design property, not a bug: if you treat it like flakiness and try to rule it out, your tests miss the system you are testing.
  • In RAG systems, most defects sit in retrieval, not in generation: an LLM that is fed the right data reliably produces useful answers.
  • According to studies, LLM-as-a-judge reaches 80 to 85 percent agreement with human ratings, about the same level two humans reach with each other.
  • Quality gates for LLM testing need acceptance thresholds instead of binary verdicts: a threshold of around 80 to 85 percent replaces the classic red/green logic.

LLM Testing Means Accepting Non-Determinism as a Feature

LLM testing starts where classic test automation stops working. If you learned test automation the traditional way, you rely on determinism and consistency: a test runs, the result is reproducible, and green stays green. With language models, that foundation is gone. An LLM does not necessarily answer the same question the same way twice.

That is why traditional software testing fails for LLM-based applications. The usual reflex is to rule such behavior out as a race condition, as flakiness or as a bug in the test setup. With generative systems, that reflex is wrong. Non-determinism is part of the system’s design, not one of its defects.

Anupam Krishnamurthy describes the break like this: what you expect from deterministic, conventional software is completely different from what you expect from a chatbot that seems to have a life of its own. That change of perspective is where every sensible test strategy for AI components begins.

Why AI Is Becoming the Test Object, Not Just the Test Tool

There are plenty of articles, opinions, tools and solutions about AI in testing. Far fewer people look at it the other way around: AI as the thing under test.

Applications keep gaining AI components. Jira and other widely used tools now come with AI built in. That shifts the question from “How do I use AI for testing?” to “How do I test software that contains AI?”

Manual review by a domain expert does not cover this. You cannot depend forever on one person rating every single answer. Once AI components are normal in products, you need automated ways to test them.

Trust Comes from Guardrails, Not from Full Understanding

Traditionally, you build trust in a system by understanding it: you know what the code does. With neural networks, that understanding is never complete, and no amount of testing will change that.

The way out is guardrails, tests and evaluations. Even if you don’t understand your system one hundred percent, a defined frame gives you confidence: as long as the behavior stays within set limits, you can use it.

This way of thinking is new only to software developers and testers. Machine learning practitioners have worked like this for a long time. The testing discipline now has to develop its own methods for it.

How to Test AI and LLM Applications: Break the System Down by Evaluation Difficulty

The first concrete step is to split the system into test tasks of different complexity. At heart, this is the same move as in classic testing with its levels of abstraction, just with different goals.

The right approach depends heavily on the use case. Image generation calls for different checks than text generation. AI is a tool, and your approach has to fit your specific solution.

For text-based conversational models, the evaluation can be arranged in tiers:

TierWhat you testHow automatable
DeterministicFactual answer in a fixed format, e.g. JSON, sorted alphabeticallyClassic, with pytest or your usual framework, via assertions
Fact-based, free textCorrectness of free-text answers with an objective coreMore effort, but verifiable
SubjectiveTone, helpfulness, precision of the answerHardest, needs calibrated evaluation
AdversarialResistance to prompt manipulationA test class of its own, not optional

Deterministic checks are quick to write and pay off immediately. If your chatbot should always return facts in the same JSON format in a fixed order, a normal automated test covers it. You want lots of tests like these.

Why Adversarial Testing Is Becoming Mandatory for Language Models

In generative systems, there is no clear line between an instruction in natural language and the behavior of the system. A user can trigger behavior through free text that was never intended.

That creates security risks that did not exist in this form before. Adversarial testing checks specifically whether manipulative inputs can push the model out of its intended scope.

If you work systematically, you cannot skip this test. It belongs next to the fact-based and subjective tests, not after them.

How LLM-as-a-Judge Works: One Model Rates Another

For free text and subjective criteria, one approach is catching on: using another LLM as the evaluator. A second model decides whether the answer of the system under test falls within an acceptable range.

That sounds easier than it is. The obvious objection: don’t two language models share the same blind spots? How do you make sure the LLM judge rates what a human would rate?

In practice, you lower that risk by using different models for the judge and the system, because they then share fewer blind spots. You can also pick a stronger model for the judge than for the system under test. This scales better and, with thousands of users, costs less to run than manual evaluation.

More important than the model choice is alignment with domain expertise. You need structured rubrics: clearly defined criteria with precise conditions and with examples. Examples noticeably improve an LLM’s judgment, the same pattern you know from everyday chatbot use.

RAG as a Test Object: Why the Responsibility Stays with You

If you use a language model straight from a provider, you can hand part of the quality assurance to that provider. With a RAG system, you can’t.

Retrieval-augmented generation uses the language model as just one component. You build everything around it yourself, so responsibility for quality stays with you, and so does the obligation to test thoroughly.

RAG is also widespread because it is low-hanging fruit: you extend an LLM’s knowledge base without retraining it. That makes it a good example of a realistic test object.

How to Isolate Errors between Retrieval and Generation

RAG splits cleanly into two parts, retrieval and generation. That split shows where an error sits.

Retrieval pulls relevant data from a vector database. This part usually involves no LLM at all, but classic machine learning methods such as cosine similarity, similarity search and semantic similarity. Generation takes that data and produces the answer from it.

“In my experience, most problems are in retrieval. If you feed an LLM the right data, you get something sensible out of it.”

(Anupam Krishnamurthy)

So it pays to build targeted tests for retrieval and, separately, for generation. That way you isolate the cause instead of rating the whole system as a black box. The principle is the same as in classic testing: test in small pieces and separate the components from their interaction.

From a Green Pipeline to an Acceptable Range

Deterministic testing knows only red and green, go or no-go. AI systems need a less rigid basis.

Instead of a binary pass or fail, criteria move within a range of values, say between 50 and 100 percent acceptance. As a quality gate, you set a threshold: if the rating is above 80 or 85 percent, the result counts as acceptable. For subjective criteria, there is no strictly right or wrong answer anyway.

You calibrate these thresholds with human experts. Studies on agreement between humans and LLM judges report agreement rates of around 80 to 85 percent. Two humans often agree no more than that. An LLM judge does not have to be more perfect than a second human opinion; it has to meet the human standard.

Why Public Model Benchmarks Give You a False Sense of Security

Public model benchmarks create an artificial feeling of safety. They claim a model is as good as a PhD and hand you a number that says little about your context.

The problem is the method. Models are trained on the standard datasets, much like a student who gets the exam questions in advance and learns them by heart. The exam then measures memorization, not fitness for your use case.

In production, this comes back to bite you when your specific context asks for something different than the benchmark does. To be fair, no vendor can anticipate your context, because it is too specific to your use case. That is exactly why you need to set your own benchmarks.

Your Existing Skills Still Count

Moving into AI testing does not devalue a tester’s experience. The analytical moves still apply: break the system down, isolate components, test in a context-specific way.

The difference is how you deal with non-determinism. That takes creative and critical thinking, but much of what you learned in years of working in quality carries over.

The real draw is somewhere else. Testing AI makes you question what you take for granted: How do we think? How do we arrive at answers? What does correctness even mean? We learn these basics intuitively and normally never question them. When you test generative systems, that is exactly what you have to do.

Frequently Asked Questions

Are different answers to the same question a flaw in the test design?

No. With language models, non-determinism is part of the system’s design, not a flaw. The reflex from traditional test automation, ruling out fluctuating behavior as a race condition or flakiness, leads to tests that miss the mark here. Instead of looking for reproducible consistency, you check whether a response stays within a defined acceptance framework.

Why isn’t manual review by a subject matter expert sufficient for AI functions?

Because it doesn’t scale. You can’t rely indefinitely on a human to evaluate every single response. Since AI components are now embedded in widely used products like Jira, the question shifts from the benefits of AI in testing to testing software that itself contains AI. This requires automated testing methods.

Which parts of a chatbot can still be tested using traditional assertions?

Anything that returns in a fixed format. You can test a factual response in defined JSON with a fixed order using pytest or your usual framework via assertions. Such deterministic tests are quick to build and deliver immediate value, so you should have plenty of them. Free text, subjective criteria such as tone or helpfulness, and adversarial testing require different methods.

How reliable is a language model in evaluating the output of another?

Studies report agreement rates of around 80 to 85 percent between LLM judges and human evaluations. Agreement between two humans is often no higher than that. A judge therefore does not have to be perfect, but rather meet the human standard. You can reduce the risk of shared blind spots by running the judge and the system under test on different models.

What security risks arise from free-text inputs in AI applications?

With generative systems, there is no clear boundary between a natural-language instruction and the system’s behavior. A user can trigger behavior via free-text input that was never intended. Adversarial testing specifically checks whether the model can be pushed beyond its boundaries by manipulative inputs. This test class is not optional.

Where do most errors occur in a RAG system?

Most often in retrieval. If a language model is fed the right data, it will generally produce useful results. That’s why it’s worth testing both parts separately: Retrieval uses classic methods like cosine similarity and semantic similarity, while generation uses those results to generate the response. This way, you isolate the cause instead of evaluating the whole system as a black box.

How do you define a quality gate when test results aren’t binary?

By setting an acceptance threshold instead of a red-green classification. The criteria fall within a range, for example between 50 and 100 percent acceptance. If the rating is above 80 or 85 percent, the result is considered acceptable. You calibrate this threshold with human experts, since there is no strictly correct or incorrect answer when it comes to subjective criteria anyway.

Are public model benchmarks useful for selecting a model?

Only to a limited extent. They convey a false sense of security because models are trained specifically on these standards, much like a student who is given the exam questions in advance. What’s being measured is rote memorization, not the model’s suitability for your use case. No provider can anticipate your specific context, so you’ll need to set your own benchmarks.

Share this page