Acceptance Test Driven LLM Development (ATDD for LLMs) is an approach in which failing acceptance tests built from real error dialogs also serve as training data for fine-tuning a large language model. New business requirements are specified as dialog-based test cases, the model is retrained, and then it is checked automatically against all previous tests to catch regressions.
Key Takeaways
- Acceptance Test Driven LLM Development applies the idea of failing acceptance tests directly to fine-tuning: a test fails, new training dialogs fix the problem, and the test suite then shows whether the model has learned without introducing regressions.
- LLM output can’t be checked by string comparison, because answers with the same meaning rarely share the same wording. Verification has to switch between strict structural comparison and semantic checking depending on the type of output.
- Real dialogs from pilot operation are the main data source: anonymized conversations show which requests the model can’t yet handle on its own and supply the material for new training and test data.
- Template-based dialogs make combinatorial testing possible: template variables are replaced with many different values, so a single test case turns into a large series of measurements.
What ATDD for LLMs Means
Test driven LLM development, in the form of ATDD for LLMs, takes the principle of acceptance test driven development and applies it to training and validating language models. Instead of training a model once and hoping it fits, the team first writes acceptance tests that the model fails at the start, and then develops specifically against them.
The term comes from David Faragó, who works at Mediform on a phone assistant for medical practices. His point: LLM development as a discipline is still in its early days. The models themselves are powerful, but the processes and tools around them lag behind.
The appeal lies in bringing two worlds together. Machine learning contributes the training set and the test set. Agile software development contributes acceptance tests as a safety net. If you derive failing test cases from real examples, you can measure afterwards whether the model has actually improved.
Why Testing Language Models Is So Hard
Verifying an LLM is hard because it isn’t deterministic and behaves like a black box. The same input can produce different outputs, and nobody can see directly why the model answers the way it does.
Faragó quotes a fitting image: LLMs are technology handed to us by aliens. They are excellent at handling natural language and drawing conclusions, yet classic methods can barely control them.
On top of that, natural language applications can’t be checked with simple string comparisons. “Did I understand correctly, you’d like to book an appointment?” and “So you want to book an appointment?” are different strings with the same meaning. That semantic check is the hard part.
Three Levers for Quality
Mediform tackles these challenges with a solid process, fast validation and tailored verification.
The process is based on CPMAI, Cognilytica’s methodology for AI projects. It combines modern software development with modern machine learning and builds in agility, which older approaches such as CRISP-DM don’t.
Validation runs in short iteration cycles with pilot customers directly involved. Patients use the bot, the team analyzes the anonymized dialogs and works out what needs to change in the next iteration.
Verification relies on a test tool that grew out of EleutherAI’s evaluation harness for language models. The same harness sits behind the well-known LLM leaderboard on Hugging Face. Mediform has extended it heavily and tailored it to its own business requirements.
Task-Oriented Dialog Sits Right in the Middle
Phone assistants for medical practices are a special case for testing, because they fall between two extremes. Faragó calls this class of applications task-oriented dialog.
At one end are tasks with a single correct answer. A string comparison for equality is enough there. At the other end is free creativity, such as a generated poem, where a person reads it over and small differences don’t matter.
A phone bot that books appointments sits in between. It has to complete a specific task but phrases its answers freely. So the check has to tell apart “literally the same” and “means the same” instead of just comparing strings.
How the Toolformer Approach Shapes Verification
Mediform uses an agent-based Toolformer approach in which the model produces two kinds of output. One is natural language text for the patient. The other is function calls.
Function calls trigger prewritten messages, so the model doesn’t have to generate long blocks of text from scratch every time. The model also produces function calls for database queries and similar actions.
The verification tool applies different levels of strictness depending on the case. With free text it can be lenient; with function calls and database queries the output has to be exactly right. That distinction decides whether a test runs as a string comparison or as a semantic check.
The Dialog Is at the Center of the Process
Mediform keeps one dialog format through the entire process. The same structure serves as training data, as test cases and as the input format for the test tool.
The workflow follows the CPMAI cycle. After deployment, the pilot customer collects anonymized dialogs that really took place. The team analyzes them during business and data understanding, measures, for example, how many dialogs reached their goal fully autonomously, and picks out the cases that went badly.
Those faulty dialogs become new acceptance tests. Because they come from real failures, they fail at first: the model can’t handle them yet. In parallel, the team generates many training dialogs in the same format.
The model is then retrained and run against both the new and the old tests. That way you measure at the same time whether the model has learned and whether it has regressed anywhere.
The Six Stages of CPMAI at a Glance
CPMAI structures LLM development in six iterative stages, and you can move back and forth between them. Mediform has merged the first and second stages for its use case.
| Stage | What it means at Mediform |
|---|---|
| Business Understanding | Merged with Data Understanding, because the same dialogs are analyzed |
| Data Understanding | Analyzing the collected dialogs, deriving new business requirements |
| Data Preparation | Writing acceptance tests, building the training set |
| Model Development | Fine-tuning the language model |
| Evaluation | Running the test tool, calculating business-oriented metrics |
| Deployment | Demo, pilot customer or production, to collect new data |
The sixth stage is the real step forward compared with CRISP-DM. Deployment closes the loop, because it delivers new, business-oriented data for the next iteration.
When Prompt Engineering Isn’t Enough: Fine-Tuning
Fine-tuning comes in when adjusting the prompts alone no longer gets you further. Prompt engineering only changes the input. Fine-tuning adjusts the weights of a base model toward the specific task.
Mediform runs two variants: a fine-tuned GPT and a fine-tuned open source model such as Mistral. Prompt engineering stays part of the solution in both cases.
The training data has one peculiarity. The dialogs come from speech-to-text, so they sometimes contain transcription errors. The test framework has to cope with input that isn’t cleanly typed text.
A language example shows how strongly fine-tuning shapes a model. A Mistral model fine-tuned on German dialogs understands French very well, but it sometimes drifts into German when it answers, because the fine-tuning pulls it toward German. A French pilot customer therefore needs French training and test dialogs.
Template-Based Tests Open the Door to Metamorphic Testing
Because acceptance tests and training data are written from templates, metamorphic testing is easy to add. You take a relevant dialog and replace the template variables with many other values.
The result is combinatorial testing for a single aspect. Instead of one case, the team runs many measurements on the same question and sees whether the model stays stable across the variations.
Before deployment there are also stress tests. The team automates the patient side with a language model of its own and runs the freshly fine-tuned model against this patient model. The resulting dialogs are then evaluated.
The acceptance tests remain the real safety net, though. Unit tests still run for the code around the model, but the behavior of the LLM itself is secured mainly by the dialog-based acceptance tests.
“You take precedents or examples, use them to specify the new requirements, and then you have a way to measure, or even a safety net, so you can really measure afterwards whether you’ve improved it.”
(David Faragó)
Why Iterations Don’t Fit Rigid Sprints
At Mediform, iteration cycles don’t follow a fixed rhythm, even though the team works in two-week sprints. A CPMAI cycle doesn’t necessarily line up with a sprint.
Sometimes several iterations fit into two weeks; sometimes a single cycle takes longer. Ideally sprint and cycle overlap, but in practice something regularly gets in the way.
Demand is the driver. When a new pilot customer needs an additional language, the team slots in a quick interim iteration and starts with the existing model before generating new training and test data. That flexibility is part of the approach, not a departure from it.
Frequently Asked Questions
Why isn’t a simple string comparison enough to verify a voice assistant’s responses?
Because semantically identical responses almost never have the exact same wording. “Did I understand correctly that you want to book an appointment?” and “So you want to book an appointment?” are two different strings with the same meaning. Furthermore, an LLM does not operate deterministically: the same input can generate different outputs. Verification must therefore be semantic rather than literal.
What’s the point of writing acceptance tests before a model is fine-tuned?
They make progress measurable. Tests derived from real error dialogs initially fail because the model cannot yet handle these cases. After retraining, the test suite reveals two things: whether the model has learned the new cases and whether it has broken old cases that were already working.
Do all outputs from an LLM-based assistant need to be checked with the same level of rigor?
No. For free-form text intended for the user, the check can be more lenient; for function calls and database queries, the output must be exactly correct. In the agent-based Toolformer approach, the model generates both types of output. This distinction determines whether a test runs as a strict structural comparison or as a semantic check.
Where can you get meaningful test cases for a dialog-based system?
From the pilot operation. The pilot customer collects anonymized, real-world dialogs; the team analyzes them and measures, for example, how many conversations reached their goal completely autonomously. The suboptimal cases become new acceptance tests, while many training dialogs in the same format are generated in parallel. A single dialog format thus serves as both training data and a test case.
What distinguishes CPMAI from older process models such as CRISP-DM?
CPMAI integrates agility and adds a sixth stage: deployment. Only deployment closes the loop, because it provides new, business-oriented data for the next iteration. The six stages are iterative; you can move back and forth between them. Business and data understanding can be combined if the same dialogs are being analyzed anyway.
When is fine-tuning more worthwhile than further prompt engineering?
When adjusting the prompts alone no longer yields results. Prompt engineering only changes the input; fine-tuning adjusts the weights of a base model to suit the specific task. The two are not mutually exclusive; prompt engineering remains part of the solution. Just how much fine-tuning influences the model is illustrated by a model trained on German dialogs that understands French but sometimes answers in German.
How do you check whether a language model remains stable across variations of the same query?
Through template-based test dialogs. You take a relevant dialog and replace the template variables with many different values. A single test case thus becomes a series of measurements on the same aspect, that is, combinatorial testing in the sense of metamorphic testing. In addition, stress tests are run in which a separate language model takes on the user’s role, and the resulting dialogs are evaluated.
Can training and testing cycles for an LLM be tied to fixed sprints?
Only to a limited extent. A complete run does not necessarily align with a two-week sprint: sometimes several iterations fit within it, while at other times a single cycle takes longer. Demand is the driver. If a new pilot customer needs an additional language, a quick interim iteration is inserted, initially starting with the existing model.


