Sherlock is an AI-based test assistant that helps staff in business departments write test cases that comply with ISO 29119. It runs on MUCGPT, the City of Munich’s own GPT platform. Domain experts without testing knowledge describe their application, and Sherlock generates structured test cases with preconditions, test steps and expected results that can be imported straight into tools such as TestLink or Jira X-Ray.
Key Takeaways
- Staff in business departments have deep domain knowledge but little testing knowledge. Sherlock starts right there and lets them produce ISO 29119-compliant test cases without prior testing training.
- Prompt engineering is the deciding factor: give Sherlock the role of “test analyst” and you get boundary value analysis and equivalence partitioning; leave it out and you end up with a math answer.
- Sherlock exports finished test cases as XML for TestLink or CSV for X-Ray, so business users can import the results directly instead of retyping them.
- Watson, Sherlock’s counterpart, automatically generates a first test report from the test data entered, based on an existing City of Munich template.
- MUCGPT has no internal knowledge about the City of Munich, so users have to load business specifications or system documents into the chat by hand until a RAG connection is in place.
Why Business Departments Struggle to Write Test Cases
AI test case generation is interesting for one group in particular: business departments, which bring deep domain knowledge but very little testing knowledge. That gap is exactly what makes writing test cases hard for them. They know their application inside out, yet face a blank page the moment it has to become a set of structured test cases.
On the technical side, things look different. If you know the software development process, you also know how a test case is put together. But the closer you get to the business departments, the further away the craft of testing is.
The obvious answer, training everyone involved, often fails in practice. Training costs time and money, and business staff have their day jobs to do. Testing happens on the side. You can’t roll out a Foundation Level course across an organization that runs a hundred projects a year.
AI Test Case Generation as a Starting Aid, Not a Replacement
The City of Munich tackles this gap with Sherlock, a test assistant built on MUCGPT, the city’s own GPT platform. Sherlock creates test cases in line with the ISO 29119 standard and gives business departments a way into testing.
The process is deliberately conversational. Ask Sherlock who it is, and you get an answer that includes a sample test case. That shows business users what a standard-compliant test case actually looks like, and from there they start iterating on their own: ask for a test case for a specific function, refine it, repeat.
Sherlock brings the format with it. Precondition, test step, expected result and postcondition are pre-structured according to the standard. The name says it all: Sherlock is the detective with the magnifying glass who investigates and writes test cases.
The division of roles still matters. The AI is the assistant, and the human stays the decision-maker. Whether a generated test case is adopted or refined is up to the business department.
Output Is Only as Good as the Input
A test assistant without context produces generic results. MUCGPT doesn’t know anything about the City of Munich yet, so for a specific business process you first have to give it the business or system specification. Within that chat, it can then create fitting test cases.
Without that context, the model makes things up. A question about the organizational structure of the City of Munich gets an answer, but not necessarily the right one. Anyone using Sherlock should factor that in and check the results against their own expertise.
A connection to retrieval-augmented generation is planned. Business departments could then feed in their own data, business concepts, system specifications and user stories and connect internal intranet pages, instead of copying specifications into the chat by hand.
AI Isn’t Deterministic, and That Can Work in Your Favor
Enter the same prompt five times and you won’t get the same test case five times. That’s the nature of the beast, because an AI model varies its output.
Rather than fighting it, you can put the variation to work. Ask for several test cases at once, around five, and review them side by side. You’ll quickly see which variants fit and can fine-tune from there.
Prompt Engineering Is the Real Core Skill
Prompting comes before the tool. In the City of Munich’s training sessions, working with Sherlock doesn’t start with the assistant itself but with prompt engineering.
The role you assign decides the result. Give the model the role of test analyst and it knows test design techniques such as boundary value analysis and equivalence partitioning. Without that role, a question about a boundary value ends up in mathematics instead of testing.
That’s what makes Sherlock tangible: a system specialist for one particular job. Besides the role, context, background information and formatting instructions matter too. Sherlock already comes with formatting that follows the ISO standard.
Getting Test Cases Into the Test Management Tool Is Still a Hurdle
The generated test cases have to go into the test management tool, and that’s where new friction appears for business users. The City of Munich has three standard tools: TestLink as an open source product, X-Ray from the Jira world and SAP Solution Manager.
Sherlock exports an importable XML file for TestLink and a CSV file for X-Ray. The current limit is ten test cases per import. SAP Solution Manager isn’t connected yet.
For business users, the tool itself is already a hurdle. Where are my test cases? In TestLink they sit under “test specification”, because that’s what the tool calls them. The automatic import takes this step off their hands, since business users no longer have to write up or enter the test cases themselves.
The limit of ten test cases has a technical reason. Depending on the model, there’s a maximum token output length beyond which the output gets cut off.
Watson Writes the Test Report
Where Sherlock investigates, Watson documents. The second assistant is still in development and automatically creates a test report from the key figures you enter.
Watson works with a system prompt into which you enter the relevant data: number of test cases executed, passed and failed, bugs found with their IDs, and the project number. Based on an existing test report template, it then produces a first draft.
The draft isn’t final. But the core data is already in it, which noticeably cuts the effort for the report.
Requirements Grow Out of Real Use
The most useful features come from business department feedback, not from the drawing board. The first version produced a single standard-compliant test case, but the departments asked for ten at once. That’s how the multi-case output came about.
Another pain point is writing defect reports. Which ticket type, which fields, how to fill them in? The plan is for users to describe the defect, for the assistant to map the information and for it to import the ticket into Mantis or X-Ray.
Hard usage figures are difficult to collect at the moment. Feedback arrives sporadically, for example through internal open space formats. In the future, users will be able to subscribe to the assistants, so the number of subscribers will at least show how widely they’re used.
More Ownership Through Roles and Permissions
Until now, any change to the assistants had to be made by the developers at the AI competence center. MUCGPT 2.0 changes that with a roles and permissions system.
From then on, the business owner of an assistant owns it. They can change the system prompt, adjust sample answers and respond to feedback without going back to the developers every time. That takes load off the AI competence center and shortens the path from idea to feature.
“In the end, the AI is the assistant, and the human is still the decision-maker.”
(Mark Menzel)
Innovation Depends on People and an Early Start
Public administration is quickly written off as old-fashioned, but that picture doesn’t hold here. The drivers are specific people, from the AI competence center to work-study students who build topics like Watson in their theses.
What really made it possible was an early decision. The City of Munich set up MUCGPT back in April 2023, jumping on the trend instead of waiting it out. Acceptance is broad: the Social Services Department, the Building Department and the Department of Education and Sports are leading the way, not just IT.
The yardstick remains the benefit. Innovation isn’t justified by itself, but by the value it adds to the work and to the city.
Frequently Asked Questions
Why Isn’t It Enough to Simply Train Departmental Staff in Testing?
Company-wide training programs fall short due to day-to-day operations. A Foundation Level program cannot be rolled out across an organization that implements a hundred projects a year, because training takes time and money, and testing is something business users do on the side. The gap remains: They know their application in detail, but are at a loss when it comes to turning that knowledge into structured test cases.
Can business users without testing training create standards-compliant test cases using AI?
Yes, if the assistant provides the right format. The Sherlock test assistant used by the City of Munich structures test cases according to ISO 29119, including preconditions, test steps, expected results, and postconditions. Business users describe their application and refine the test cases through dialogue. The division of roles remains clear: The AI acts as an assistant; the decision to accept or refine the results rests with the business department.
What happens if a language model lacks the technical specifications for the system being tested?
Then the model makes things up. MUCGPT has no internal knowledge about the City of Munich: A question about the organizational structure may yield an answer, but not necessarily the correct one. For a specific business process, the business or system specification must therefore be uploaded to the chat. An integration via Retrieval-Augmented Generation is planned to eliminate the need for this manual copying.
Why does the same request to an AI test assistant yield a different test case every time?
AI models do not operate deterministically; their outputs vary. The same prompt five times yields five different test cases. Instead of fighting this, you can take advantage of the variation: Request several test cases at once (say, five) and review them together. This way, you can quickly identify which variants work and fine-tune from there.
Why does a question about boundary values sent to the language model end up in mathematics instead of testing?
Because the role is missing. If the model is assigned the role of test analyst, it understands boundary value analysis and equivalence partitioning and applies them to the described application. Without this assignment, it interprets the term mathematically. That’s why, in training sessions, working with a test assistant doesn’t start with the tool itself, but with prompt engineering: role, context, background information, and formatting specifications.
How do AI-generated test cases get into a test management tool?
Via export formats that the target tool can import. Sherlock generates an XML file for TestLink and a CSV file for X-Ray, so business users don’t have to retype anything. The limit is a maximum of ten test cases per import because models have a maximum token output length and truncate beyond that. SAP Solution Manager is not yet integrated.
Can the test report also be generated using an AI assistant?
Partially. The Watson assistant works with a system prompt into which key metrics are entered: the number of test cases executed, those that passed and failed, bugs found with their numbers, and the project number. Based on an existing test report template, this generates a first draft. It’s not final, but the core data is already included, and the effort required is noticeably reduced.
Where do the most useful features for an AI assistant in a testing environment come from?
From actual use, not from the drawing board. The first version of Sherlock generated exactly one standards-compliant test case, but the business departments requested ten: that’s how the multiple-output feature came about. The next pain point turned out to be creating defect reports, which is why the plan is to let users describe the defect and have the assistant submit the ticket to Mantis or X-Ray.


