Skip to main content

Search...

Prompt Engineering for Test Cases: Review Remains Mandatory

Prompt engineering for test cases gets better results with few-shot examples and embedded specs. Still, two of eleven generated tests remained wrong.

• • Updated: • 11 min read
Cover of the expert talk on 'Prompt Engineering for Test Cases: Review Remains Mandatory' with David Faragó and Richard Seidl.

Using prompt engineering for test generation means coaxing better test cases out of a large language model step by step, with targeted prompt patterns. Proven techniques include few-shot examples for robustness, embedded library specifications for correctness and the ReAct pattern for stable conversation sequences. Correctness remains the central unsolved weakness.

Key Takeaways

  • Prompt engineering doesn’t replace human review: even after iterative improvement with several prompt patterns, two out of eleven generated test cases were still wrong, because the model can’t reliably guarantee correctness.
  • Few-shot examples in the prompt make results noticeably more robust: they stop the model from producing completely different or absurd output after tiny changes to the prompt.
  • Putting domain knowledge directly into the prompt improves correctness more than general prompt tuning: adding a concrete library specification clearly raised the share of correct test cases.
  • The ReAct pattern (reasoning plus action) was the most robust prompt pattern in David’s experiments: in his handful of trials, no hallucinations occurred and the prompts stayed stable when varied.
  • Testers are better equipped for prompt engineering than many other roles, because they are used to checking correctness, specifying requirements and building representative examples.

Prompt Engineering Is What Makes Language Models Useful for Quality Work

Prompt engineering means carefully designing the instructions you give a language model so that it reliably handles a specific task. It’s the alternative to fine-tuning, where you retrain a model with collected and labeled data. With prompt engineering, the model stays as it is; you steer it purely through the text of your request.

David Faragó has been working this way since GPT-3 came out. His first serious attempt was about rating commit messages. Instead of going to the trouble of fine-tuning a model, he wrote a specification based on existing guidelines and added examples. The model then graded the quality of commit messages, with surprisingly good results.

For testers, the topic is closer to home than it first sounds. Many prompt patterns are about quality and non-functional aspects. That’s everyday bread and butter in testing.

Why You Can’t Rely on the Correctness of Generated Tests

Language models don’t guarantee correct results when they generate tests. That’s the central limitation, and it matters a lot as soon as the task is specific rather than creative.

David describes two moments that cooled his initial enthusiasm. In the commit message experiment, the model first gave a poor grade with a convincing explanation. A tiny change, a comma turned into a semicolon, flipped the verdict entirely: the same weak commit message suddenly got top marks. The result was simply wrong.

The second setback came from ChatGPT on a statistics problem. The answer sounded plausible all the way through, but contained an error of around ten percent in the middle. David applied the suggested method and spent one or two days in a dead end because he hadn’t questioned the result enough.

With creative tasks like a poem, a deviation in content is acceptable. With test cases, it isn’t. When a generated test fails, you look for the bug in the system under test, while it’s actually sitting in the test.

“The thing can do a lot, but I wouldn’t swear that it always delivers a correct result.”

(David Faragó)

How to Improve a Prompt Step by Step

A good prompt comes about iteratively: you apply proven prompt patterns one at a time and compare the results. Much like design patterns in software development, there are documented prompt patterns, and new ones appear every week.

David picked about five to ten patterns that have worked well in his practice and added them one after another to a single test generation example. He kept the example small on purpose: string processing with Python and pytest.

With each iteration, one aspect got better and another got worse. Test coverage and the structure of the tests improved noticeably. Correctness remained the sore spot. The final result was much better than the first attempt, but not flawless.

Two Levers Decide Quality: Robustness and Correctness

Two levers had the biggest effect on test quality. Robustness describes whether the model returns a comparable result when you change the prompt slightly or just repeat it, instead of producing something completely different.

Concrete examples in the prompt increase robustness. This approach is called few-shot learning. It doesn’t make the tests always identical or always correct, but it rules out the absurd cases: the model copying the entire system-under-test code into the test file, adding unnecessary imports or writing long explanations before and after the test code.

The second lever is correctness, and a simple measure helped here. David was working with a dependency called difflib. The model knew the module and, when asked, described the underlying algorithm, but stayed vague. Only when he pasted the difflib specification, about four to five paragraphs of text, directly into the prompt did the correctness of the generated tests improve noticeably.

The rule of thumb: don’t count on the model to recall its own knowledge precisely. Supply the relevant specification, even if the model theoretically knows it.

Conversation Beats a Single Prompt: The Model as Its Own Reviewer

A single prompt isn’t the end of the road. With coverage criteria in the prompt, David generated eleven test cases from the string example. The previous step, without the extra coverage requirement, had produced five tests, all correct. With the instruction to increase test coverage, eleven tests came out, some of them harder, and two were wrong again.

Instead of optimizing that one prompt further, David switched to a conversation. There’s a dedicated pattern for this: you let the model judge for itself whether its previous output did the job it was given.

That makes intuitive sense. A generative language model produces its output word by word. Judging a finished result after the fact is an easier task than producing it without mistakes. It’s the same with people: telling whether something is good or bad is easier than creating something good yourself. In David’s case, the model recognized the two faulty test cases as faulty.

Which Approaches Actually Help Against Hallucination

Correctness is most likely to improve through well-designed prompt sequences, less through specialized models. David tried Claude, the constitutional AI model from Anthropic, a company founded by former OpenAI employees. Asked whether it hallucinated, the model replied that it was a constitutional AI model and did not hallucinate. The test cases it generated were about as good as David’s first, weak ChatGPT attempt. On correctness, the result was disappointing.

He was more impressed by the ReAct pattern. ReAct stands for reasoning and action and has nothing to do with the front-end framework. The reasoning part is similar to chain of thought, where you tell the model to think step by step before giving a result. The action part teaches the model to make external calls or fetch information from outside.

In a chatbot experiment for MediForm, David saw very high robustness with a cleanly written ReAct pattern. He could correct a single aspect in the prompt and everything else stayed stable. In the handful of experiments he ran, no hallucinations occurred.

The ReAct pattern also shows up in popular tools. Auto-GPT heads in the direction of GPT-based agents and chains a whole sequence of prompts with plugins and external API calls. LangChain, a library for programming such agents, uses a ReAct-like pattern under the hood.

Testers Have an Edge as Prompt Engineers

Testers bring exactly the mindset prompt engineering needs. Correctness is the models’ central weakness, and quality is the core skill in testing.

Many prompt patterns target quality and non-functional aspects. Picking good few-shot examples for a prompt or writing a clean requirements specification comes naturally to testers. A critical, skeptical look at a result is an advantage in prompt engineering, not an obstacle.

How to Start Small

Starting with small, everyday tasks already pays off. The return on investment is high and the effort low.

David uses ChatGPT on the side for manageable problems: refreshing his memory of a library he hasn’t used in a while, or making sense of a long, unfamiliar stack trace. Often the model names the cause of the error right away. Where you used to switch to Stack Overflow or Google, you can now switch to the language model.

The dialog is the real advantage over a search query. When a topic needs more back-and-forth and you want to clarify details with follow-up questions, the model is more useful than a single Google search. Along the way you build experience that will carry over to systematic prompt engineering later.

Where the Field Is Heading: Open Models and More Efficient Fine-Tuning

Much of the progress comes from open models and new fine-tuning techniques. Meta released Llama as open source, although not for commercial use. Unlike pure prompt engineering, you have the model itself and can fine-tune it to improve properties such as correctness.

Model sizes range from about three or six billion parameters up to 13, 30 or 60 billion. Fine-tuning the larger models is hard on the hardware side. At the same time, new optimizations come out every week: quantization, fine-tuning only certain parameters, or storing weight differences instead of the weights themselves.

A leaked Google document titled “We have no moat” argues that neither Google nor OpenAI has a lasting lead, because open-source development and openly published research are closing the gap fast. David considers the document one-sided but sees a grain of truth in it.

In this disruptive phase, it’s hard to predict how much more it will take to build a system robust and correct enough to use with a clear conscience in a safety-critical area. Some things with a big wow effect turn out to be useless in practice, while less flashy approaches become genuinely useful.

Frequently Asked Questions

What distinguishes prompt engineering from fine-tuning a language model?

In prompt engineering, the model remains unchanged; control is exercised solely through the text of the prompt. Fine-tuning retrains a model using collected and labeled data and requires access to the model itself. Meta released Llama in 2023 as open source, though without commercial use, thereby enabling fine-tuning to improve characteristics such as correctness.

Why do errors in generated test cases carry more weight than in creative texts?

In a poem, a deviation in content is tolerable; in a test case, it is not. If a flawed test fails, you look for the error in the system under test, even though it lies within the test itself. An experiment evaluating commit messages demonstrated just how unreliable such judgments can be: Changing a comma to a semicolon turned a poor grade into a top grade.

How can you prevent a model from producing completely different results when prompts are only minimally altered?

Including concrete examples in the prompt increases robustness, an approach known as few-shot learning. It does not guarantee identical or always correct tests, but it does prevent absurd outliers: such as the model copying the entire code of the system under test into the test file, adding unnecessary imports, or writing long explanatory texts before and after the test code.

Is it worth copying a library specification into the prompt if the model is familiar with the library?

Yes. When generating tests using the Python dependency difflib, the model recognized the module and described the algorithm, but remained vague. It wasn’t until four or five paragraphs of the difflib specification were included directly in the prompt that the correctness of the tests improved noticeably. Don’t rely on a model to accurately retrieve its own knowledge.

Can a language model meaningfully evaluate its own output?

Yes, there’s a specific prompt pattern for this: you have the model assess whether its previous response fulfilled the given task. Evaluating a finished result is easier than generating it word for word without errors. In an experiment with eleven generated test cases, the model identified the two flawed tests on its own in the second step.

What sets the ReAct pattern apart from a single prompt?

ReAct stands for Reasoning and Action and has nothing to do with the front-end framework. The reasoning part is similar to Chain of Thought; the model thinks through each step before outputting a response. The action part retrieves information from external sources. In a chatbot experiment, a well-written ReAct pattern remained stable despite changes to individual aspects; no hallucinations occurred in this handful of trials.

Do specialized models with built-in principles solve the hallucination problem?

No. During testing in 2023, Claude from Anthropic responded to a question about hallucinations by stating that it was a Constitutional AI model and did not hallucinate. The generated test cases had then a quality roughly equivalent to that of a first, weak attempt by ChatGPT. Correctness improved more through well-thought-out prompt sequences than through the choice of a heavily promoted model.

What prior experience helps with prompt engineering?

Testing experience. Correctness is the models’ central weakness, and quality is the core competency in testing. Many prompt patterns focus on quality and non-functional aspects. Selecting good few-shot examples and formulating a clear requirements specification are familiar tasks. A critical, skeptical eye toward a result is an advantage here, not an obstacle.

Share this page