Skip to main content

Search...

BDD Testing: Given-When-Then for Business and Developers

BDD testing uses Given-When-Then scenarios so business stakeholders and developers share one understanding of a feature before any code is written.

• • Updated: • 10 min read
Cover of the expert talk on 'BDD Testing: Given-When-Then for Business and Developers' with Pascal Moll and Richard Seidl.

Behavior Driven Development (BDD) describes how a feature should behave in a structured language that the business side and the development team understand equally well. The core pattern is Given-When-Then: precondition, action, expected result. The format bridges different mental pictures, serves as the basis for automated tests and keeps requirements and test code consistent in one shared place.

Key Takeaways

  • The Given-When-Then syntax of BDD gives business and development a shared language that heads off misunderstandings in feature descriptions.
  • Cucumber links readable text descriptions directly to executable test code, so even people without a technical background can see what a function is for without knowing how it’s implemented.
  • In BDD, test data can live as data tables right in the feature file or in external files. An incomplete data set from the start is the most common pitfall.
  • Functions in test code shouldn’t have more than eight executable lines. If a function grows beyond that, it’s a candidate for splitting.
  • BDD tests need adjusting far less often than UI tests, which justifies the extra effort of introducing them in the long run.

What BDD Testing Actually Describes

BDD testing starts before the first line of code. Behavior Driven Development describes how a feature behaves, in a language that business and technical people understand equally well. The focus isn’t the code but the question of what a feature is supposed to do.

Without BDD, the classic process is lopsided. A feature gets written up, and at first mainly the product owner knows what it’s about. The development team joins later and has to work out what matters.

BDD steps in before that. It makes everyone involved describe the desired behavior so that no technical jargon is needed and nobody drifts off. The result is a common basis that the business side and the technical side can both refer to.

How the Given-When-Then Structure Works

BDD scenarios follow a Given-When-Then structure that breaks a test into three clear phases. The Cucumber tool uses the Gherkin language for this.

For a browser test, it looks like this:

  • Given: The browser is open and the page under test is loaded. That’s the starting situation.
  • When: The login form is filled in and the OK button is clicked. That’s the action.
  • Then: With valid data, the page behind the login appears. With invalid data, an error message appears. That’s the expectation.

This form is compact and still readable for everyone. A developer sees what needs to be done. Someone from the business side who has thought it through understands what it’s about without any technical knowledge.

That overlap is the whole point. Prose creates a different picture in each reader’s head, and those pictures rarely match. Often you don’t even notice the difference. The Given-When-Then structure narrows that room for interpretation and clears misunderstandings out of the way.

BDD Is Teamwork, Not a Matter of Roles

Nobody has the exclusive right to write BDD scenarios. Both the business side and a test analyst can write one. There’s no fixed ownership.

In practice, analysts often provide a first draft and testers review it. Some of it can be taken over as is, some needs to be added to or adjusted. Every change goes back to the business side: is this still right, or has something drifted?

That back-and-forth is the real value. Talking and agreeing with each other, not the format itself, creates shared understanding.

From Description to Executable Test

Once a scenario is worded well enough, implementation follows. In Cucumber, the steps of a scenario can be generated as a skeleton that is then filled with code, the so-called glue code.

The glue code holds the programmed functionality and is typically written by test automation engineers or developers. The Given, When and Then steps reappear in the function names, which creates a one-to-one mapping.

The effect: even someone from the business side could look at the test code and see which functionality maps to which step. Exactly how it was implemented is secondary.

The finished tests run on a build server such as Jenkins. Depending on the test strategy, they run as soon as a commit comes in.

How BDD Handles Test Data

In BDD, test data can be stored directly as data tables in the feature file. The tables are filled in and feed into the automated test at runtime. Who maintains them, the business side or the testers, is open.

Alternatively, the data can sit in external files and be pulled in from there. Both approaches work.

The groundwork is what matters. Test data should be meaningful and complete before you start. Extending it later is possible, but it often means reworking the tests. Test data management deserves serious attention up front.

One Source of Truth via Jira and the Build Server

BDD only pays off when scenarios and code stay in sync. In theory the business side could write directly in Cucumber, but that’s not common.

One practical route runs through Jira. Scenarios are stored there and downloaded into the feature files via a plugin, or uploaded again when they change. The exchange works in both directions and can be driven by the build server.

When a developer pushes a change, the build server pulls the updates and syncs the current state back. The business side always sees the latest version of the feature files in Jira.

That solves an old problem. Several storage locations quickly raise the question of which version is current. BDD instead relies on one source of truth in one place.

Where BDD Testing Fits and Where Its Limits Are

BDD is flexible because you can put practically anything into the glue code. It isn’t limited to UI tests. It also works for integration and acceptance tests, extended with additional libraries.

Even manual tests are possible. In that case, the Given-When-Then statements are prepared without executable glue code behind them. They serve as standardized documentation that a manual tester works through.

Load testing is a clear limit. Technically it can be done, but BDD isn’t the right tool. Reach for something else instead of bending BDD out of shape.

Why Getting Started Begins with the Right Test Case

The best starting point is a test case that is neither trivial nor overly complex. It should have a visible impact, but it shouldn’t drag on for months.

Behind this is an honest insight: BDD comes with overhead. Coding a test case directly is faster than also creating feature files and a well-designed Gherkin structure. The extra effort only pays off later, and that has to be communicated clearly from day one.

“This BDD implementation has a certain overhead that only pays off later. You have to be aware of that, and you have to be careful about it from the very beginning.”

(Pascal Moll)

Maintenance: BDD Tests Age More Slowly Than UI Tests

BDD tests need maintenance like any other test, with one extra step. When circumstances change, the Given-When-Then statements have to be updated to match.

The good news: in practice, that happens much less often than with pure UI tests. UI tests take noticeably more maintenance.

A handy way to avoid redundancy in Cucumber is the Background keyword. It lets you define steps that always run at the start of a test. That saves duplicated code and keeps the scenarios easier to read.

Test Code Is Code: Clean Code Pays Off

BDD scenarios produce real code, and that code needs the same quality as production code. The complexity of a programmed function shouldn’t go beyond a healthy level.

A practical rule of thumb: once a function goes past eight lines, it’s worth thinking hard about whether it can be split. A subfunction adds clarity. This heuristic is easier to apply than formal complexity models, where you first have to identify nodes and edges in the code.

Reviews and walkthroughs help, too. A four-eyes principle or a look by the whole team uncovers improvements. Scheduled refactoring days, say one or two days set aside for improving code, demonstrably reduce complexity.

Where AI Realistically Supports BDD Today

AI can assist with BDD, but it doesn’t replace human review. There are plugins that make suggestions during testing, such as predictions based on data.

One concrete approach: a test’s history, meaning how often it passed or failed, can be used to derive predictions. When a developer pushes code to the repository, the goal is to predict which defects this code could trigger and which tests might fail.

It also seems likely that AI will suggest Given-When-Then structures in the future. The real work will then be checking the suggestion: does it fit, does it match expectations, does it need small adjustments? That will come as an aid, but not as a replacement for the people involved, at least not for now.

Frequently Asked Questions

Why do requirements formulated in prose often lead to misunderstandings during testing?

Prose text conjures up a unique image in every reader’s mind, and these images rarely coincide. Often, the difference isn’t even noticeable. Instead, the “Given-When-Then” structure breaks down a requirement into a precondition, an action, and an expected result. This reduces the room for interpretation, ensuring that both the business side and the development team are referring to the same description.

Who on the team is allowed to write BDD scenarios?

There is no fixed responsibility. Both the business team and a test analyst can write a scenario. In practice, analysts often provide a first draft, which the testers review: some parts can be adopted as is, while others are supplemented or adapted. Any changes are sent back to the business team for clarification.

Does BDD work even without test automation?

Yes. The Given-When-Then statements can be prepared without any executable glue code behind them. They then serve as standardized documentation that a manual tester works through. BDD can also be used beyond UI testing, for example for integration and acceptance tests, because virtually any functionality can be incorporated into the glue code.

Is BDD suitable for load testing?

No. Technically, it’s feasible, but it’s not the right tool. For load testing, you should use specialized tools instead of forcing BDD to do something it wasn’t designed for. The strength of this approach lies in behavior that can be described in business terms, that is, in UI, integration, and acceptance tests that can be meaningfully structured into preconditions, actions, and expectations.

Why should test data be complete before testing begins?

An incomplete data set is the most common pitfall. While it’s possible to add data later, this often requires reworking the tests. The data can be stored directly in the feature file as data tables or written to external files from which it is retrieved during execution. Both approaches work; the key is the preparatory work.

Is BDD worth the effort despite the additional work?

In the long run, yes; immediately, no. Writing a test case directly in code is faster than creating additional feature files and a well-thought-out Gherkin structure. This overhead only pays off later and should be communicated openly from the start. The payoff comes through maintenance: experience shows that BDD tests need to be updated significantly less often than pure UI tests.

How do you ensure the maintainability of automated test code in the long term?

Test code deserves the same quality as production code. A practical rule of thumb: If a function exceeds eight executable lines, it’s worth giving serious thought to splitting it up. This is easier to apply than formal complexity models, which first require nodes and edges in the code. Reviews based on the four-eyes principle and scheduled refactoring days (about one to two days) further reduce complexity.

Can AI take over the creation of BDD scenarios?

As an assistant, yes; as a replacement for the people involved, not yet. There are plugins that make suggestions, such as predictions based on a test’s history: How often did it run successfully, how often did it fail, and which tests might a new push break? Suggestions for Given-When-Then structures are also obvious. Verifying the suggestion remains a human task.

Share this page