Skip to main content

Search...

Selenium Automation Testing: A Browser Tool First

In Selenium automation testing, Selenium only drives the browser. Good selectors, page objects and the vendor drivers in Selenium 4 keep tests stable.

• • Updated: • 13 min read
Cover of the expert talk on 'Selenium Automation Testing: A Browser Tool First' with Boris Wrubel and Richard Seidl.

Selenium is an open source framework for browser automation that remotely controls all common browsers through standardized WebDriver interfaces. Test automation with Selenium is mainly about web applications: page elements can be read, clicked and filled in. Selenium supports many programming languages and has become much more stable since version 4, because browser vendors now ship the WebDriver themselves.

Key Takeaways

  • Selenium is a browser automation tool rather than a testing tool: it remotely controls any browser through the WebDriver, regardless of vendor or operating system.
  • Teams that know CSS and XPath selectors well and use the Page Object pattern consistently keep maintenance effort low, even when the UI changes often.
  • Writing tests in the same programming language as the product makes developers far more willing to fix failing tests themselves and write their own test cases.
  • Combining Cucumber with Selenium shows the product owner which business functionality a test covers, without anyone having to read the code.

What Selenium Really Is: Browser Automation, Not a Testing Tool

Selenium is an interface for automating browsers, not a testing tool in the strict sense. When people talk about Selenium automation testing, that distinction often gets lost. Tenders and job ads list Selenium as “the test automation tool”, while the Selenium project itself describes it as a way to control browsers remotely.

Even so, its main practical use is test automation. Since so much software is built as web applications, it makes sense to control a browser precisely and run test cases through it.

Technically, the WebDriver sits between Selenium and the browser. Each browser has its own WebDriver, which does the actual remote control. Selenium talks to that WebDriver and hides the differences between Chrome, Firefox, Safari or Opera. You write your test once and do not have to deal with the details of each browser.

Since the WebDriver standard, browser vendors are responsible for making this connection work. Anyone who releases a web browser ships the matching WebDriver support with it. The result is more stability, more efficiency and better performance than in the early days.

Which Programming Languages Work with Selenium

Selenium comes with bindings for many programming languages, and that is what makes it so flexible. Tests can be written in scripting languages such as Ruby or Python, and just as well in Java, Kotlin or C#. There are more implementations than any single developer could ever master.

You integrate Selenium through the library for your language. In Java or Kotlin, you add the dependencies to your project and use the Selenium library to drive the browser.

That freedom has a concrete consequence for teams: you can pick the language the project already uses. Test automation then stays inside the technical world of the product instead of sitting off to the side.

How Selenium Automation Testing Controls the Browser

Selenium works on two levels, much like a person using a browser. At the browser level, you open tabs, navigate, reload pages and change settings. At the page level, you read information and interact with the web elements.

At the element level, you query states, click, type text, select values or drag and drop. That covers most of what a user does in a browser.

Browser settings can be controlled the same way. A typical example: a web application offers a PDF for download, and the browser opens it in its built-in viewer by default. A setting can force the browser to download the file instead. Configurations like this are among the adjustments that keep a test run stable.

Good Selectors Decide the Maintenance Effort

Your choice of selectors decides how much maintenance your test automation needs. For simple cases, an ID or tag name is enough. Once things get more complex, you need solid knowledge of CSS selectors and above all XPath selectors. Combine them well and the maintenance effort stays small.

A common mistake is the absolute XPath from the browser’s context menu. Right-click, “Copy XPath”, and a path from the root element all the way to the target element ends up in the test code. If one small thing on the page changes, the test breaks, and nobody can tell anymore which element was meant. Selectors should be readable.

There are libraries that automatically swap a failing locator for a new one. But a test that fails after a UI change is not a flaw. It is a signal. The test noticed that something changed, and that is exactly its job.

If you follow the Page Object pattern, you change a locator in exactly one place, even if a hundred test scripts use it. A clean structure keeps maintenance to a minimum.

Selenium Is Development Work

Many problems that seem to come from Selenium actually come from the programming language in use. Test automation with Selenium is development work, not a click script.

“When you have a lot of problems that you think you have with Selenium, they usually aren’t with Selenium itself, but with the programming language you used.”

(Boris Wrubel)

That leads to a clear recommendation for choosing the language: use the language the product is built in. It brings the developers on board. When a test run fails, they find the cause faster and can step in themselves, both to fix it and to write new tests.

This approach fits the Scrum idea that there is no separate tester standing apart from the team. Quality becomes a job for the whole team instead of something handed off to one person.

Bringing in the Business Side with Cucumber

The business side comes in through Cucumber. Selenium combined with Cucumber shows product owners and the business what a test checks, without them having to read any code. Well-written scenarios show the tested business functionality directly.

Who writes the scenarios is a question of its own. Leaving it entirely to the business department is not ideal. It works better to write them together, so the business view and the technical feasibility meet.

Cucumber steps need a sensible granularity. Endless scenarios in which every single UI step is spelled out with heavy parameterization become hard to follow. At that point it is fair to ask what Cucumber is for at all. If every click is a step, you might as well write the flow cleanly in the test class. Cucumber pays off only when the scenarios express business functionality, not mouse movements.

How to Scope UI Tests

UI tests belong at the top of the test pyramid, which means there should be deliberately few of them. You pick one business case and run through it at the business level. Variants that change nothing important on the UI belong in integration and unit tests.

Early in a project, things can look different. While the software is still small, teams like to test several variants through the UI and check whether things are displayed correctly, just to have something in place. You can delete or archive those tests later, once the lower levels take over the coverage.

This discipline keeps the slow, expensive UI layer from filling up with a hundred scenarios that would be better off elsewhere.

Why Good Reporting Decides Whether the Team Accepts It

Reporting is the lever that gets developers involved. When a Selenium test fails, the report has to show the cause quickly. Otherwise test automation stays an island.

A useful report shows the step where the test failed and the reason: an element that could not be identified, or a backend service that was not available. It includes a screenshot, the page’s HTML and, depending on the need, one or two extra pieces of information. Videos of test runs are possible but usually only worth it as an exception, when normal reporting is not enough.

The key question a report has to answer is simple: is the bug in the application or in the test? If you can tell quickly, the barrier for the team stays low.

Jenkins offers one of the best Cucumber report engines for pipeline integration. Cucumber writes its results to a JSON file, and there are open source tools on GitHub that turn that file into a readable report and sometimes generate a PDF from it, including screenshots and other attachments.

SeleniumCucumberGrow: Up and Running in Under an Hour

A common objection is that setting up from scratch takes too much effort, even though Selenium itself costs nothing. The open source project SeleniumCucumberGrow tackles exactly that problem. Its name stands for Selenium, Cucumber and Grow.

The tool is a project generator. You enter a project name and a package name, and it creates a runnable Selenium project with an example that searches a Wikipedia page for “Software Testen”. That removes the tedious renaming of packages and classes that used to make copying an old project such a chore.

The project’s promise is in the title of its conference talks: Get started with less than one hour. Within an hour, the first test case runs against your own website. The project is kept up to date with current Selenium versions and dependencies and uses Cucumber’s built-in reporting, extended with the information that matters for fast troubleshooting. An extension for accessibility testing is in the works that highlights the element that failed an accessibility check.

What You Need for a Clean Start

The biggest pitfall is having no concept for your test automation. A working example alone is not enough. You need an idea of how to approach automation before you swap the Wikipedia demo for your own application.

Two things carry the start:

  • Follow the Page Object pattern. It keeps locators in one place and the tests maintainable.
  • Know your CSS and XPath locators. Without that knowledge you end up with brittle absolute paths that break at the slightest change.

Modern IDEs such as Eclipse or IntelliJ help with good plugins that speed up test development. Combined with the Page Object pattern, that really does get the first test case against your own website running in a short time.

What Is New in Selenium 4

Selenium 4 brought real change after a long announcement phase. Version 3 was around for a long time, and the move to 4 took a while, but it is well established now.

Relative locators are new. You search for elements to the left of, to the right of, above or below another element. For certain layouts that makes locating elements simpler.

The bigger relief is WebDriver handling. You used to need the matching ChromeDriver for your operating system for every Chrome version, put the binary in the right directory and set the path. Boni Garcia’s WebDriver Manager, originally a separate GitHub project, has been built into Selenium as Selenium Manager. Since the 4.6 and 4.8 releases, driver handling works out of the box, and you no longer have to worry about downloading the right driver.

Where Test Automation with AI Is Heading

Artificial intelligence will change test automation, most likely as support at first. Early approaches hand parts of testing to AI, especially visual checks.

In mobile testing, there are libraries that take an instruction such as “press the login button” and work out on their own which element the login button is, whether through machine learning or other methods. This kind of assisted automation is clearly on its way.

The vision of a test bot that you simply hand your website to, which then tests on its own and sends back a report, is further off. And it immediately raises the question of trust: how sure can you be that the bot is checking the right things? Or, put more sharply: who tests the test automation?

Frequently Asked Questions

Why does Selenium need a separate WebDriver for each browser?

The WebDriver handles the actual remote control. Selenium communicates with it, thereby hiding the differences between Chrome, Firefox, Safari, or Opera, so that a test only needs to be written once. Since the WebDriver standard was introduced, browser vendors have been providing the necessary support themselves. This has resulted in greater stability and better performance compared to earlier solutions.

Can browser settings also be changed using Selenium?

Yes. Selenium operates on two levels: At the browser level, you open tabs, navigate, reload pages, and set settings; at the page level, you read elements and interact with them. A typical example is a PDF that the browser would otherwise display in its built-in viewer. You can force the browser to download it by adjusting a setting. Such configurations help stabilize test runs.

Why do automated UI tests fail even with minor changes?

The most common cause is absolute XPath expressions that are copied from the browser into the test code via right-click and “Copy XPath.” They span from the root element to the target, fail at the slightest change, and make it impossible to tell which element was intended. Readable CSS and XPath selectors, combined with the Page Object pattern (which centralizes each locator in a single location) keep maintenance efforts to a minimum.

Which programming language should you choose for test automation?

Ideally, the language used to build the product. Selenium is available as a library for Ruby, Python, Java, Kotlin, C#, and other languages, so the choice is yours. Using the same language gets developers on board: they find bugs faster during test runs, fix them themselves, and write their own test cases. Test automation with Selenium is development work.

How many scenarios belong at the UI level?

Few. UI tests sit at the top of the test pyramid: You take a business case and run through it at the business case level. Variations that don’t significantly change the user interface belong in integration and unit tests. At the start of a project, it’s acceptable to perform more testing of the UI just to ensure basic coverage. Such tests can be deleted or archived later.

What should be included in a report so that a team takes failed tests seriously?

The report must show the step at which the test failed and the reason, such as an unidentified object or an unavailable backend service. It should also include a screenshot and the page’s HTML code. The key question is: Does the application have a bug, or is it the test? Videos of the test runs are usually only worthwhile in exceptional cases.

Do you have to provide the appropriate WebDriver binaries yourself?

This used to be necessary: For each Chrome version, you had to obtain the appropriate ChromeDriver for the respective operating system, place the binary in the correct directory, and set the path. Boni Garcia’s WebDriver Manager, originally a separate GitHub project, was integrated as the Selenium Manager. Starting with versions 4.6 and 4.8, driver management was included out of the box in 2023.

Can artificial intelligence take over testing of web applications?

Initially, it will serve more as a support tool, particularly in the visual domain. In the mobile sector, libraries existed in 2023 that could accept a command such as “press the login button” and determine on their own which element was intended. A bot that you simply feed a website to, and that performs independent testing and delivers a report, was still a long way off. The question of trust remains: Who tests the test automation?

Share this page