Skip to main content

Search...

Shift Left vs Shift Right: You Don't Have to Pick One

Shift left parallelizes testing before go-live, shift right watches production. One insurer runs the same UI tests in both worlds. How that works.

• • Updated: • 12 min read
Cover of the expert talk on 'Shift Left vs Shift Right: You Don't Have to Pick One' with Björn Scherer and Richard Seidl.

“Shift left on the right” answers the shift left vs shift right question with a combined testing approach: development and testing run in parallel early on, and after go-live the same tests keep an active eye on production. Test cases are written before coding starts, run several times during the sprint and are then reused as synthetic monitoring in production. That way one pool of test artifacts catches defects before customers notice them.

Key Takeaways

  • Synthetic monitoring with reused UI test cases detects failures in production before real customers run into them.
  • Test automation before go-live and synthetic monitoring in production share the same artifacts, so nobody has to build and maintain two separate test suites.
  • Fast feedback in the sprint comes from a deliberately trimmed test set: the tests for the current feature plus prioritized regression tests, up to the time limit the team has chosen for itself.
  • In production, tests may only use predefined synthetic data, and every run has to identify itself as monitoring so that customer analytics and web tracking stay accurate.

Shift Left vs Shift Right: Why You Can Use Both

Shift left vs shift right is not an either-or choice. Shift left means moving testing earlier and running it in parallel with development. Shift right moves part of the observation into production, after go-live. At Cosmos Direct, the two work together: functional test cases are written before coding starts, and later the same UI tests keep running as synthetic monitoring in production.

Why Testing Still Piles Up in Agile Teams

Even in agile teams and under DevOps, testing activities pile up right before go-live. This isn’t a waterfall problem that agility made go away. Each feature, taken on its own, follows a linear path: at some point the team starts on it, at some point it goes live, and in between it is developed and tested.

A team doesn’t work on every feature at once from day one. It starts with one, brings it to maturity and moves on to the next. That is exactly what creates the testing backlog before go-live. And in the end, testing usually gets the blame for not being finished.

Björn Scherer, who comes from the insurance world at Cosmos Direct, describes two opposite responses to this pattern: shift left pulls activities forward, shift right moves part of the observation into production. Both have their appeal. The interesting question is how to take the best from each.

What Shift Left Means: Working in Parallel, Not Just Starting Earlier

Starting earlier on its own achieves nothing. If you simply move the same mechanics forward, the bottleneck forms somewhere else. The real lever is cutting work into pieces small enough to be handled in parallel.

Instead of finishing a user story as one block and then handing it on within the team, get individual acceptance criteria far enough along that the first can already be looked at while the second is still being built. As soon as the first is tested, the second gets developed.

A bigger step is having several disciplines look at a feature at the same time. At code level, that leads toward test-driven development; at the business level, toward acceptance test-driven development. The principle stays the same: develop and test in parallel instead of one after the other.

Business Logic Can Be Written Early, Glue Code Cannot

The valuable part of a test is its functional definition, and you know that from the start. Before anyone builds a feature, it is clear what the feature should do. You can write that part in executable form right away, or in a form that becomes executable later.

What you can’t bring forward is the technical part. Selecting a field in Selenium only works once the UI exists. But that glue code is the last step, not the first. The business-driven test cases come before it.

The goal is test cases that contain no technology at the top level and are driven purely by the business requirements. A keyword-driven approach gets you there. Whether a team goes with Gherkin or keyword-driven testing is up to the team.

Fast Feedback Needs Test Sets Sized to the Team’s Pain Threshold

Fast feedback for development teams requires more frequent deployments, at least to the test environment. Deploying opens up higher test levels: not just unit tests in the build pipeline, but API tests and, beyond that, UI or end-to-end tests. They catch bugs that unit tests miss.

Nobody runs thousands of UI tests in ten minutes, so each team works with a subset. The core is the test cases for the current feature. The set is then topped up with regression tests until it reaches the pain threshold for fast feedback.

The team sets that time limit itself: ten minutes, fifteen, half an hour, depending on its tolerance. The size of the test set follows from it, and it can be recut every sprint. Regression cases are picked through tagging and prioritization, not generated automatically.

Larger applications can stagger this: the focused set during the day, a broader regression set overnight and, if needed, an even larger one every week. Feedback stays fast without permanently giving up breadth of coverage.

Shift Right with Synthetic Monitoring: Testing When No Customers Are Online

Classic monitoring analyzes what real users do on the system. It writes logs and builds reports on top of them. The catch: it only works while customers are active. During the lunch break or at night, there is no data.

Synthetic monitoring therefore runs use cases proactively, on a schedule, whether or not anyone is using the system at that moment. A defect shows up before the first customer notices it. In the best case, it is fixed before anyone notices anything at all.

The key difference from pure backend tools is perspective. Tools like Dynatrace show which service has a fault or where things are slowing down. From a business point of view, that still leaves open where the customer is actually affected.

“We come at it from the opposite direction. Because we run it from the customer’s point of view, we know that straight away, and then we use the other tools to analyze the error.”

(Björn Scherer)

Reusing Existing Test Artifacts in Production

Reuse is the heart of the approach. The artifacts created to the left of go-live shouldn’t be left behind at release. They should keep paying off to the right of it, in production.

From the regression set, which has to be ready by acceptance at the latest, the team picks one to three especially relevant use cases. These then keep running as synthetic monitoring. By acceptance, the tests have to match the application anyway; otherwise you’d be putting an application live whose tests no longer work.

The logic comes full circle: if a use case matters enough to run in production, it matters enough for the fast feedback loop. These use cases are the first to be updated when something changes. Most applications have one or two core use cases that qualify.

At Cosmos Direct, these are the online application journeys. At the end of a journey, the insurance policy is actually concluded, not just an application sent off. For car insurance, there are around 1.4 million paths through the journey. Testing all of them is impossible. For operations, it is enough to prove that the happy path and two or three variants run through.

Why UI Tests, Even Though They Are Considered Fragile

The obvious question is why use fragile UI tests instead of more stable API tests. There are several reasons, and stability is something the team has worked on a lot, independently of monitoring.

  • Reuse: The UI and end-to-end tests already exist from development. Nothing new has to be built.
  • User perspective: UI and end-to-end tests map real user journeys. That perspective is much harder to get from API tests alone.
  • No extra tooling: An additional monitoring tool would need someone in the cross-functional team who knows how to use it. If that is one person, the bus factor is one. Sticking with the existing test stack removes that risk.
  • Early start: The tests can be built into the standard test process early, so problems with the monitoring use cases surface early.

Test flakiness remains the big issue. Anyone extending tests into production has to invest in their stability first. Slower response times show up indirectly: if the UI responds more slowly than the configured timeout, the test fails. The tolerances were raised on purpose so that flakiness doesn’t set off constant false alarms.

Test Data in Production Is the Real Hurdle

Getting tests to run in real production was a long internal battle. Real customer data is off-limits. For every use case, the team defines which data it needs and what has to exist in the system. This synthetic data is loaded into production once, and the tests may only work with it.

The use case has to fit the data. An application journey behaves differently for an existing account than for a new customer. Before a use case is even allowed into production monitoring, it has to prove in pre-production that it won’t break anything.

For that, the applications must be able to tell synthetic traffic from real traffic. The use case signals that it is the monitoring run. Otherwise you skew the numbers: for a niche product, monitoring every 15 minutes can generate several times the real number of customers and distort every business statistic.

Web tracking from e-commerce must not be skewed either, which can be handled by switching it off via cookies. In practice, your test framework needs a monitoring mode. The same test case uses different data in the test environment than in production. It takes some tuning, but it is doable.

Alerting and Reporting Make the Results Actionable

Synthetic monitoring is only as useful as the response to it. A central control center is running anyway and receives the alerts, built into the existing company processes. That pays off especially at night, when the team isn’t on site.

The next step reverses the route and alerts the teams directly. The team that develops an application is responsible for it. A fast lane sends the problem straight back to them, while the official alerting path stays in place.

Whether a team uses synthetic monitoring at all is up to the team. The idea behind the DevOps migration is: take responsibility for your application. If a team says it has everything under control, that’s fine too. The central job is to provide the infrastructure and methodology that teams can plug their use cases into.

A dashboard shows a central on-call service at a glance whether a single application is affected or there’s a wildfire. The next stage goes deeper: assessing causes from the user’s point of view, spotting the steps where things go wrong and aggregating that over time.

Frequently Asked Questions

Does the testing bottleneck disappear right before release when a team works using agile or DevOps methods?

No. Even in agile teams, testing activities pile up before go-live because each individual feature follows a linear process: it’s initiated, developed, tested, and then goes live. A team doesn’t work on all features simultaneously from the start; instead, it brings one to maturity and then moves on to the next. This is exactly what causes the bottleneck at the end.

Is it enough to simply start testing activities earlier?

Starting earlier alone just shifts the bottleneck to another point. The key lies in making smaller adjustments: advancing individual acceptance criteria to the point where the first can be tested while the second is still being developed. The next step is to examine a feature from multiple disciplines simultaneously: at the code level toward TDD, and from a business perspective toward ATDD.

Which parts of an automated test can be written before the software exists?

The business definition can be done in advance, but the technical part cannot. What a feature is supposed to do is determined before coding begins and can be formulated in a way that is either executable now or executable later. Selecting a field in Selenium only works once the UI exists. This glue code is the final step. Whether to use Gherkin or keyword-driven testing is up to the team.

How large can an automated test suite be while still ensuring feedback remains fast within the sprint?

The team sets the time limit itself: ten minutes, a quarter of an hour, or half an hour, depending on their tolerance. The size of the test set is based on this limit. The core consists of test cases for the current feature, supplemented with regression testing up to this threshold. Selection is based on tagging and prioritization, not automatic generation, and can be tailored anew for each sprint.

What does synthetic monitoring offer that traditional monitoring cannot?

Traditional monitoring analyzes what real users do and therefore requires active customers. At night or during lunch breaks, there is no data available. Synthetic monitoring runs use cases cyclically and independently of actual usage, so that a defect is detected before the first customer notices it. Backend tools show the affected service, but not where the customer is experiencing an issue.

Aren’t more stable API testing methods better suited for monitoring in production than UI tests?

UI and end-to-end testing have four advantages: They already exist from the development process; they map real user journeys; they avoid the need for an additional tool (and the single point of failure associated with having only one person who knows how to use it); and they can be integrated early into the standard testing process. This requires work on stability, as well as intentionally increased timeout tolerances to prevent false positives caused by flakiness.

What data can tests use in a production environment?

Only predefined synthetic data; real customer data is off-limits. For each use case, it is specified which data is required and which must be available in the system. This data is loaded into production once. The use case must match the data context, since an application workflow behaves differently for an existing account than for a new customer. Prior to deployment, it must be verified in pre-production that nothing breaks.

Can monitoring runs in production skew business analytics?

Yes, if the application cannot distinguish synthetic traffic from real traffic. For a niche product, monitoring every 15 minutes can easily generate many times the actual number of customers, thereby distorting all business statistics. The use case must therefore indicate that it is a monitoring run, and web tracking must be disabled, for example via a cookie. To do this, the test framework needs a monitoring mode.

Share this page