Skip to main content

Search...

Shift Left: Quality Starts with the User Story

The shift left approach builds quality in early: vertically sliced user stories, feedback in the IDE, solid unit tests and contract testing with PACT.

• • Updated: • 14 min read
Cover of the expert talk on 'Shift Left: Quality Starts with the User Story' with Alexander Vukovic and Richard Seidl.

Shift left means building quality assurance into the development process as early as possible instead of bolting it on at the end. It starts with well-sliced user stories, continues with tool-supported code quality right inside the development environment and rests on broad unit test coverage as the foundation for every test level above it.

Key Takeaways

  • Unit tests are the precondition for any shift left strategy: without this fine-grained safety net at code level, enough coverage on higher test levels cannot be reached at a reasonable cost.
  • Consumer-driven contract testing with tools such as PACT lets teams check API compatibility between services locally against automatically generated mocks, without setting up a full integration environment.
  • Quality feedback inside the IDE, such as cyclomatic complexity shown inline, beats a later build server report because the developer fixes the problem on the spot instead of going through an extra correction loop.
  • Teams where every developer builds a user story alone and testers only step in at the end of the sprint recreate a mini-waterfall inside the sprint and risk finishing no story at all.
  • A blanket target of 100 percent unit test coverage backfires. It works better to mark critical business logic explicitly and require full coverage for exactly that code.

What the Shift Left Approach Means in Quality Assurance

The shift left approach secures quality as early as possible in the development process rather than at the end. Left and right refer to the timeline that runs from the requirement to delivery. Moving quality to the left simply means doing the quality work as early as you can.

The idea works in layers. Quality can already be secured in the user story. The team then builds broad coverage at unit test level. The other quality levels build on top of that: they interlock, complement each other and overlap as little as possible. A continuous integration and delivery pipeline then ships the result with its quality checks in place.

None of this is new. Decades of software development experience led to agile processes in which developers and testers work together as a matter of course again. With DevOps, operations is now growing into the same integrated approach.

The Requirement Is the Biggest Lever for Quality

The biggest lever sits at the very beginning: requirements that are well described and can actually be implemented. Agile processes have long had established methods for slicing user stories, reviewing them and checking them against quality criteria, exactly at the point where that is cheapest.

The real sticking point sits one level above the individual stories, in the team’s mindset. In many teams, structures that grew over years decide how stories get sliced. A classic pattern: at the start of the sprint every developer grabs a story and builds it alone, and two testers check everything at once in the last two days. That is a mini-waterfall inside the sprint.

The reasoning behind it sounds harmless, and that is the problem. “We don’t want to get in each other’s way” is an attitude that makes real collaboration hard. Stories end up sliced by what one person can do alone, not by the value the result delivers.

Slice User Stories Vertically, Not for Convenience

The value of the result should drive the slicing, not developer convenience. A vertical slice cuts through every layer of a feature and delivers something that runs, can be tested and can be used by the end of the sprint.

A horizontal slice produces dead intermediate states. If a story contains only the user interface of a feature and the database behind it is built three sprints later, you have something that can be neither tested nor used. It adds no value.

Every iteration should end with a finished piece of value: quality-assured, working and ready to ship. Vertically sliced stories are easier to secure, because a complete feature can also be tested as a complete feature.

Scrum works with a prioritized sprint backlog. In theory, the whole team works on the top-priority story until it is done, then moves to the next one. Of five committed stories, maybe only three are finished in the end, but those three are really done. With the mini-waterfall, there may be nothing finished at all, because every story was started at once and none was completed.

“Stop starting, start finishing.”

(Alexander Vukovic)

Why Teams Resist Working Together

Clinging to the old way of working usually comes from fear, not from a lack of knowledge. Often it is the biggest knowledge holders, people who have been with the company for decades, who worry that sharing what they know will make them replaceable.

To them, a cross-functional team looks like a threat, because others are supposed to learn the same things. Changing that mindset takes support over several iterations and showing, in practice, that the fear is unfounded.

The argument does not hold up anyway. A junior will never make up in three months what someone has built up over thirty years. But if the junior can take over certain tasks when the experienced colleague is out, everyone benefits. Strictly speaking, this is the Scrum Master’s job. In practice, many Scrum implementations optimized that role away right at the start.

The IDE Brings Quality Straight into the Code

The IDE matters more than many people expect. A lot of quality can go into the code through the development environment, long before the build server even starts.

Refactoring features are one example. Renaming a class method consistently across the whole project, instead of editing every file by hand, is a small thing modern IDEs simply do. In the Java world IntelliJ is the most common choice, on the .NET and C# side it is Visual Studio, and in between sits the free, lean and fast Visual Studio Code.

Plugins give immediate quality feedback in the editor. One plugin for Visual Studio Code measures the cyclomatic complexity of a method inline while you write it. Add too many nested conditions and it tells you, right above the code, that the complexity is too high.

That feedback is part of shift left. If the same static analysis only ran on Jenkins or GitLab, you would go through one more loop: wait, look at the report, get told off, rework the code. In the editor you see the problem as you type, and the whole process saves time.

Tools such as SonarQube are well established for static code analysis in enterprise environments. But when the same feedback is available instantly, it should be used instantly instead of waiting for server runs and filled dashboards.

You Cannot Test Code Quality in Afterward

Code quality is either there or it is not. No amount of testing can test it into the software, however good the testing is. If you ignore quality criteria such as naming, coding conventions or avoiding duplicated code, you build on a base that will be hard to keep working on later.

The mechanism is easy to predict. When a team is pushed to deliver as much functionality as possible in a short time, it ships whatever it can, quality or not. That works, but not for long. At some point it backfires, and it is usually the developers who pay for it, because the pressure to deliver stays and even grows if the software succeeds.

That leads to a simple rule for daily work: whatever quality you can get with tool support and without significant extra time, invest in it. It is a basic requirement for software that keeps its value over the long term.

How Much Unit Test Coverage Makes Sense

Unit tests are the base of shift left, the fine-grained safety net. They are small regression tests that run in isolation and make sure the code still does what it did before. With that safety net in place, you can make changes, even large architectural ones, knowing that any errors you introduce will show up.

No other test level delivers this coverage. Reaching the same coverage on higher levels would take a disproportionate investment that no business can justify.

The ideal unit test coverage is not 100 percent. Many things cannot be tested sensibly at unit level. A unit test should run in isolation, check one small unit and finish in milliseconds or at most seconds, not hours. End-to-end and integration tests belong on the levels above.

The standard example is getters and setters. If you test a setter just to hit the coverage target, you have successfully tested the assignment operator of Java or .NET. That adds nothing. It is waste.

Then there is legacy code. If a team has worked for ten years without unit tests, it starts at zero percent coverage. No management and no customer will pay to bring years of grown source code under full coverage after the fact. A 100 percent target makes no sense there, and going from seven to ten percent motivates nobody.

The better route goes through critical code. Define explicitly which parts of the code must be covered no matter what, for example business logic that calculates monetary values. Annotate that critical code, measure it and require 100 percent coverage there. That is independent of legacy code and of getters and setters, and it resolves the conflict between the two goals.

API Tests and Contract Tests Close the Gap Between Unit and UI

In architectures built on microservices and APIs, the next step is to secure the API level. An API test does not check the graphical interface but the programming interface: it sends API calls and checks the results.

Between unit tests and API tests you can add another layer: consumer-driven contract testing. Every API has a provider that offers it and one or many consumers that use it. So that both sides can evolve independently, they agree on a contract that defines which requests are allowed and which responses are expected.

The open source tool PACT records these contracts from the consumer side and stores them centrally. From the contracts it can generate mocks that simulate either the consumer or the provider. That lets you test against the contract on a developer machine or in the CI system, without a heavy setup with databases and Docker images.

The payoff is early feedback. You learn right away whether all consumers still work with a changed provider. That interaction is central in microservice architectures and in enterprise environments with an Enterprise Service Bus or Kafka, where many services exchange data.

This layering splits the responsibilities sensibly:

LevelWhat It Is For
Unit testFine-grained safety net, small isolated units, critical code
Consumer-driven contract testCompatibility between provider and consumer, early and without a full setup
API testSecuring the backend, checking variants and combinations
UI testEnd-to-end business flows and layout, in small numbers

Why the Test Pyramid Beats the Ice Cream Cone

Shift left also means building a test pyramid instead of the testing ice cream cone. The pyramid has many unit tests at the bottom and little UI automation at the top. The ice cream cone, an anti-pattern, flips that around: everything automated through the UI at the top, nothing underneath.

Modern user interfaces are mostly built with React, Vue or Angular, with Svelte on the rise. They fetch their data from the backend through HTTP, REST or GraphQL calls. Often a backend for frontend bundles these calls into a UI API. At that API level, one step below the interface, you can test combinations without clicking through a hundred variants in the UI. That is faster and breaks less often when things change.

The UI itself still needs testing, both for correct layout and for function. But UI automation is only a small building block. If you take the lower levels seriously, the top level can stick to what it is meant for: end-to-end tests of business flows.

Automated Delivery Needs the Whole Chain

Frequent releases are only manageable with a fully automated continuous integration, delivery and deployment pipeline. The ideal state that large providers have reached looks like this: at the push of a button the entire quality assurance runs, the result is deployed to production automatically, checked and monitored there again, and you know right away whether it works.

This level of automation brings new risks. CI systems have become a target for attacks because they give direct access to the code. Securing automated pipelines is a big topic of its own, with its own OWASP Top 10 list by now. No advantage comes without a potential downside, but the direction is clear.

Frequently Asked Questions

Why does it cause problems when each developer implements their own user story independently during a sprint?

This pattern creates a mini-waterfall within the sprint: All stories start at the same time, and the testers don’t check everything all at once until the last two days. By the end of the sprint, potentially nothing is finished, because while everything was started, nothing was completed. If, on the other hand, the team works according to prioritization, perhaps only three out of five stories are finished, but those are truly complete.

How can you tell if a user story is broken down sensibly?

A sensible breakdown is vertical: It encompasses all levels of a feature and delivers something functional at the end of the sprint that can be tested and used. A horizontal division creates dead ends, such as when the user interface of a feature is assigned to one story and the associated database isn’t developed until three sprints later. The benchmark is the value of the result, not the question of what one person can accomplish on their own.

What should you do if experienced employees don’t want to share their knowledge with the team?

This is usually driven by fear, not a lack of knowledge. Long-time experts, in particular, fear that sharing their knowledge will make them replaceable and view a cross-functional team as a threat. This argument can be refuted: A junior developer cannot catch up in three months to what someone has built up over thirty years. What’s needed is guidance over several iterations, which is actually the Scrum Master’s responsibility.

Is static analysis on the build server sufficient, or is quality feedback in the IDE necessary?

Feedback in the editor is superior because it eliminates an entire correction cycle. If the same analysis runs first on the build server, it triggers an additional loop of waiting, reviewing, criticism, and reworking. A plugin that reports a method’s cyclomatic complexity inline as you write immediately highlights the problem. Server-based tools like SonarQube were well-established in the enterprise environment by 2023, but they do not replace this immediate feedback.

Can poor code quality be compensated for with more testing?

No. Code quality is either there or it isn’t, and it cannot be retroactively “tested into” software, no matter how thoroughly it is tested. Those who ignore naming conventions, coding standards, or code duplication under delivery pressure may get by in the short term, but will pay the price later. The rule for day-to-day work: Anything that can be done with the help of tools and without a significant time investment is worth the effort.

Is 100 percent unit test coverage a reasonable goal?

No, a blanket requirement of 100 percent is counterproductive. A test for a setter method only verifies the assignment functionality of the programming language and is a waste. For legacy code that has evolved over ten years without unit tests, no one is going to pay for full coverage retroactively, and an increase from seven to ten percent doesn’t motivate anyone. A better approach: explicitly define and annotate critical code (such as financial business logic) and require 100 percent coverage there.

How do you test compatibility between services without setting up a complete integration environment?

Through Consumer-Driven Contract Testing. Providers and consumers agree on a contract specifying which requests are allowed and which responses are expected. These contracts can be used to automatically generate mocks that simulate the other side, allowing testing to take place right on the development machine or in the CI system without a database or Docker images. In the 2023 conversation, Alex Vukovic named PACT as the best-known open-source tool for this.

What risks does a fully automated deployment pipeline entail?

The high degree of automation makes the pipeline itself a target for attack, because CI systems allow direct access to the code. The security of automated pipelines was already a major topic in 2023, with its own OWASP Top 10 list. This does not change the overall direction: Frequent releases cannot be managed without end-to-end automated integration, delivery, and deployment.

Share this page