Skip to main content

Search...

Teaching Software Testing: Students Need Real Complexity

Software testing is hard to teach at university because small exercises hide the benefit. Complex projects, TDD and testable architecture make it click.

• • Updated: • 13 min read
Cover of the expert talk on 'Teaching Software Testing: Students Need Real Complexity' with Kai Renz and Richard Seidl.

Teaching software testing at university is hard because students lack real-world context: small, manageable exercises barely show what tests are good for. More complex projects with real infrastructure, techniques such as test-driven development and behavior-driven development, and a deliberate approach to AI-generated code build the understanding of testing and code quality that students need.

Key Takeaways

  • Students only see the point of testing once the assignments are complex enough. Simple exercises don’t create real motivation, because the benefit never becomes tangible.
  • Test-driven development works particularly well in teaching, because its red-green cycle gives beginners a clear, checkable structure they can apply right away.
  • Poor software architecture makes good testing structurally impossible: if domain logic and database logic are mixed, you can’t test business logic without a running database.
  • Code coverage metrics such as a 75 percent quality gate in SonarQube tempt students to write tests just to hit the number rather than to secure quality.
  • Anyone who uses AI-generated code without understanding it loses control over what runs in the system, which becomes a real problem in safety-critical domains such as banking or air traffic control.

Why Teaching Software Testing Is So Hard

Students often don’t see the point of tests, and that is the core problem in teaching software testing: the purpose doesn’t come across as long as the assignments stay trivial. If you write a simple function, have thought about it for a long time and know it works, you see no reason to test it as well.

Kai Renz, who has been a professor of software engineering at Darmstadt University of Applied Sciences for eight years, describes this as a double hurdle. At the start, many students can’t really program yet. Writing tests then feels doubly abstract, because the foundation is missing. Once that first hurdle is cleared, the next one follows: small, manageable examples don’t give a convincing reason to write tests.

The classic argument doesn’t land at this stage. Telling students that tests will later protect them from regressions and from uncertainty when the software evolves in the real industry stays anecdotal. It is a story about a future that doesn’t exist yet for learners.

Generative AI makes this worse. If ChatGPT produces the correct solution anyway, the question of why you’d write a test becomes even more pressing.

Complex Projects Provide the Reason Trivial Examples Can’t

The way to overcome the lack of motivation is complexity. As soon as a project becomes realistic enough, the need for tests appears on its own.

In Darmstadt, a software engineering lab course simulates a pizza shop for exactly this purpose. Each group runs its own Kubernetes cluster on real infrastructure, which makes integration tests necessary as well. The students are deliberately thrown in at the deep end.

The result is mixed. For some, it clicks, and they understand why they are testing. Others mainly ask what they need to do to pass the course. For them, the spark doesn’t catch. That split isn’t a failure of the method. It is the reality of a diverse student body.

Conceptual Understanding Beats Tool Knowledge

The most important thing students take away is conceptual understanding, not mastery of a particular tool. The industry moves so fast that any specific tool is soon out of date.

Specific techniques and tools are still taught, from unit tests through integration and end-to-end tests to UI tests. The focus is on how a test is structured: applying the AAA pattern consistently and deciding clearly what the system under test actually is. A single class? A complete feature? Which level am I working at right now?

Experienced developers make exactly these decisions almost automatically. Do I need a test database? Do I mock a framework or not? If you know how the components interact, you decide deliberately. If you don’t, you need clear guidance to get started.

Test-Driven Development Works Well in Teaching

Test-driven development follows a simple, easy-to-grasp cycle that students can follow well in the lecture hall. First the test, then the functionality. The test starts out red, and at first maybe not everything even compiles. Then comes the code.

The most common question from students is: how do I know what to test if the function doesn’t exist yet? That leads to half-philosophical discussions about what you need to know before you can do something.

The answer is pragmatic: if you know what you would program next, you also know what should come out. And then you know what the test looks like. TDD helps here because it narrows the focus to one small step toward the functionality.

Pair programming with a driver and a navigator reinforces this, as do coding sessions in the lecture hall run on mob programming principles. One person codes, everyone watches. It works, but it remains demanding, because very strong and very inexperienced students sit in the same room.

Behavior-Driven Development Brings the User’s Perspective into Testing

Behavior-driven development shifts the focus away from the implementation toward the application and the user. Instead of asking how something is implemented, BDD asks what should be tested and how a test case can be described.

This technique is used in an elective called “Professional Testing”. Over fourteen weeks, participants work through the whole spectrum, from backend tests across different programming languages and tools to UI tests with Playwright. The learning curve is steep, and the knowledge they take away is broad.

The effect shows up on the job. Former students report that they could put what they learned to use right away. Some arrive at companies where testing isn’t yet standard practice and know more about it than the colleagues already there. Fresh graduates are often the only way to bring quality thinking into a company that has managed without it for twenty years.

How Software Architecture Determines Testability

Good testability depends heavily on the software architecture. Systems that are hard to test usually mix things that belong apart.

The clearest example is separating domain logic from controller logic. Approaches such as hexagonal or functional architecture only show their strength when you apply them consistently. The price is visible overhead: lots of helper classes that copy objects from the database layer into the business logic.

This is exactly where students often lose interest. They already have a class that is coupled to the database and ask why that should be a problem. The answer becomes obvious as soon as you want to test the business logic: if you first have to spin up a database to test the business logic, the two aren’t cleanly separated.

What to Do When an Existing System Can’t Be Tested

The typical trap is paralysis. The team faces a system that is hard to test and sees nowhere to start. It needs refactoring, but refactoring is hardly possible without tests as a safety net, and the tests are hardly possible because of the architecture. A genuine vicious circle.

The pragmatic way in is through what’s new. New functionality gets built and tested cleanly and consistently. That lets the team experience what a good architecture brings: easier refactoring, tests at every level, static analysis with no open findings.

Over time, that understanding seeps into the legacy code. With very large, old systems, it never gets implemented completely. And even in new development, there is a risk of repeating the same mistakes. That is why the team needs a shared way of thinking about quality.

Why Code Coverage Shouldn’t Be a Goal in Itself

Code coverage of 75 percent is not an end in itself. Teams that write tests only because a red quality gate needs to turn green miss exactly that.

In the lab course, SonarQube serves as the static code analysis tool, with a quality gate that requires more than 75 percent coverage. If teams don’t reach it at first, some simply write any tests at all until the number fits. That pushes the mindset in the wrong direction.

The principle is: you don’t write a test to turn a number green, you write it to improve quality. What the tool says about code quality, complexity for example, is more interesting than coverage alone anyway.

Non-Functional Testing Is Deliberately Taught Elsewhere

Security, usability and performance hardly fit into the software engineering modules, because the curriculum there is already packed. Darmstadt handles them in specialized courses.

IT security has its own course and a strong research group behind it, on both the technical and the usability side. For usability, there is also a separate course on human-computer interaction, which asks how functionality can be implemented in a meaningful way at all.

Closer integration of these modules would be desirable, but two obstacles stand in the way. Lecturers like to work autonomously, and coordination takes effort. And students would face a much bigger workload if one module required another.

Performance testing is only placed within the agile testing quadrants, nothing more. A load test with 500 pizza orders has nothing to do with a real system and would remain fake. Exercises should have real-world impact instead of simulating tasks that mean nothing outside the classroom.

AI at University: Responsibility Stays with the Student

For today’s students, generative AI is a given. Teaching responds with clear limits where they matter and openness everywhere else.

The introductory programming course ends with a practical exam on a computer without internet access. Anyone who only had the solutions generated beforehand won’t have the basic programming skills the exam requires. In the lab courses, on the other hand, the rule is: AI may be used, but responsibility for what is submitted stays with the students. They have to be able to explain what they did.

One incident shows the risk clearly. While developing a new feature for the pizza project, the team found that the functionality they wanted was already in the code. The commit history contained a commit that had replaced almost the entire project. Under deadline pressure, a student had given the code to ChatGPT and asked it to find bugs. The AI added the functionality without being asked.

“If you worked at a bank, in air traffic control or for a brake manufacturer: could you sleep well if you had committed code you didn’t even know was there?”

(Kai Renz)

The student owned up honestly, and that is what made the conversation so valuable. The lesson: AI helps, but you have to know what you’re doing yourself.

The Skills That Last

Being able to use ChatGPT isn’t a skill that convinces an employer. Students who openly solve their assignments purely with AI hear that message loud and clear.

In theory, you could get a degree with AI without understanding much, as long as you somehow get through the exams. But prompting on its own is a short-term skill that will probably get simpler again as talking to AI becomes more natural.

Two abilities last. The first is communicating with people: really grasping what the other person wants. That is hard to delegate. The second is keeping control over what the AI produces. Is it plausible? Do I understand it? Do I still have the big picture?

But you only get the big picture with experience. Without experience, there is no big picture. That is the real dilemma of a generation that starts with AI before it has built the judgment to check the results.

AI Doesn’t Understand the Context on Its Own

Some constraints have to be given to the AI explicitly, because it won’t take them into account by itself. Logging is a good example.

Logging frameworks and how to write good log messages can be taught quickly, and the benefit is obvious. Then comes the follow-up question: are you even allowed to log everything? If someone asks for their data to be deleted, but the system has logged which pages that person viewed, will you ever get that out of the log again?

Limits like these come from data protection and copyright law. So does the question of where a company’s source code is sent at all and which server is reading along. An AI doesn’t grasp these implications automatically. You have to spell them out.

Frequently Asked Questions

Is it enough to practice testing on small programming tasks?

No. With trivial tasks, the benefit remains imperceptible: Anyone who writes a short function, has thought long and hard about it, and knows that it works sees no reason to perform testing. The reference to future regressions in the industry remains a story about a future that doesn’t yet exist for learners. Only realistic complexity creates the need on its own.

What are the benefits of working with real infrastructure in educational projects?

It creates the opportunity for test types that otherwise wouldn’t be justified. In a software engineering lab course, groups simulate a pizza shop and each runs their own Kubernetes cluster for it, making integration testing necessary. The response is mixed: some students are enthusiastic about the approach, while others mainly ask what’s required to pass the course.

How do you address the objection that you can’t perform testing on something that hasn’t been programmed yet?

Pragmatically: If you know what you’d program next, you also know what the output should be, and thus what the test looks like. This is precisely the most common question beginners ask about test-driven development. The technique helps because it narrows the focus to a single small step toward the desired functionality.

How can you tell if an architecture is hindering testing?

A clear sign: You first have to create a database to test business logic. In that case, domain logic and database logic aren’t cleanly separated. Approaches like hexagonal or functional architecture only demonstrate their strength when implemented consistently, and they incur visible overhead: many helper classes that copy objects from the database layer into the business logic.

Is 75 percent test coverage a reasonable goal?

Not as a number alone. A quality gate that requires coverage above 75 percent tempts developers to write any tests necessary until the number is right. This shifts the mindset: a test is written to improve quality, not just to turn a number green. Metrics on code quality, such as complexity, are more meaningful than coverage alone.

How risky is it to adopt AI-generated code without understanding it?

You lose control over what’s running in the system. In one student project, a commit was found in the commit history that replaced nearly the entire project: Under pressure to meet a deadline, a student had given the code to ChatGPT for debugging, and the AI added functionality without being asked. In fields such as banking, air traffic control, or brake manufacturing, something like this immediately becomes a problem.

Which skills remain valuable when AI writes a large portion of the code?

Two. First, communication with people, that is, truly understanding what the other person wants. That’s difficult to delegate. Second, control over the output: Is it plausible? Do I understand it? Am I keeping track of everything? Prompting alone is a short-term skill that will likely become simpler as communicating with AI becomes more natural.

Yes, it doesn’t take such implications into account on its own. Logging illustrates this well: The technical structure of good log messages is easy to explain, but the real question is whether everything is allowed to be logged. If someone requests the deletion of their data while there is a log of which pages they viewed, that becomes difficult. This also includes where a company’s source code is sent.

Share this page