Skip to main content

Search...

Game Testing at Tibia: Exploratory Testing Before Checklists

Game testing for Tibia means weekly patches for a 25-year-old game with 500,000 players. Exploratory testing and big external tests keep it stable.

• • Updated: • 12 min read
Cover of the expert talk on 'Game Testing at Tibia: Exploratory Testing Before Checklists' with Oliver Heldt and Richard Seidl.

Game testing for an MMORPG means checking the quality of a game that has grown for decades and gets new patches every week. Exploratory testing sits alongside checklists, god spells replace tedious leveling, and twice a year up to 70,000 real players test new features in a separate test environment before the official update goes live.

Key Takeaways

  • In testing the MMORPG Tibia, exploratory testing matters more than simply checking against requirements documents, because many risks only show up between the lines of the design documents.
  • A bug that kills players unfairly is not a cosmetic problem: in Tibia, players lose progress worth a week of hard play, and that drives them away.
  • In the external test with up to 70,000 invited players, the critical mass of users uncovers bugs and balancing problems that the internal team could not foresee.
  • Without in-game test tools, the so-called god spells, it would be practically impossible to set up specific game states and test quests in the middle of their sequence without playing for hours first.
  • The release decision at Tibia rests not only on green checkmarks in checklists but also on the gut feeling of the testers who tested each feature.

Why Online Games Need Game Testing

Without game testing, an online game puts its players at real risk, and in the end that costs trust and money. In a massively multiplayer game like Tibia, which has been running since 1997, every bug weighs heavily, because players invest real time in their characters.

When a character dies because of a bug, the player pays a high price. Oliver Heldt puts it like this: you play toward something for a week, and then you die because there’s a bug in the code. Anyone who goes through that loses interest in the game.

So the case for testing is economic and emotional at the same time. A studio wants to keep its players and deliver a product with reliable quality. The argument “a game isn’t vital, so you don’t need to test it” falls short as soon as a community has been attached to a product for years.

What Makes a 25-Year-Old Game Testable, and What Doesn’t

The biggest challenge with an old game is the sheer amount of content that has built up and has to be considered with every change. Tibia gets a patch every week, plus interim releases and big updates such as a summer update.

An update like that brings new quests, new features and new game logic that all have to work alongside what already exists. The test team spends about seven weeks on a big update and splits up the work so that the required quality stays within reach in that time.

The central question is “how deeply do we test the old parts when something new comes in?” rather than “have we tested everything?” You can’t retest a game that has grown for over 25 years from scratch with every release. So you decide deliberately where the risk is and how far you go into it.

Why Exploratory Testing Is Key in Game Testing

Exploratory testing carries a lot of weight in game testing, because you can rarely read all the risks off a requirement. There are basic design and implementation documents that the team tests against, but that is only one side of it.

A single requirement leads to many exploratory tests. The real work is reading between the lines of the documents and finding out where a risk is hiding. Just checking whether something is red or green is not enough for a living online game.

Product knowledge helps, but it doesn’t replace intuition and fresh eyes. New colleagues bring a different perspective, and exactly that input uncovers things a well-established team overlooks.

Checklists as Inspiration, Not a Mandatory Program

In game testing, checklists serve mainly as inspiration, not as a program to work through. For clearly defined cases, such as a new outfit or a new mount, the team uses checklists built on experience that cover typical sources of bugs.

There are smoke test lists as well, but they have a different purpose. They are meant to prompt the question of how deeply a new feature should be tested. The goal is not to collect as many green checkmarks as possible at the end.

Once the checklist for a new outfit has run through cleanly, the test isn’t over. The checklist gives you a good foundation, nothing more. The rest depends on how much test depth you want.

A Good Bug Report Needs Repro Steps Your Grandma Understands

The most important part of a bug report is a reproducible step-by-step description that works without follow-up questions. Oliver Heldt is clear about his own standard.

“It would be cool if you could give it to your grandma and she’d roughly know how to reproduce it.”

(Oliver Heldt)

Repro steps like that can get long, because a bug sometimes involves several players performing specific actions with specific timing. “Press a button three times” won’t do as a description.

A report includes step-by-step lists and screenshots, videos less often. Videos tend to come from the players. The standard stays the same: the developer should need to ask as few questions as possible, and another tester should be able to reproduce the bug as well. “Internal error on the website” as the only information helps nobody.

God Spells: Why Testers Are Allowed to Cheat

Testers work with so-called god spells, built-in ways to intervene in the game that make testing practical in the first place. From a player’s point of view, this would be cheating. In testing, every one of these spells has a clear purpose.

They range from teleporting instead of walking long distances to setting a specific level and querying or setting quest states. That way you get a character into exactly the situation you want to test, quickly, and save a lot of time.

Without these tools, some tests would not be possible at all. In an area full of monsters, a tester would simply die, because they don’t have the level of experienced players. Some of those players have been playing Tibia for ten or twenty years. The point of testing is to examine a piece of game logic under control, not to outplay them.

On top of that, there are test characters and log files. A character with a huge number of items is good for a first run, because more can go wrong. Prepared characters are ready for targeted tests. Everything that helps judge a behavior as right or wrong belongs in the same toolbox of manipulation and information.

How 50,000 Players Become Part of the Test

After the internal test, real players come in, in an external test that runs for three to four weeks. Twice a year, between 50,000 and 70,000 players are invited, out of roughly 500,000 active players in total.

The external test has its own test server and its own website. Players don’t get a checklist. They are asked to play the way they normally would. From inside the game, they file a bug report with a right-click, which captures the location and a description of the bug directly.

Mostly small things come back: a typo, a spot on a house wall where something isn’t right, or a place where you get stuck. In other words, exactly the bugs that were left alone internally as minor. But the critical mass of players also uncovers things nobody had on their radar in the internal test.

The external test also serves balancing. Is a boss too strong, are there too many monsters, does new game logic make one of the four vocations superior? Metrics allow adjustments before the release, on points that are harder to judge internally.

Players Try Everything, so You Have to Think of Everything

Players actively look for ways to outsmart the system, and testing has to anticipate exactly those attempts. If it can be done, it will be done.

One example: when an NPC is about to take an item from them, players try throwing the item on the ground first so the NPC can’t collect it. Tricks like that are meant to get money or advantages they aren’t entitled to.

These departures from the real world are what tests fail or pass on. Often the honest reaction is: nobody thought of that. That’s part of the job.

When a Release Is Ready Enough

In game testing, “done” is a state of data and confidence, not complete coverage. There is a due date, and usually by then the quality has reached the desired level, with every feature tested as deeply as planned.

The team makes the decision together, but it relies heavily on experience. The team lead explicitly asks the testers who worked on a feature how they feel about it. That feeling counts alongside the green checkmarks and closed issues.

If the confidence isn’t there, there is a retest. Then the team has to be honest with developers and product managers and say it needs a few more days, because something was missed in the estimate. Since the testing is done internally for the studio’s own product, this step is possible without having to argue with an external customer.

Even After 25 Years, There Is Plenty of Technical Testing Left

An old game doesn’t stand still. New technology arrives regularly and brings its own test tasks. Tibia stays 2D, and there are no plans for a leap to new graphics. Even so, the technical scope keeps growing.

Examples include a dedicated app with game information and a full soundtrack for a game that used to have no sound at all. The soundtrack was a big test task, because sound doesn’t simply switch on and off.

Sound follows its own logic and has to fit the location, which sounds different on a coast than at a well. Players can control the sound in fine detail, for example muting only weapons or only other players’ spells. In testing, you then work a lot with log files and manipulation to judge whether the sound at a given spot is the right one, without waiting hours for the next piece of music.

Classic Bugs and What They Teach Us

Even carefully tested games don’t catch every bug, because you can’t test everything. That basic rule of testing shows up in concrete cases from production.

  • A whoopee cushion could be stacked. If a player stacked too many, a data type overflowed and the client or server crashed.
  • A cube let one player hit another over the head and then vanish, instead of being penalized as usual.
  • An event island that the team had worked on for weeks suddenly became unreachable after a legitimate change to its access. Nobody noticed until the external test phase.

These cases share a pattern: a sensible change in one place breaks something else that nobody had in mind when making the change. Areas touched late in the release can’t be fully retested, so something can slip through. The weekly patch rhythm softens the blow, because such bugs can be fixed quickly.

Frequently Asked Questions

What are the consequences of a bug that causes a player character to die?

The player loses progress that took them a real-world week to achieve. This isn’t just a cosmetic issue, it’s a reason to quit the game. In an MMORPG like Tibia, which has been running since 1997, players become attached to their characters over the course of years. That’s why the rationale for testing is both economic and emotional.

Is it enough to test a game against its design documents?

No. Basic design and implementation documents cover only one side of the story; a single requirement often requires many exploratory tests. The real work lies in reading between the lines and identifying risks that no one has written down. Product knowledge helps with this, but it doesn’t replace intuition: New colleagues see things that a well-established team might overlook.

Are checked-off checklists a sign of sufficient testing?

No. Checklists serve as inspiration, not as a mandatory program. For clearly defined cases, such as a new outfit or a new mount, there are experience-based lists of typical sources of errors, as well as smoke test lists. Even if such a list is completed without issues, the test isn’t over yet. It provides a foundation; the depth of testing is decided separately.

What should be included in a defect report so that developers have to ask as few questions as possible?

Reproducible step-by-step instructions, along with screenshots and, less commonly, videos. The standard: Even a completely uninvolved person should be able to roughly understand how to reproduce the bug. Some repro steps can be lengthy because multiple players must perform specific actions at specific times. A report like “internal error on the website” doesn’t help anyone.

What tools do testers need within a running game?

Built-in intervention options (called “god spells” in Tibia), such as teleporting instead of walking, setting a specific level, and querying and changing quest statuses. Without them, it would be impossible to deliberately recreate specific game states, and in an area full of monsters, a tester would simply die. In addition, there are pre-prepared test characters, such as one with a large number of items, as well as log files.

What does an open test with real players offer that internal testing does not?

The critical mass of users. In Tibia, twice a year, 50,000 to 70,000 of the roughly 500,000 active players are invited to a dedicated test server for three to four weeks. They play without a checklist and report bugs by right-clicking within the game. The test also provides metrics for balancing, such as whether certain bosses are too powerful.

Can a decision to release a feature be based on the testers’ gut feeling?

Yes. The team lead explicitly asks the testers who have thoroughly tested a feature for their gut feeling, and this feeling counts alongside checkmarks and closed issues. If there isn’t enough confidence, a retest follows. The team then openly tells developers and product managers that they need a few more days.

Why do bugs slip into a release despite careful testing?

Because not everything can be tested, and changes made late in the process can’t be fully retested. A well-intentioned change can then break something else: A change to the access path made an event island inaccessible, but this wasn’t noticed until external testing. The weekly patch cycle cushions such cases because fixes are delivered promptly.

Share this page