Skip to main content

Search...

Speech Recognition Testing: Every Fix Means Retraining

Speech recognition testing means testing a black box: every fix costs days of retraining. Golden test sets and generated variations find errors early.

• • Updated: • 14 min read
Cover of the expert talk on 'Speech Recognition Testing: Every Fix Means Retraining' with Olaf Thiele and Richard Seidl.

Speech recognition testing, and testing audio AI in general, means checking speech-to-text and text-to-speech models for errors without being able to see inside the model. Every fix requires retraining that takes several days and costs real money. Golden test sets and AI-generated test variations are the main ways to uncover recognition and pronunciation errors systematically.

Key Takeaways

  • Audio AI models can’t be debugged from the inside: finding a misrecognized word means retraining the whole model, which takes several days and costs anywhere from a few hundred to several thousand euros.
  • Golden test sets are currently the most important quality assurance method for audio AI, because classic test automation runs into the black-box nature of AI models.
  • The data gap hits German especially hard: of the 600,000 annotated hours of audio behind the Whisper model, only about 2,000 to 3,000 are German, which puts a structural cap on recognition quality for German speakers.
  • A lack of diversity in the training data goes straight into model quality: train on 90 percent male voices and you get a model that recognizes women poorly and makes synthetic female voices sound male.
  • ChatGPT can generate synthetic test data for audio AI, because it produces variations in wording that people simply don’t think of when they write test cases by hand.

Speech Recognition Testing Means Testing a Black Box

The hardest part of speech recognition testing is that nobody can really look inside the model. A model for transcription or speech synthesis is a black box. You can’t step through its logic line by line the way you would with a classic function. All you see is what goes in and what comes out.

With a classic function, you find a rounding error, fix the decimal places, and the test turns green. An AI model has no such point where you can step in. If it mispronounces a word or fails to recognize it, you can’t just change a rule. You have to retrain the model, and depending on its size, that takes several days and costs anywhere from a few hundred to several thousand euros.

Olaf Thiele has worked with German-language audio AI for years, covering both speech-to-text and text-to-speech. His assessment: the models work these days, but when it comes to testing them, the industry is only getting started.

Why AI Training Can’t Be Reliably Reproduced

At present, you can’t reliably reproduce an AI training run, because the hardware alone changes the result. With the common libraries PyTorch and TensorFlow, the same dataset can produce slightly different results on different chip architectures, even when every other variable stays the same.

That becomes a problem as soon as someone needs to reproduce a result exactly. Most training runs on hyperscalers such as AWS, Google Cloud, Azure or OVH. There, you don’t know whether you’ll get the same physical machine next week, say a V100 or an A100. Without that, the basic conditions for an identical result simply aren’t there.

For a long time, nobody even asked. As long as the goal was just to get a model running, exact reproducibility wasn’t an issue. It’s only now, as customers grow larger and demand governance, that it has moved to the top of the list.

When to Stop Training Is a Gut Call

There is no formula that tells you when to stop training. A model goes through several stages and improves on its internal test set with each one. That’s the trap: the longer you train, the better the model fits the training data, but at some point it gets worse at generalizing to new, unseen data.

The learning curve flattens out logarithmically. Whether you stop at step nine, ten, eleven or twelve often comes down to gut feeling. On top of that, the data isn’t evenly spread. A male speaker may be recognized better and better as training goes on while a female voice gets worse at the same time. There is no clear-cut answer for where to stop.

The “test set” used during training has nothing to do with testing in the software engineering sense. It is a small amount of data the model uses to check itself internally, not proof of quality in the classic sense.

What a Golden Test Set Is and Why It Helps

The most common tool in audio AI testing is a golden test set: a carefully curated collection of examples against which you can measure the model’s quality. Ideally, the customer builds this set, because they know best what they want to achieve. In practice, that rarely happens.

A good golden test set covers gradations. For dialects, that could mean very strong Bavarian next to mild Bavarian. With a few thousand samples, you can then watch how the model changes from one training run to the next.

The test set also helps soften errors you’ve found. If a model doesn’t recognize a certain word, you feed it more examples of that word. That is no real fix in the sense of a targeted bug fix, more a way of steering through the amount of data.

The AI Doesn’t Know What It Doesn’t Know

The core problem in finding errors is that an AI has no point at which it says, I don’t know. It produces a result even when it has no real answer. In audio recognition, that means you often don’t know which words a model drops or swallows.

Loanwords are a good example. Some German speakers pronounce “Budget” the hard German way, others the French way. To teach a model to recognize both, you’d need thousands of examples of exactly that variant. Where would you get them?

The same gap affects testing the output. If you want to check synthesized speech automatically, you end up going in circles: you’d have to run the generated audio through a recognition model again. If that model was trained on the same data, it will miss exactly what it would have missed in the original. The underlying problem just repeats itself.

Gold In, Gold Out: Data Quality Decides the Outcome

A model can only give back what went into its training data. That rule shapes every application. If you want to synthesize a young female voice but most of your input is audio of older men, you’re fighting the data.

This is where a bias creeps in that runs all the way through to the result. The freely available Mozilla speech data consists of about 90 percent male voices. As a result, women’s voices are recognized less accurately, and synthetic female voices sometimes sound male.

The effect gets stronger when a model is over-optimized for a small number of speakers. Train it on 300 hours of one well-known voice and you get a model that understands that person perfectly, while other speakers, such as an older woman, are barely recognized.

“What you put into a model like this is what you get out of it. That’s why I’m very cautious with applications that are supposed to work for a diverse range of people.”

(Olaf Thiele)

Why German Suffers from the Data Gap

There is far less usable training data for German than for English. One well-known model was trained on 600,000 hours of annotated audio, around 550,000 of them in English. German got only about 2,000 to 3,000 hours.

A German model on par with English would need a comparable base, somewhere around 50,000 hours of German audio. Right now it’s mainly the hyperscalers that collect data at that scale, and their models are getting measurably better.

Freely available sources are scarce. Mozilla’s project for collecting speech has wound down. Many obvious sources fail on licensing: sessions of the German Bundestag, radio broadcasts, audiobooks recorded from Project Gutenberg texts. Audiobooks have a second problem too, because a reading voice doesn’t sound like natural conversation.

A positive counterexample is the openly licensed German voice of Thorsten, who recorded around 20 hours with a good microphone. We’d need more open datasets like that, including from women and from different regions.

Dialects Are in Demand but Hard to Deliver

Models that handle dialects almost always fail for lack of data. Dialects vary a great deal by region, sometimes from one valley to the next. In Switzerland, someone is building a separate model for each valley with public funding.

In Germany, development is driven more by commercial interests, and the data is clearly lopsided. There tends to be more data from the south and less from the east. As a result, people from eastern Germany are recognized less reliably than people from the south. A Saxon model would need around 1,000 hours of Saxon speech, and that data just doesn’t exist.

How to Test a Voice Chatbot: Several AIs, Several Test Problems

A voice skill isn’t built on one AI but on at least three. A pizza order is a good way to separate the stages and see how testable each one is.

ComponentTaskTestability
Speech-to-TextTurns audio into textHard, because it stays unclear who said what and how
Natural Language UnderstandingRecognizes the intent in the textEasy to test: text in, intent out
Text-to-SpeechTurns text into audioHard, because automated checks need another AI

The middle component, natural language understanding, can be tested the classic way. You feed in a text and check whether the right intent comes out. You write these tests as usual, even if platforms such as Alexa don’t offer them out of the box.

The two audio ends remain the problem. On the recognition side, you only see the recognized text. You can’t tell whether an error came from a dialect, from mumbling or from a poor microphone. On the output side, the question is again how to check hundreds of synthesized answers without having a hundred people listen to them.

Using Generative AI to Test Speech Variability

Language models such as ChatGPT help most directly wherever testers need a lot of text variations. If you want to test the responses of a voice skill, you normally have to write every variant by hand, and after a few hours you simply run out of ideas.

A language model gives you 50 or 100 variants of a response on request. The temperature setting controls how freely it phrases them. These texts are then synthesized into audio, and testers listen to a sample. This way you catch errors that used to slip through, such as a wrong pronunciation or two words that come out wrong together because the synthesis works with phonemes.

The same approach works on the input side. Ask for 100 ways to order a pizza, and you get sentences you would never have come up with yourself. Each extra sentence in the test dataset costs almost nothing and widens the range you cover considerably.

One limit remains: for sensitive or high-risk use cases, the error rate of generative models is not acceptable. For internal and preparatory work, though, the approach works well.

MLOps and Standards Are Coming, but Audio Lags Behind

MLOps tools are picking up speed, but so far they are built for text, and a lot is missing for audio. The reason is the sheer amount of data. Ten seconds of audio are cut into 20-millisecond windows, the full frequency band is analyzed for each window, and a sliding window pulls in the neighboring windows as well. The result is huge vectors. A typical training dataset easily reaches 400 to 500 gigabytes.

At that size, it’s hard to keep track of data cleanly across training runs. If you change the mix, for example more women or fewer fast-talking men, you’d need to track exactly which data went into which run. That is exactly the tracking today’s tools don’t provide.

Hugging Face works like a GitHub for AI models and hosts tens of thousands of them. Its library lets you pull, say, 20 percent of a dataset, but it doesn’t hand over the matching metadata. The only thing required so far is technical run data, meaning what it takes for a model to start at all. Which data went in and which parameters were used for training stays unknown.

Standards would help the field, mainly because they would make models comparable. Customers today run whole collections of models, a so-called model zoo. Quality assurance and governance across many models built in different ways are hardly possible as long as everyone brings their own conventions. Whoever wants to set a standard would be in the best position with the reach of a platform like Hugging Face.

Frequently Asked Questions

Why can’t a detected error in an AI language model simply be corrected?

An AI model doesn’t have a point of intervention like a traditional function. You can fix a rounding error by adjusting the decimal places, but a misidentified or mispronounced word cannot be rewritten using a rule. The only real correction is retraining, which in 2023 took several days depending on the model size and cost several hundred to several thousand euros. In practice, adjustments are made instead by adding additional training examples.

Can the result of AI training be precisely verified?

There is no guarantee of reliability. Even the hardware affects the result: With the PyTorch and TensorFlow libraries, the same dataset produced slightly different results in 2023 on different computer architectures. Furthermore, anyone training on hyperscalers’ infrastructure doesn’t know whether they’ll be assigned the same physical machine (such as a V100 or A100) a week later. The issue of reproducibility only came to the forefront when customers began demanding governance.

What makes a good golden test set for speech models?

Examples with clear gradations. For dialects, a very strong Bavarian accent should be included alongside a mild Bavarian accent, so it becomes clear at what point the model starts to fail. With a few thousand samples, you can observe how the quality changes over multiple training runs. Ideally, the customer compiles the golden test set themselves, because they know best what they want to achieve.

Why do speech recognition models often have a harder time recognizing women’s voices?

Because the training data is predominantly male. In 2023, the freely available Mozilla speech data consisted of about 90 percent male voices. The result: women’s voices are recognized less accurately, and synthesized female voices sometimes sound masculine. This effect is exacerbated by over-optimization. If you train the model for 300 hours using a single prominent voice, you end up with a model that perfectly understands that person, while, for example, an older woman is barely recognized at all.

How much training data is missing for German-language audio models?

A well-known model was trained on 600,000 annotated hours of audio, of which about 550,000 were in English and only about 2,000 to 3,000 in German. To achieve a level of German comparable to that of English, roughly 50,000 hours would be needed. The availability of free sources was low in 2023: the Mozilla collection project had ended, and Bundestag sessions, radio recordings, and audiobooks faced licensing issues.

Can voice assistants even be tested using traditional methods?

To some extent. A voice skill consists of at least three AI components, and only the middle one is easy to test: With natural language understanding, you input text and check whether the correct intent is recognized. The two audio-related components remain challenging. With recognition, you only see the transcribed text and can’t determine whether a dialect, mumbling, or a poor-quality microphone caused the error.

Can a language model be used to generate test cases for voice applications?

For text variations, yes. A language model can generate 50 or 100 different phrasings of a response or a pizza order on command; the “temperature” setting controls the flexibility of the output. These texts are synthesized as audio and listened to on a random basis. This reveals errors that previously slipped through, such as word combinations that the phoneme-based synthesis mispronounces. For sensitive or dangerous use cases, the error rate is too high.

Is it possible to trace which data was fed into an audio model?

Hardly. The tools commonly used in 2023 did not track exactly which data was included in which training run; only technical run data was required, that is, the minimum necessary for a model to start. The Hugging Face library could extract about 20 percent of a dataset but did not provide the associated metadata. The sheer size of the data adds to the difficulty: audio training datasets can easily reach 400 to 500 gigabytes.

Share this page