Fuzzing, also called fuzz testing, means testing software with random or deliberately manipulated input to expose unexpected behavior such as crashes or security holes. Classic fuzzing knows nothing about the system under test. Custom fuzzing builds in knowledge of input formats and protocols, so it can generate millions of test inputs automatically that are syntactically valid but push the content to its limits.
Key Takeaways
- Fuzzing without format knowledge gets stuck at validation: a random IBAN gets its check digits right in only about one case in 100, and the country code has to be valid too. The fuzzer spends a whole day producing a few valid IBANs and never reaches the actual program logic.
- Custom fuzzing combines known syntax rules with controlled random variation, so only selected fields such as the transfer amount are left open while everything else stays valid and gets processed.
- Every program that processes external data has to cope with invalid, negative and mathematically extreme input such as NaN or negative amounts, because values like these can spread through downstream systems like a virus.
- LLMs are a poor choice as the main source of test data: they learn from existing patterns and therefore miss exactly the unknown edge cases that expose bugs.
- The official test suite for the German e-invoice format XRechnung contains only a few dozen cases for around 150 data fields, which makes automated test generation a necessity, not an option.
What Is Fuzzing?
Fuzzing means bombarding a program with invalid, unexpected or completely random input to see what happens. Does it crash? Does it hang? Can you throw it off balance? The technique goes by several names, fuzz testing, fuzzing tests or, in some search queries, fuzzy testing, but the idea is always the same.
The term goes back to Bart Miller, an American professor, in the late 1980s. Miller worked on his mainframe from home, over a plain telephone line and a modem. Wisconsin gets a lot of thunderstorms, and the electrical discharges interfered with the connection, so the bits and bytes arrived at the mainframe altered.
He noticed that many programs simply couldn’t handle this garbled input. So he turned it into a programming assignment for his students: generate random, meaningless input and use it to test the common programs of the day. The results were sobering. Most UNIX programs of that era could be knocked over this way.
The paper was initially turned down by the reviewers. Their argument: why would anyone ever send input like that to a program? In the late 1980s, Miller knew every user of his mainframe personally. Only with the internet and anonymous users did it become clear how real the danger was.
Every Program Needs to Withstand Malicious Input
A program that processes external data cannot assume that the data is well formed. That is the central lesson of fuzzing.
An attacker can poke at a system remotely with random data. If the system stops responding at some point, there may be a weakness there, one that can be widened step by step until the system can be manipulated. That is why fuzzing is now standard practice for any software that processes data from third parties.
So the topic sits where security and robustness meet. It is about security holes, and it is just as much about keeping a system stable when the input is unexpected.
How a Single Input Can Break a System
Small changes to valid data can have a big effect. A bank transfer makes this easy to see.
A normal transfer reads: from Andreas to Richie, one euro. In 99.9 percent of cases, that’s exactly what happens. Things get interesting when you change specific values:
- Negative values: A transfer of minus one euro. If a script somewhere forgot that amounts can be negative, you have a hole.
- Infinity: An amount of infinitely many euros. If that gets through, the bank has a problem too.
- Not a Number (NaN): A special floating-point value. Anything you calculate with NaN turns into NaN. A single accepted NaN transfer could infect the daily total, the balance sheet and the share price.
- Buffer overflow: An email address with 100,000 characters instead of the usual 256. If the target buffer overflows, the extra characters can be used to overwrite critical data.
- Command injection: Instead of a name, the field contains a smuggled-in command that deletes the database on the other end.
Banks are protected against attacks like these. The question is whether the same holds for every fintech startup and every piece of code an intern wrote. Fuzzing starts with random data and prods here and there until the system shows a weakness.
“Any program that processes data from outside cannot assume that this data is well formed. It has to cope with faulty, random and malicious data.”
(Andreas Zeller)
Fuzzing Runs for Hours, Not Minutes
Nobody fuzzes by hand. A fuzzer runs automatically over long periods, typically 24 hours, often days, sometimes weeks, and generates hundreds of thousands to millions of attempts along the way.
The advantage of pure randomness is that the fuzzer needs no prior knowledge and has no bias. It tries one input after another and also finds bugs that no known attack catalog lists. That is why the odds of a hit are good, even if a single run takes a long time.
Why does logging matter so much? Because the classic beginner’s mistake is not recording anything. If you send random data for a week and then can’t tell which input caused the crash, you have nothing to show for it. Log every input.
A protected environment matters just as much. Injected commands can do anything, up to deleting your own file system or the very database that holds your test data. Never fuzz a production system.
Time Bombs Are Especially Hard to Find
Not every bug shows up right away. A malicious value can be slipped in early and only cause trouble many interactions later. This separation of cause and effect is called a time bomb.
Bugs like these are particularly hard to debug because the trigger and the visible effect lie far apart. For an attacker, that is a nasty tool: plant the data, then wait until it does damage weeks later.
That is why good fuzzing records the entire sequence, not just the last interaction. Once you have a reproducible failing sequence, automated techniques can shrink it. Out of a thousand interactions, they filter out the two or three that actually cause the failure.
Why Pure Randomness Fails with Structured Formats Like the IBAN
Random data doesn’t get far with complex input formats. A bank rejects arbitrary bytes immediately, because a transfer has a strict structure: a well-formed document, checksums, valid fields.
The IBAN puts numbers on the problem. Random data hits the two check digits after the country code in only about one case in 100. On top of that, the two-letter country code has to be valid, out of 26 times 26 possible combinations. So random input produces a valid IBAN only in a tiny fraction of attempts.
And a transfer needs two valid account numbers, yours and mine. The probabilities multiply. A fuzzer spends a whole day just producing a handful of valid IBANs. It never gets to the actual logic, such as the checks on the amount field.
The point is that data packets are built to withstand simple disturbances. The same mechanisms that protect a bank transfer on a line full of thunderstorm noise also stop a naive fuzzer.
Custom Fuzzing: Randomness Plus Knowledge of the Format
Custom fuzzing combines the randomness of fuzzing with conventional testing techniques that rely on knowledge of the system. Instead of generating everything at random, you teach the fuzzer the rules of the input format.
Tell the generator how an IBAN is built: two letters, two check digits, then the account number. From then on it only produces data that follows this pattern, and the share of syntactically valid input jumps.
The key step is to generate data that is syntactically valid and still nonsense in content, like an English sentence that is grammatically perfect and means nothing. The data packet has to be accepted first, otherwise a manipulated value like minus one euro never reaches the deeper processing.
This gives you precise control. You build the transfer by the rules, with sender, recipient, account numbers and amount, and then drop all checks in one or two places. For the amount, say, you run arbitrary values through. The result is a million transfers in one minute, nearly all of them valid, and a far more efficient test than pure randomness.
Fuzzing in Security vs. Fuzzing in Software Testing
Security and software testing use fuzzing for different goals, and that shapes the method.
In security, you often don’t want to assume any knowledge of the system you are attacking. Breaking in with insider knowledge counts for less. So security sticks to the classic approach: throw random data at everything, with as little prior knowledge as possible.
When you test your own software, it’s the other way around. You know the system and want it to hold up under bad input, so you put your knowledge of the format and of likely weak spots to work. If the IBAN check is known to be solid, you skip it and point the fuzzer at a spot that has just changed.
This is where custom fuzzing pays off. Existing test cases are tuned to known problems and rarely find new ones. A generator that targets a new feature quickly puts its finger on the sore spot.
Complex Data Formats Can’t Be Tested Without Generation
Modern data formats have grown more complex than anyone can test by hand. Germany’s electronic invoice format, XRechnung, has been mandatory since 2025 and is a good example.
XRechnung has around 150 data fields, including constructs such as sales outside the EU. The official test suite contains only a few dozen test cases. With so few handwritten invoices, there is no way to cover 150 fields with their extreme and negative values.
Nobody types in that many cases by hand. Testing at this scale only works with automated test data generation.
Then there is the insecure channel. E-invoices arrive by email, and anyone can claim to be someone else. Where a person used to read the paper invoice and only a limited range of variants was possible, brand-new, barely proven implementations now face an open door.
Open Problems: Coverage, Oracles and the Limits of LLMs
Fuzzing doesn’t solve everything. Three problems remain.
Coverage: Many special cases live in the processing, not in the data format. Whether a recipient is on a blocklist can’t be derived from the input, because there is no field for it. A good fuzzer detects which branches in the program haven’t been executed yet and focuses on them. Even good fuzzers reach only part of all branches, though.
The oracle problem: Did the system do the right thing? Common fuzzers only check for generic failures such as crashes or hangs. Many ignore error messages, because security people are looking for the uncontrolled crash. As a tester, you want to know whether the error message is appropriate and whether the system does the right thing from a business point of view. Building good checks for that is still hard.
The limits of LLMs: Machine learning learns from existing data and produces more of the same. That works for job applications, code and images, things other people have made before. With test data, you often want the exact opposite. Ask an LLM to turn 100 existing transfers into 100 new ones, and the new ones look like the old ones. You might as well test with the training data.
For individual data fields, such as generating valid IBANs, LLMs are quite useful. For the creative task of producing something nobody has tested yet, they are about the worst tool you could pick. That job still belongs to a creative tester who finds the bugs before an outside attacker does.
Frequently Asked Questions
Where does the term “fuzzing” come from?
The term was coined by American professor Bart Miller in the late 1980s. Miller was working on a mainframe via a modem and a telephone line, and thunderstorms were corrupting the transmitted bits. Many programs couldn’t handle the corrupted input. He turned this into a student assignment: Most UNIX programs at the time could be thrown off by random data.
Why are negative numbers or NaN values in input data so dangerous?
Such values spread like a virus. Anything that is calculated with a NaN becomes a NaN itself: A single accepted NaN transfer can infect the daily total, the balance sheet, and the stock price. Negative numbers affect scripts that were never designed to handle signed values. Infinite values, excessively long fields (such as an email address with 100,000 characters instead of 256), and smuggled-in commands have a similar effect.
How long does a fuzzing test typically run?
A fuzzer typically runs for 24 hours, often days, sometimes weeks, generating hundreds of thousands to millions of attempts in the process. This cannot be replicated manually. The advantage of pure randomness: The fuzzer requires no prior knowledge and has no bias, so it can also find flaws beyond known attack catalogs.
Is it okay to run fuzzing on a production system?
No. Injected commands can trigger any behavior, including deleting the system’s own file system or the very database where the test data is stored. A protected environment is required. This includes logging: If you send a week’s worth of random data and, when the system crashes, no longer know the input that triggered it, you’re left empty-handed.
How many attempts does a fuzzer without format knowledge need to generate a valid bank transfer?
A great many, because a bank transfer requires two valid account numbers and the probabilities multiply. The two check digits of an IBAN alone match random data in only about one out of 100 cases, and the two-letter country code has to be valid as well, out of 26 times 26 possible combinations. The fuzzer spends an entire day producing a few valid IBANs and never gets to the actual logic, such as checking the amount field.
Do security professionals fuzz differently than software testers?
Yes, because their goals differ. In security, the goal is to assume as little prior knowledge as possible about the targeted system, since a breach involving insider knowledge carries less weight; in that context, pure chance remains the method. When testing one’s own software, however, one actively uses knowledge of formats and vulnerabilities: if the IBAN validation works reliably, one targets the fuzzer at a field that has just been changed.
Is the official e-invoice test suite sufficient for a robust test?
No. The X-Rechnung standard recognizes around 150 data fields, including constructs such as sales outside the EU, while the official test suite comprises only a few dozen test cases. Outliers and negative values across 150 fields cannot be covered with manually written invoices. The format has been mandatory since 2025, and invoices are sent via email, that is, through an insecure channel.
Can large language models generate the test data for such checks?
Only to a limited extent. Machine learning trains on existing data and generates similar results. If you ask an LLM to create 100 new bank transfers based on 100 existing ones, the new ones will look just like the old ones; in that case, you might as well use the training data for testing. LLMs are useful for individual data fields, such as valid IBANs. For unknown borderline cases, a creative tester is still needed.


