Differential testing, online
Differential testing feeds the same generated inputs to two implementations that are supposed to agree and treats any disagreement between their outputs as evidence of a bug in one of them.
JavaScript, TypeScript and Python run in your browser. Nothing to install, no account.
Why it works without an oracle
Most testing needs an oracle: to assert that a function is correct you must already know the right answer. Differential testing does not. It needs only a second implementation that must produce the same answer. That suits cases where writing expected values by hand is impractical: a numeric routine, a parser, a compiler pass, an optimised rewrite of something already trusted.
The cost of skipping the oracle is that a shared bug is invisible. If both implementations are wrong in the same way, they agree, and a differential test reports agreement. It finds differences, which is not the same thing as finding errors.
Differential testing, differential fuzzing, property-based testing
- Property-based testing generates inputs and checks that some stated property holds: the output is sorted, the round-trip returns the original. You write the property.
- Differential testing generates inputs and checks a single implicit property: two implementations agree. You write no property, only the second implementation, which you usually already have.
- Differential fuzzing is the same comparison driven by a fuzzer, generating large volumes of input and often mutating towards inputs that reach new code. The difference is scale and input strategy; the idea is the same.
Equivl is differential testing with generated inputs from a seeded PRNG. The seed means every run is reproducible and every share link reproduces the exact run the sender saw.
Execution trace divergence
Comparing return values catches implementations that end up somewhere
different. It misses implementations that end up in the same place by
different routes: sometimes fine, sometimes a bug that will surface later on
an input you have not tried. A snippet can call emit(value) to
record intermediate state, and the recorded traces are then compared step by
step, so the run reports where the two implementations first diverged
as well as that they did.
Shrinking a counterexample
A generated failing input is usually far bigger than the bug needs. Shrinking repeatedly replaces the failing input with a smaller one that still fails, until nothing smaller does. Three details of Equivl's implementation:
- Probes run in batches, so minimizing costs a handful of executions instead of one per candidate. That matters most for the compiled languages, where each probe costs a sandboxed container and, on a metered server, a unit of quota.
- “Still fails” is answered by the comparison engine's own divergence check, so a run that failed because the backend was unreachable cannot be shrunk into a meaningless minimal input.
- Every accepted candidate must be strictly smaller by a size measure, which guarantees that shrinking terminates.
Nothing to install
JavaScript and TypeScript run in a Web Worker in your browser; Python runs in Pyodide, compiled to WebAssembly and served with the app rather than from a CDN, so it works offline. For those three, differential testing costs you a paste and a click. Java, Kotlin, C#, Go, Ruby and Rust execute through a sandboxed server that you run locally. That backend is not hosted publicly today.
Questions
What is differential testing?
Differential testing runs two implementations that are supposed to agree against the same inputs and treats any difference in their outputs as evidence of a bug in one of them. It needs no expected values, only a second implementation.
What is the difference between differential testing and property-based testing?
Property-based testing checks a property you write, such as the output always being sorted. Differential testing checks one implicit property instead (that two implementations agree), so you supply a second implementation rather than a specification.
What is differential fuzzing?
Differential fuzzing is differential testing driven by a fuzzer, generating large volumes of input and often mutating them toward new code paths. It differs from differential testing in scale and input strategy rather than in principle.
Can differential testing miss bugs?
Yes. If both implementations are wrong in the same way they agree, and the comparison reports agreement. Differential testing finds differences between implementations, which is not the same as finding all errors in either one.