Check whether two functions produce the same output
Execution-based code equivalence checking tests whether two differently written implementations behave the same way by running both on the same inputs and comparing the results, instead of analysing their source.
JavaScript, TypeScript and Python run in your browser. Nothing to install, no account.
Black-box equivalence checking, in the browser
Equivl treats the first function defined in each snippet as the entry point and
reads its parameter list to decide what to generate; you can annotate a
parameter with a @param {type} name comment when the name alone is
ambiguous. Both functions then receive identical positional arguments, and
their return values are compared after cross-language canonicalization.
This is black-box checking: nothing is inferred from how the code is written, so the two implementations can share nothing but a signature. Code that a static analyser can say little about matters less here, because the check never inspects the implementation.
Four verdicts
- identical: the results matched exactly, leaf for leaf.
- equivalent: they matched under the strictness rules you chose, for instance treating
1and1.0as the same number. - near-equivalent: they matched, but something in the comparison had to be approximated, so the match cannot be reported as plain equivalence.
- divergent: they disagree, and you get the inputs they disagree on.
The third verdict guards against overstating a match. Anything a comparison has to resample, truncate or approximate (rescaling an image so two sizes line up, skipping edit distance on an enormous input, flattening transparency) raises a flag, and a flagged result is capped at near-equivalent however perfect its score, so a qualified match is not reported as plain equivalence.
Comparing across languages
Two languages can differ in ways that are not behavioral differences:
Python's None against JavaScript's null, a tuple
against an array, 1 against 1.0, dictionary key
order, set order. Each is a toggle you set according to what you meant. With
all of them on, the match is strict; turning one off says that particular
difference does not matter for your purposes.
Return values, or every step along the way
By default only the return value is compared, which is black-box equivalence in
the strict sense. A snippet can also call emit(value) to record
intermediate values, and those traces are then compared step by step, which
catches two implementations that reach the same answer while diverging
internally and converging again. Trace comparison is a toggle, because
sometimes that internal difference is exactly the change you were making.
What this is not
Equivl does not do formal verification. Agreement across generated inputs is evidence, not proof. A run can show that two implementations differ and can build confidence that they do not, but it cannot establish equivalence for all possible inputs, which is what a symbolic equivalence checker attempts. In exchange, it runs real code in the languages Equivl supports, including code a symbolic tool may be unable to model.
Limits
- Equivalent is evidence, not proof. An input the generator never produced can still separate the two, and a bug both share reads as agreement.
- Only return values and
emit()values were compared. Console output, running time, memory and anything written outside the function were not. - A divergence is a fact about the inputs that separated the two; it does not say which side is right.
Questions
Is there a tool to check if two functions have the same output?
Equivl does this in the browser. Paste both functions and it generates inputs from the entry function’s parameters, runs both implementations on the same arguments, and reports each generated case where their results differ.
Does agreement on generated inputs prove two functions are equivalent?
No, it is evidence rather than proof. Differential testing can definitively show that two implementations differ, but agreement across a sample of inputs cannot establish equivalence for every possible input.
What does black-box equivalence checking mean?
It means the check relies only on inputs and observed outputs, never on the internal structure of the code. Two implementations can share no code, no algorithm and no language and still be compared.
Which languages can I check?
JavaScript, TypeScript and Python run entirely in your browser, TypeScript transpiled with sucrase and Python through Pyodide compiled to WebAssembly. Java, Kotlin, C#, Go, Ruby and Rust run through a sandboxed execution server you run locally; they are not hosted for public use today.