Equivl

How to verify a refactor did not change behavior

To verify that refactored code still produces the same output as the original, run both versions against the same set of inputs and compare their results. The original is the specification, so any input on which the two disagree is a behavior change the refactor introduced.

Compare two snippets

JavaScript, TypeScript and Python run in your browser. Nothing to install, no account.

The original is the specification

A refactor is supposed to preserve behavior, so you already have an oracle: the code you started from. You do not have to work out what the function should return, because whatever the old version returned is the answer, bugs included. That is the setup differential testing needs.

Why your test suite is not quite the same check

Your suite checks the cases somebody thought to write down. A refactor can break behavior at edges nobody wrote down: the empty input, the duplicate key, the negative zero, the value that was exactly at a boundary the loop used to handle differently. Generated inputs can reach cases that hand-written ones miss, and the check needs no new assertions because the old implementation supplies every expected value.

This complements a test suite and does not replace one. The suite encodes what the code is meant to do; this checks that whatever it does now, it did before.

Refactors where this helps

What it will not tell you

It reports only what it observed. If a parameter's type is never generated as the shape that breaks (a string where the generator only produced integers), the divergence is not in the sample and the run reports agreement. Agreement means no difference was found across the cases that ran. Raising the case count buys more confidence, but agreement is never a proof of equivalence.

Questions

How do I verify that refactored code produces the same output as the original?

Run both versions against the same generated inputs and compare their results. The original version acts as the oracle, so any input on which the two disagree is a behavior change introduced by the refactor.

Why not just read the diff?

A text diff shows how much the source changed, which is almost unrelated to how much the behavior changed. A pure reformatting can produce a huge diff, and a single changed comparison operator can produce a one-character diff that breaks a whole class of inputs.

Does passing mean the refactor is safe?

It means no difference was found across the cases that ran. That is evidence rather than proof, and how much confidence it buys depends on how many cases ran and whether the generated inputs covered the shapes that matter.

Related