Compare two audio files and see how similar they are
An audio comparison lines two recordings up in time, compares their spectra moment by moment, and reports how similar they sound as a score from 0 to 100%, with the moments where they differ most.
Runs entirely in your browser; nothing is uploaded.
What the score means
Both recordings are decoded by the browser, mixed to mono and resampled to 16 kHz. Each is cut into frames of 32 milliseconds, and each frame is described by the loudness of 40 frequency bands. The two recordings are lined up by where their sounds begin, and the score is the average match between lined-up frames. The score measures similarity and does not say the recordings are the same: a re-encode, a remaster and a different take all land somewhere on the same scale.
The scale has three bands. Below 40% reads as different recordings; 40% to 95% as the same material, audibly changed; 95% and above as the same recording, re-encoded. On the test recordings the bands were chosen from, a 64 kbps MP3 of a master scored 98.8% against it, the master with noise mixed in at a 17 dB signal-to-noise ratio scored 77.1%, a copy shifted up a semitone scored 89.0%, and three pairs of different pieces scored 5% to 10%. Those recordings were synthesizer loops, so the band edges are a guide measured on one kind of sound and may not hold for others.
The page reports Identical only when the two files hold the same samples: equal bytes, checked before anything is decoded, or equal samples on every channel at the files' own sample rate. Any score short of that is capped at Near-equivalent, however close it is, because the comparison discarded sound above 8 kHz and the difference between channels to reach it.
Offsets, tempo, and a clip inside a longer track
A recording that starts later than the other is found at its offset, and the score covers the stretch where both play. A copy played up to 10% faster or slower is matched by trying tempos from 0.90× to 1.10× and keeping the one whose onsets line up best; the tempo is chosen by that alignment and never by the score. A short clip is found where it sits inside a longer recording, and the result says how much of the longer one it covers rather than claiming the clip came from it: the same loop in two songs would read the same way.
Pitch is not searched. A copy shifted by a semitone scores lower and reads as audibly changed rather than as the same recording.
What you see
- Both waveforms on one time axis, lined up by the offset and tempo that were found.
- A strip under them shaded by how little each second matched, and up to five moments that differ most, each of which plays both recordings from that point.
- A play button for each recording, and a button that plays the difference between them when their tempos match.
- The score's lossy steps named as flags: resampled to 16 kHz, mixed to mono, stretched to a tempo.
Practicing a line: voice likeness
Audio mode also asks a second question: how close is my reading of a line to a target reading of the same line? Choose Voice likeness, drop or record the target and your take, and each side is measured for pitch, loudness and pauses, then the two readings are lined up word for word. The result is one row per dimension, each with a sentence on how your take differs: intonation (the tune of the line), pauses and emphasis are scored from 0 to 100%, and pitch and pace are reported as readouts, because an impression may choose its own pitch and speed. The two pitch lines are drawn over each other, each measured from its own speaker's middle pitch.
An optional voice model adds a sixth row, timbre: how alike the two voices sound, from a speaker model (ECAPA-TDNN, trained on VoxCeleb) that downloads once, about 57 MB, from a button that states the size, and then runs in your browser. With it loaded, a recording that sounds like two voices reads "Can't tell" on every row instead of being scored. An impression usually scores low on timbre even when the delivery is right, and the model hears a change of pitch as a change of voice.
There is no overall score and no verdict. This is practice feedback on how a line is spoken, not identification: no row says whose voice either recording is. Without the voice model the sound of the voices is not compared, so two different voices reading with the same melody and timing score high, and a second voice in a recording is not detected; with it, two people who sound alike can still pass as one. Each side needs at least 3 seconds of speech and at most 30 seconds in all. A pair that looks like two different lines gets a warning, and the check can miss.
Limits
- Nothing above 8 kHz is compared. Two MP3s at 320 and 128 kbps differ mostly above that, so they can score the same.
- Stereo is mixed to mono, so a swapped or missing channel does not change the score.
- One offset lines up the whole recording. A pair where one has a stretch cut out or inserted scores low after the cut, because nothing follows the edit.
- A recording with no sound in it (quieter than −60 dBFS throughout) gets no score: two silent files would match perfectly and mean nothing.
- The formats are whatever this browser decodes: WAV, MP3, FLAC, Ogg, Opus and M4A in current browsers. A file it cannot decode is named, with a way to compare the two files as bytes instead.
- Files up to 300 MB and recordings up to 30 minutes are compared. A 10-minute pair took about 3 seconds in a desktop browser.
Questions
Are my recordings uploaded?
No. Both files are decoded and compared inside your browser and never leave the page. Files up to 8 MB are remembered in this browser for your next visit; saving a comparison to history keeps the names, settings and score, never the recordings.
Does a high score mean it is the same recording?
No. A high score means the two sound alike once lined up, as measured at 16 kHz in mono. Only two files holding the same samples are reported as identical; everything else, however close, is reported as near-equivalent at best.
Can it tell whether a clip was taken from a longer track?
It can find where a clip matches inside a longer recording and say how much of it the clip covers. It cannot say the clip was taken from it: the same sample or loop used in two songs would match the same way.
Can I share a comparison as a link?
No. A link cannot carry two recordings, so Audio mode has no share link.
Can voice likeness tell whether two recordings are the same person?
No. It compares how two readings of one line are spoken (melody, pace, pauses and loudness) and gives no verdict about who is speaking. The optional timbre row says how alike two voices sound to a speaker model, and it shifts with pitch, microphone and room, so a close result is not evidence that a recording is a particular person.