Compare two files, zips or folders byte for byte
A byte-for-byte file comparison reads every byte of two files and reports either that they are identical or the offset of the first byte where they differ, together with a SHA-256 checksum of each file.
Runs entirely in your browser; nothing is uploaded.
Two files
Any two files can be compared: installers, firmware images, disk images, archives, or two copies of anything whose contents should match. Each file is read once, in a background worker, and that one pass computes its SHA-256 and compares it with the other file. The result is Identical when every byte matched, or Different with the offset of the first differing byte, how many bytes differ and in how many runs, and a hex view of the bytes around each difference. When one file matches the start of the other exactly, the result says that one is a prefix of the other instead of pointing at a byte.
There is no similarity score. A percentage for two binaries invites "99.99% the same" about a firmware image that does not boot, so the answer is a word and an offset.
Checking a download against its checksum
Drop one file and paste the checksum published beside the download. The
algorithm is read from the checksum's length: 64 hex digits is SHA-256, 128 is
SHA-512, 40 is SHA-1 and 32 is MD5. A bare checksum, a sha256sum
line with the file name after it, and the BSD form
SHA256 (file) = … are all understood, and capitals and spaces are
ignored.
A match shows the file is the one the checksum was computed from, byte for byte. It does not show that the file is safe: that depends on where the checksum came from, and a checksum published on the same page as a tampered download proves nothing. SHA-1 and MD5 are broken for security, so a match on either shows only that the file was not corrupted in transit.
Two zips or two folders
Two zip files, two folders, or one of each are paired by path and shown as a tree of what changed: each file is changed, only in A, only in B or identical. The tree opens on the changes, with identical files folded away and counted. Sizes are compared first, then, for two zips, the CRC-32 each zip records; a difference in either proves a change without unpacking anything. A matching size and CRC-32 proves nothing, because a CRC-32 collision is easy to make, so those files are unpacked and compared by SHA-256 before they are called identical.
A changed file can be opened in the mode that reads it: a JSON file in Data mode, a CSV in Tables mode, a PDF in Documents mode, an image in Image mode, a recording in Audio mode, and anything else as bytes. A switch, off by default, pairs a file only in A with a file only in B when their contents are the same bytes, and reports it as moved.
Limits
- A zip or folder is compared up to 20,000 files a side, and a zip up to 2 GB unpacked a side. The unpacked size is counted as the bytes come out, not taken from the zip's own headers, which a zip bomb lies in.
- Zip entries compressed with stored or deflate are read, which covers zips made by Windows, macOS and most tools; other methods and password-protected entries are listed as not read, never as identical.
- A zip inside a zip is compared as one file; it is not opened.
- Empty folders are not compared: a folder appears only as the path of a file inside it.
- Folders can be dropped or chosen in a desktop browser. Phone browsers generally cannot pick a folder; files and zips work there.
- A file opened from inside a zip in another mode is unpacked into memory, up to 100 MB.
Questions
Are my files uploaded to compare them?
No. Both files are read and hashed inside your browser, in a background worker, and never leave the page. They are not kept between visits; saving a comparison to history keeps the names and counts, never a path or a byte.
How do I check that a download arrived intact?
Drop the downloaded file and paste the checksum published with it. A SHA-256, SHA-512, SHA-1 or MD5 checksum is recognized by its length, and the result says whether the file matches it. A match shows the file is unchanged from the one the checksum was made from, not that it is safe.
Can two different files have the same checksum?
For SHA-256, no such pair has ever been found. Two plain files are compared byte for byte, so that verdict does not rest on the hash at all; files inside two zips or folders are called identical when their sizes and SHA-256 checksums match. For MD5 and SHA-1, yes: collisions can be made on purpose, which is why a match on either is reported as ruling out corruption, not tampering.
What happens with a very large file?
It is read in 4 MB pieces, so memory stays flat whatever the size, and there is no size limit on two plain files. Two 1 GB files took about 8 seconds in a desktop browser, most of it spent computing the two checksums, and a progress bar has a Stop button.