Compare two images
Image comparison measures the difference between two images with several independent metrics (per-pixel error, structural similarity, perceptual hash distance and a visual diff map) because two images can be numerically close and visibly different, or the reverse.
Runs entirely in your browser; nothing is uploaded.
The images never leave the page
Decoding, resampling and every metric run in your browser. Nothing is uploaded, nothing is stored on a server, and there is no quota, because no server is involved in the comparison.
What each metric tells you
- Pixel error: mean and maximum per-channel difference. It is exact, so a one-pixel shift of otherwise identical content produces a large error, which is correct if you are checking an encoder and useless if you are checking a layout.
- SSIM (structural similarity) compares local luminance, contrast and structure in windows rather than pixel against pixel, so it tracks what a viewer notices more closely than raw error does.
- Perceptual hash (dHash and pHash) reduces each image to a short fingerprint and compares the fingerprints by Hamming distance. It is cheap, tolerates scaling and re-encoding on images with detail, and suits “is this the same image” better than “how different is it”.
- Diff map: a heat map of where the difference is. Two images with the same error score are very different problems depending on whether the error is spread thinly everywhere or concentrated in one region.
Two limits to check before you trust a score
A perceptual hash is unreliable on low-detail images
A pHash works by keeping the low-frequency DCT coefficients and thresholding them against their median. On a smooth image (a gradient, a flat background, a simple logo) nearly all the energy sits in a handful of coefficients and the rest land within a rounding error of the median, so those hash bits are decided by resampling noise. A gradient does not hash stably even against itself. The algorithm itself causes this, which is why the hash is one signal among several rather than the verdict. When the hash is what a score rests on, Equivl measures how far those coefficients actually sit from the median and flags the result when the margin is rounding noise, and a flagged result reads as near-equivalent at best however perfect the number looks.
Lossy compression sets a noise floor
Two JPEGs of the same scene differ everywhere at low amplitude from quantization. A diff visualization that scales to the maximum difference will therefore render compression noise as a bright full-frame difference and bury the actual change inside it. Part of what an image diff shows is the codec.
Different sizes, and the flag that follows
Two images of different dimensions cannot be compared pixel to pixel until one is resampled, and resampling invents information. That is done when needed, and it raises a flag, after which the result is capped at near-equivalent however high the score. The same applies to flattening transparency against a background.
Resemblance, optionally, with a model that runs in your browser
A further score, semantic similarity, is available on
request. It embeds both images with the vision half of CLIP
(clip-vit-base-patch32, quantized) and reports the cosine between
the two vectors. The model runs in your browser: it is a one-time download of
about 103 MB that the browser keeps, and the images still never leave the
page. Nothing is fetched until you press the button that states that size;
choosing the metric, or switching to image mode, downloads nothing. The pixel,
structural and hash metrics need no download at all.
That score is a claim about resemblance: whether the model finds two pictures alike, including across a change of size that leaves a pixel comparison nothing to line up (an image against its half-size copy measured 0.94). It is not a claim that two renderers agree, or that any pixel matches. Two screenshots that differ in the one number you care about can score almost perfectly. Unrelated images do not score near zero either; pairs with nothing in common measured between 0.50 and 0.74, so a score only means something well above that floor. The model also looks only at the central square of a non-square image, scaled down to 224 pixels, and the result says so when that happened. A semantic image score is therefore reported as near-equivalent at best, however high the cosine (even for two byte-identical files), and an image that failed to decode is a non-match under this metric as under every other.
Limits
- Images are compared as decoded pixels. File metadata such as EXIF, the embedded color profile and the compression settings is not compared, so two files can differ in bytes and still match.
- A pixel score says how much changed, not whether the change matters. A moved decimal point in a chart label is a few hundred pixels out of hundreds of thousands.
Questions
Are my images uploaded when I compare them?
No. Decoding and every metric run in your browser, so the images are never uploaded or stored on a server.
What is the difference between pixel error and SSIM?
Pixel error compares each pixel directly, so a small shift of identical content scores as a large difference. SSIM compares local luminance, contrast and structure in windows, which tracks what a viewer notices more closely than pixel error does.
Why does the perceptual hash disagree with the other metrics?
Perceptual hashes are unreliable on low-detail images. A gradient or a flat background puts nearly all its energy into a few coefficients, leaving the remaining hash bits to be decided by resampling noise, so such an image does not hash stably even against itself. When the hash is carrying the score, Equivl checks how much margin those bits were decided by and flags the result if there was none, so a hash match on a flat image is never reported as plain equivalence.
Can I compare two images of different sizes?
Yes, but one must be resampled first, and resampling invents information. The result is flagged and capped at near-equivalent rather than being reported as a clean match.
Does the semantic image similarity score use AI, and are my images sent anywhere?
It uses the vision half of the CLIP model, which runs inside your browser after a one-time download of about 103 MB that starts only when you press the load button, so the images still never leave the page. The score is a cosine between the two image embeddings. That is a claim about resemblance, not about two renderers agreeing or any pixel matching, so it is reported as near-equivalent at best.