Compare two PDFs by text and by page
A PDF diff extracts the text of two PDF files page by page and compares it line by line, then draws each page at a fixed resolution and compares the paired pages as pictures, so a changed clause shows in the text diff, an inserted page is reported as inserted, and a moved figure can show as pixel differences on its page.
Runs entirely in your browser; nothing is uploaded.
Text: what changed, with page markers
The text of every page is read out of the file with pdf.js, line by line as the PDF lays it out. The lines of each document are joined and compared with the same line diff and score that Text mode uses, and a marker in the diff shows where each page of each file begins, so a page that exists on one side only is visible as a run of added lines under its own marker. Line breaks are kept as the PDF has them: two exports that wrap a paragraph differently show changed lines, with the moved words marked, rather than being silently rejoined.
Text pulled out of a PDF is not the PDF. Layout, fonts, images and anything
drawn as a picture are not in it, and the reading order of a two-column page
or a table is the order the file stores it in, which is not always the order a
reader would use. Every text result therefore carries the flag
document-text-extracted, and a perfect text match is reported as
near-equivalent, never as equivalent. Only two files with exactly the
same bytes are reported as identical, and that is checked before
either file is parsed.
Pages: each pair as a picture
Pages are paired by content, not by position. Each page of one file is scored against each page of the other by the words they share (or, for pages with no text, by a small fingerprint of how they look), and a sequence alignment picks the pairing. A page inserted in the middle of a document is reported as inserted, and the pages after it still pair with their counterparts, instead of every later page reading as changed.
Every paired page is drawn at 96 dots per inch and compared the way Image mode
compares two pictures, with a structural similarity score and a map of where
the pixels differ. The document's page score is the mean of the pair scores
over the longer document's page count, so a page one side lacks costs a full
page's share. Because a pixel at 96 DPI is about a quarter of a millimetre, a
hairline rule, a font one size smaller or a shifted baseline can vanish into
it; every page result carries rendered-at-fixed-dpi and is
reported as near-equivalent at best. Two pages of different sizes,
such as Letter against A4, are not resampled to fit and count as changed.
Scanned PDFs
A scanned PDF is a picture of text with no text inside it. Its pages yield no text, so the text comparison stops and says which file is scanned, rather than comparing an empty string and calling the result a score. Two different scans are never reported as matching text. The pages can still be compared as pictures, and a scan measured against the file it was printed from scores low even when it is faithful, because skew, paper tone and compression change every pixel. There is no OCR: Equivl does not read text out of pictures.
Locked and damaged files
A password-protected PDF asks for its password on the page. The password is used in your browser to open the file and is never stored, remembered or put in a link; a file that opened with one is noted in the result. A PDF that opens without a password but forbids copying is read normally. A file that is not a readable PDF, because it is damaged or was cut off while downloading, is named as such, and a file that fails part-way is never compared as the pages that did read.
Limits
- Each file is read up to 100 MB and 2,000 pages, checked before it is parsed, because everything runs in the browser tab.
- A PDF's own JavaScript is never run, and the fonts and character maps pdf.js needs are served from this site, so opening a document fetches nothing from anywhere else.
- Form field values, annotations' comments and attachments are not compared.
- There is no option to ignore line breaks.
Questions
Are my PDFs uploaded to compare them?
No. Both files are opened, read and drawn inside your browser and never leave the page. They are not kept between visits either; saving a comparison to history keeps the file names, page counts, scores and flags, never a page or a line of text.
Can it compare a scanned PDF?
Only as pictures. A scanned page has no text layer, so the text comparison stops and says so. The pages can be compared as images, and a scan against the original scores low even when it is faithful. Equivl has no OCR.
Why is a perfect match reported as near-equivalent rather than equivalent?
Because both comparisons discard something. Extracted text drops layout, fonts and images, and a page drawn at 96 dots per inch drops anything finer than a pixel at that size. Only two files with exactly the same bytes are reported as identical.
What happens when a page was inserted?
Pages are paired by content, so the inserted page is reported as inserted and the pages after it still pair with their counterparts. The document page score counts the inserted page as a full page of difference.