Compare two texts
Text similarity comparison measures how close two pieces of text are using several independent metrics (character edit distance, shared-word overlap, weighted-term cosine similarity and a line-level diff) because each one captures a different kind of closeness.
Runs entirely in your browser; nothing is uploaded.
Nothing leaves your browser
Text comparison runs entirely client-side. There is no upload, no backend, no account and no quota. The text you paste never leaves the page. If the text is something you would not paste into a stranger's server, check whether any tool, this one included, uploads it.
Four metrics, because “similar” is not one thing
- Edit distance (Levenshtein) counts the single-character insertions, deletions and substitutions needed to turn one text into the other. It is the right measure for near-identical text (a typo, a changed date) and a poor one for two paragraphs that say the same thing in different words.
- Token overlap (Jaccard) measures the share of distinct words the two texts have in common. It ignores order entirely, so a reordered sentence scores as identical.
- Weighted cosine similarity (smoothed TF-IDF) weights words by how distinctive they are, so agreement on the and and counts for little and agreement on a rare term counts for a lot. It is the one of the four built for “are these about the same thing”.
-
Line diff (longest common subsequence) shows which lines
were added, removed and kept: the familiar diff view, for when you want to
see the change rather than score it. Past about 1,000 lines a side it first
anchors on lines that appear once in each text, as
git diff --patiencedoes; a stretch with nothing to anchor on shows as removed and added, and the diff says so.
Two texts can score high on one and low on another, and that spread is informative: a high token-overlap score with a low edit-distance score suggests heavy rewording of the same content; the reverse suggests a small edit somewhere that matters.
Meaning, optionally, with a model that runs in your browser
A fifth score, semantic similarity, is available on request.
It embeds both texts with a small sentence-embedding model
(all-MiniLM-L6-v2, quantized) and reports the cosine between the
two vectors. The model runs in your browser: it is a one-time download of
about 38 MB that the browser keeps, and the text still never leaves the page.
Nothing is fetched until you ask for it, and the size is stated on the button.
That score measures whether two passages sit near each other in a learned space, which tracks “are these about the same thing” better than any count of shared words. It is not evidence that they say the same thing. Two sentences that contradict each other can score highly, and the model reads only the first few hundred words of a long text, which the result says when it happens. So a semantic score is reported as near-equivalent at best, however high the cosine. It can never produce a plain match, by design.
Normalization is optional
Before comparing, text can be normalized: line endings unified, trailing whitespace dropped, Unicode forms folded. Whether that is correct depends on your question. Two files that differ only in line endings are the same document to a reader and different bytes to a checksum, and only you know which you meant.
When a comparison has to approximate
Edit distance on very large inputs is quadratic, so past a size threshold it is skipped rather than left to hang. When that happens the result is flagged, and a flagged result is reported as near-equivalent at best, never as a clean match.
Limits
- A score measures how much of the text changed, not what the change means. Changing a refund window from 14 to 30 days in a four-line policy still scores 99.2% by edit distance. Read the differences, not only the score.
- The texts are compared as plain text. Formatting, links and anything else lost when they were pasted is not in the comparison.
- Near-equivalent means the score cleared the chosen threshold. A different threshold gives a different verdict for the same texts.
Questions
Is my text uploaded anywhere?
No. Text comparison runs entirely in the browser with no backend involved, so the text you paste never leaves the page.
Which similarity metric should I use?
Use edit distance for near-identical text such as a typo or a changed value, token overlap when word order does not matter, and weighted TF-IDF cosine similarity to judge whether two texts are about the same thing. The line diff shows the change rather than scoring it.
Why do the metrics disagree with each other?
Because they measure different kinds of closeness. High word overlap with a large edit distance suggests the same content reworded heavily, while the opposite pattern suggests one small but significant edit.
Does the semantic similarity score use AI, and does it send my text anywhere?
It uses a small sentence-embedding model that runs inside your browser after a one-time download of about 38 MB, so the text still never leaves the page. The score is a cosine between the two embeddings, which is a measure of resemblance and not proof of identical meaning, so it is reported as near-equivalent at best.