What this tool does
Paste two versions of something and the first thing you see should not be forty green lines caused by a Windows line ending. Normalisation runs before anything is compared — CRLF against LF, trailing whitespace, tabs, a byte-order mark — because that single step removes more false differences than any other. What is left is a real diff, at line or word granularity, with a unified diff output carrying genuine @@ hunks and line numbers.
How it works
The core is a longest common subsequence over tokens. Every token that appears in both texts in the same order is kept; everything else is a deletion or an insertion. The subtlety is memory. A full m by n table of numbers would be four bytes per cell, which for two 100 KB files is about ten gigabytes, so only the values live in two rolling rows and the directions live in one byte per cell. The direction table is unavoidably m by n — there is nothing to walk the answer back through without it — but at one byte per cell it stays near four megabytes. Past roughly four million cells the longest common subsequence is skipped entirely and a cheaper prefix-and-suffix comparison is used instead, and the result says which strategy ran rather than pretending a line-level match was found.
Before any of that, both sides are put into the same shape. A byte-order mark is removed, CRLF and bare CR are collapsed to a single newline, and per-line whitespace is optionally collapsed and trimmed. This is the step that fixes the complaints a naive diff tool generates: a file saved on Windows and the same file saved on Linux differ in every single line by an invisible carriage return, a file with a trailing newline and one without differ in the last line, and a CSV exported with a BOM has a first column called \uFEFFname rather than name. Tabs against four spaces is the same story. Each of these is a formatting difference, not an edit, and treating it as an edit is what makes people stop trusting diff tools.
The diff runs at two granularities. Line level answers "which lines changed", and word level answers "which words inside the unchanged lines changed", so a single edited word glows inside a paragraph that otherwise shows no change at all. The word level runs the subsequence over whitespace-separated words and then slices the original spacing back out of whichever side a run came from, so re-wrapped text does not read as a rewrite and the parts still concatenate back into exactly what you pasted. The line splitter keeps each line's own trailing newline inside the token, the way git does, which is what makes a file that ends with a newline and one that does not show up as the real difference it is instead of sliding past unnoticed.
The output is a genuine unified diff: --- and +++ headers, @@ -start,count +start,count @@ hunk headers with true line numbers, context prefixed by a space, additions by + and deletions by -. Deletions are written before additions inside a hunk, which is the order a patch file requires, and a file with no trailing newline gets the \ No newline at end of file marker where git puts it. That makes the output close enough to git apply to be saved and used. The similarity figure is unchanged tokens over total tokens as a percentage to one decimal place, and it is worth being blunt about what it is: a measure of how much of the text is in the same order, not a plagiarism score and not a similarity judgement about authorship.
Worked example
Checking a config change between two versions, and finding out whether the rest of the file really is unchanged.
- Left: three lines ending in a newline. Right: the same three lines with "line two" changed to "line 2"
- Normalisation runs first: both files end with a newline and both use LF, so nothing is hidden and nothing is invented
- Line diff reports 1 added, 1 removed, 2 unchanged, 50.0% similarity
- Unified diff output: --- a / +++ b, then @@ -1,3 +1,3 @@ with " line one" as context, -line two, +line 2, and " line three"
- CRLF on the left only is normalised away first, so the same comparison does not report all three lines as changed
One changed line, real line numbers in the hunk header, and a patch that could be saved and applied rather than a screenshot of coloured text.
Accuracy and limitations
- The similarity percentage is unchanged tokens over total tokens. It measures how much text is in the same order, and is not a plagiarism score, a duplicate-content check or any kind of judgement about authorship.
- Above roughly four million table cells the longest common subsequence is skipped in favour of a prefix-and-suffix comparison, and the middle is reported as one deletion followed by one insertion. The result says which strategy ran, but a very large pair will not give you fine-grained word-level matching.
- Normalisation is aggressive by design. If a difference in whitespace, case or line endings is genuinely the change you are looking for, turn those options off — the defaults are chosen for people who want to see content edits, not byte-level ones.
Frequently asked questions
- Why did my diff show every single line as changed?
- Almost always because of a line-ending difference. One file saved on Windows uses CRLF and the other on Linux uses LF, and a character that renders as nothing is still a character, so every line differs. This tool normalises line endings, trailing whitespace and the byte-order mark before comparing, which is why it usually shows one changed line where a naive tool shows a screen of green.
- Is the similarity percentage a plagiarism score?
- No. It is unchanged tokens divided by total tokens, as a percentage to one decimal place, so it tells you how much of the text appears in the same order in both inputs. It says nothing about where the text came from, who wrote it, or whether it was copied, and no tool of this kind can answer that without a corpus to compare against.
- Can I save the output as a patch file?
- Yes, the unified diff output is real patch format: --- and +++ headers, @@ hunk headers with real line numbers, context lines prefixed with a space, and the \ No newline at end of file marker where git puts it. Copy it into a .patch file and it is close enough to git apply to be useful, which is a better outcome than a screenshot of coloured text.
- What is the difference between line diff and word diff?
- Line diff tells you which lines changed, which is what you want for a code or config change. Word diff compares the whitespace-separated words inside those lines, so a single edited word lights up inside a paragraph that otherwise shows no change at all. Word mode also means a re-wrapped sentence does not read as a full rewrite, because the words themselves are in the same order.
- How large a document can this handle?
- The comparison is bounded at about four million table cells, which is a couple of thousand lines each side. Past that the middle is handled by a cheaper prefix-and-suffix comparison and the result tells you so, rather than quietly returning a less precise answer. For very large files this is a browser-side tool working in memory, not a diff service.
- Is my text uploaded anywhere?
- No. Both texts are processed in your browser tab and nothing is sent, stored or logged. That matters here more than for most tools, because the things people diff are frequently unpublished drafts, credentials that slipped into a config file, or customer records.