What this tool does
A cryptographic hash answers the wrong question for images. It tells you whether two files are the same bytes, so re-saving a JPEG, resizing it, or pulling it out of a CMS that stamps EXIF makes it a different file with total confidence, even though the picture has not changed. A perceptual hash fingerprints what the image looks like instead, which is what makes it usable for finding the same supplier's photo on three different products in a catalogue.
How it works
aHash is the cheap one: an 8 × 8 greyscale reduction, then one bit per cell for brighter or darker than the mean of those 64 cells. It is also the fragile one, and the reason is that the threshold moves with the image. Raise the brightness of the whole picture and the mean rises with it, but not quite in step, so cells sitting near the mean swap sides and bits flip all over the hash. It is still the fastest option, and a flat image always hashes to all zeros, which is a useful way of spotting a blank background.
dHash looks at 9 × 8 cells and records, for each row, whether each cell is darker than the one to its right: eight comparisons per row, 64 bits. Nothing is compared to an absolute level, only to a neighbour, so adding 40 to every pixel leaves every comparison pointing the same way. That is why it survives a brightness or contrast change where aHash does not. The cost is that it sees only horizontal structure, so a purely vertical gradient is invisible to it.
pHash is the robust one: a 32 × 32 greyscale reduction, a two-dimensional DCT, and then the top-left 8 × 8 block of coefficients, which is where the slow-changing structure lives. The DC term at position (0,0) is mathematically just the average brightness, and it is discarded on purpose, because keeping it would reintroduce exactly the thing a perceptual hash exists to ignore. The remaining 63 coefficients are each compared to their median, which is the whole payload, so a pHash really carries 63 useful bits and one pad.
The distance is the Hamming distance: how many of the 64 bits differ. A number on its own means nothing, because two unrelated photographs land about 32 bits apart purely by chance — 64 bits split evenly between same and different, so 32 is the line between coincidence and signal. Every band is read against that: 0 to 4 is a re-save, 5 to 10 an edited or recompressed version, 11 to 15 worth eyeballing, and past 25 you are back in the range random images occupy. One limitation has no workaround: all three algorithms read the image in a fixed order, so a mirrored photo hashes as maximally different. Hash the image and its flip.
Worked example
A catalogue has the same supplier photo attached to three different SKUs, and one variant is mirrored.
- The two candidates pHash to 0f7c1a3e5b9d0247 and 1f6d0b52a4c8e935
- XOR then count set bits: 31 of 64 bits differ
- 31 sits on the 32-bit line two unrelated photographs occupy, so this is a different image
- Hashing the first image mirrored gives 3e5b9d2470f7c1a3, which is 30 bits from the original
- A one-character edit, 0f7c1a3e versus 0f7c1a3f, is a distance of 1, which is a re-save
A distance of 1 is the same picture stored twice, and a distance of 30 or 31 tells you nothing — those are the numbers random images produce. The mirrored comparison is what catches the third variant, because without the flip a true duplicate reads as unrelated.
Accuracy and limitations
- A perceptual hash matches what an image looks like, not what it contains. It is a catalogue de-duplication tool, not proof that two files are the same picture or that either is authentic.
- Groups come from single-link clustering, so A can merge with B and B with C while A and C are nothing alike. Each group reports its own maximum pairwise distance; a value above the threshold means the group was chained.
- Crops and rotations are not handled by any of the three, because they change the low-frequency layout the hash is reading.
Frequently asked questions
- What is a perceptual image hash?
- A 64-bit fingerprint of what an image looks like rather than what it contains. Re-saving, recompressing or resizing the same picture keeps the hash nearly the same, which is exactly what a cryptographic hash refuses to do and why SHA-256 is the wrong tool for finding duplicate photos.
- What is the difference between aHash, dHash and pHash?
- aHash thresholds 8 × 8 cells against their mean and breaks on a brightness change. dHash compares neighbouring cells, so it survives brightness and contrast. pHash runs a DCT and keeps the low-frequency block, discarding the average-brightness term, which makes it the most robust and the slowest of the three.
- What Hamming distance counts as a duplicate?
- Use 0 to 4 for a re-save, 5 to 10 for an edited or watermarked version, and 11 to 15 for a candidate worth looking at. The review thresholds used here are 10 for aHash, 10 for dHash and 12 for pHash, and anything above that is unlikely to be real.
- Why does a distance of 30 not mean anything?
- Because two unrelated photographs land about 32 bits apart out of 64 purely by chance, since the bits split evenly between same and different. A reported distance of 30 or 31 is indistinguishable from coincidence, and only a number well under 25 carries information.
- Why does my mirrored duplicate not match?
- All three algorithms read the image in a fixed order, so left and right are not interchangeable and a flipped photo lands around 30 bits away. Hash the image and its mirror image as well, and treat a match against the flipped hash as a mirrored duplicate rather than a different picture.
- How does it group duplicates?
- Single-link clustering: A merges with B when they are close, and B merges with C when they are too, so a chain of near-misses can put two unrelated images in one group. Each group reports its own maximum pairwise distance, and a value above the threshold is the visible sign that it was chained rather than matched.