What this tool does
Text can carry characters that render as nothing at all, and everything still works, so nothing tells you they are there. A zero-width space between two words looks like a space, compares as different, and breaks find-and-replace. A right-to-left override can make a line of code read backwards in an editor while looking correct in a browser. Text pasted from an AI assistant, a website, a PDF or a messaging app can arrive carrying a dozen of them without you noticing. This tool finds every one, names it, and removes it on request.
How it works
Every character in your text is checked against the Unicode ranges that occupy no space: the zero-width family, the word joiner, the deprecated Mongolian vowel separator, bidirectional embeddings and overrides, directional isolates, the invisible mathematical operators, tag characters, variation selectors and interlinear annotation marks. Each one is reported with its codepoint and its official name, so a finding can be checked rather than taken on trust.
Not everything invisible is safe to delete. A zero-width joiner holds together the parts of an emoji sequence, so a family emoji is really four people joined by three invisible characters, and stripping the joiners leaves four unrelated people. The same joiner builds conjuncts in Devanagari, Bengali, Arabic and Persian. Joiners are therefore reported but left alone unless you explicitly ask for them, and a blank Hangul filler or Braille pattern is treated the same way because those are real letters from another script.
Unusual spaces get different treatment again, because they are not invisible. A non-breaking space or a thin space occupies a position and is visible, so removing it would weld two words together, which is a worse outcome than the odd spacing. Those are rewritten to a normal space instead, which fixes string comparison and search without changing what the text says.
The last check is for lookalike letters, which are the security-relevant case. A Cyrillic 'а' and a Latin 'a' are different characters that render almost identically, so a domain, a filename or a shell command can be made to look like something it is not. The tool maps the common collisions back to the Latin character they imitate and reports each substitution, which turns an invisible trick into a visible diff.
Worked example
An invoice total keeps failing a validation check even though the copied string looks correct, and a command copied from a forum behaves unexpectedly.
- Paste the invoice text and the tool reports 6 findings, 4 of them zero-width characters
- The non-breaking spaces are reported separately from the invisible ones, because they are visible
- Enable removal and the zero-width characters are deleted, the spaces become normal spaces
- The pasted command is flagged as containing a Cyrillic 'р' and 'о' among Latin letters
The cleaned invoice compares equal to what the portal expects, and the command now contains only the characters it appears to contain.
Accuracy and limitations
- Removing a character that was carrying meaning changes the text. A joiner inside an emoji sequence or an Indic conjunct is why those are left alone by default.
- This tool finds characters. It cannot tell you whether a particular one was intentional, which is why it reports before it changes anything and shows a preview of the result.
- A string can be made to look like almost anything using a combination of overrides, so treat text from an untrusted source as data rather than as instructions.
- Lookalike correction covers the letter pairs that collide in common interface fonts. A full Unicode confusables table is far larger and would report pairings nobody actually encounters.
Frequently asked questions
- What are invisible characters and why are they in my text?
- They are code points that render as nothing, most often zero-width spaces, byte-order marks, joiners and variation selectors. They arrive whenever text is copied from a website, a PDF, a messaging app or an AI assistant, and they survive every ordinary edit because there is nothing on screen to select.
- How do I find a zero-width space?
- Paste the text in. Anything invisible is reported with its codepoint, so U+200B tells you it is specifically a zero-width space rather than a byte-order mark or a joiner, which are different characters with different effects.
- Will removing these break my emoji?
- Not with the default settings. Emoji sequences are held together by zero-width joiners, and those are reported but not removed unless you tick the box to include them, precisely because removing them splits one emoji into several.
- Why does my string not equal the one on the website?
- Almost always because of an invisible character or an unusual space. A non-breaking space looks identical to a space but is a different code point, so a strict comparison fails even though nothing appears different. The normalise option rewrites those to a real space.
- What are lookalike letters?
- Characters from other alphabets that render almost exactly like Latin ones, such as a Cyrillic 'а' that looks like 'a'. They matter because a domain, filename or command can be built from them and still look correct, which is a common way of spoofing both. This tool reports each one and can convert it to the Latin letter it imitates.
- Is my text uploaded anywhere?
- No. Every check runs as ordinary JavaScript in your browser tab. The text is not sent anywhere, not stored, and not logged, which is the only way a tool like this can be trusted with a draft you have not published yet.