Invisible Character Detector & Remover
Find and remove hidden Unicode characters, zero-width spaces, non-breaking spaces, BOM characters and other invisible characters from text.
Your text never leaves your browser.
Character inspector
Hidden characters will appear here as labelled tokens, right where they sit in your text.
What are invisible characters?
Invisible characters are Unicode code points that take up no visible space or look exactly like an ordinary space. They exist for good reasons — controlling where lines break, joining emoji, shaping scripts such as Persian or Hindi, or marking the byte order of a file.
When text is copied between PDFs, web pages, word processors, chat apps and spreadsheets, these characters often come along for the ride. You can't see them, but software can: they change string lengths, break searches and make identical-looking values unequal.
HelloZWSPworldlooks like “Helloworld”, but has 11 code points
Common invisible characters
Not every invisible character is a mistake. Joiners, directional marks and variation selectors carry meaning in emoji and many languages, so CleanGlyph marks them as review recommended instead of removing them automatically.
| Character | Code point | Typical use | Default in CleanGlyph |
|---|---|---|---|
| ZWSPZero Width Space | U+200B | An invisible separator that can provide a word-breaking opportunity. Often unwanted in copied data or code. | Default: Remove |
| ZWJZero Width Joiner | U+200D | Joins emoji sequences (👩💻) and shapes scripts like Devanagari. Often legitimate. | Default: Review first |
| ZWNJZero Width Non-Joiner | U+200C | Prevents letters from joining in Persian, Arabic and Indic scripts. Often legitimate. | Default: Review first |
| NBSPNon-Breaking Space | U+00A0 | A space that prevents a line break. Common from web pages and word processors. | Default: Replace with space |
| BOMByte Order Mark / Zero Width No-Break Space | U+FEFF | Marks file encoding at the start of a file. Breaks CSV headers and JSON when copied. | Default: Remove |
| SHYSoft Hyphen | U+00AD | An optional hyphenation point, frequent in PDF and e-book text. | Default: Remove |
| WJWord Joiner | U+2060 | Prevents a line break at its position without adding visible space. | Default: Review first |
| MVSMongolian Vowel Separator | U+180E | A Mongolian script control, rarely intended outside Mongolian text. | Default: Review first |
Why do invisible characters cause problems?
Because they are hard to see, hidden characters usually surface as confusing symptoms rather than obvious errors. Typical places they cause trouble:
Copied PDF text
PDF extraction often inserts soft hyphens, non-breaking spaces and zero-width characters where lines were wrapped.
Spreadsheets & Excel matching
VLOOKUP, XLOOKUP and MATCH fail when one cell holds a hidden character that the other does not.
JSON & code
Zero-width characters inside keys, identifiers or strings can cause parse errors and bugs that are very hard to spot.
CSV data
A BOM at the start of a file can end up inside the first column header, breaking imports and column lookups.
How to remove invisible characters
- 1
Paste
Paste text into the input or use the Paste button. Analysis starts immediately in your browser.
- 2
Inspect
Hidden characters appear as labelled tokens in the inspector. Select one to see its name, code point and position.
- 3
Review
Check the detected list. Choose Remove, Replace or Keep for each type — joiners are kept unless you decide otherwise.
- 4
Clean
Press Clean text to apply your choices. Only the characters you selected are changed.
- 5
Copy
Copy the cleaned text and compare the before and after Unicode code-point counts.
Frequently asked questions
What is a zero-width space?
A zero-width space (U+200B) is a Unicode character that takes up no visible width. It tells software where a line may break, but in copied text it usually ends up as an unwanted hidden character inside words, numbers or identifiers.
How can I find invisible characters in text?
Paste the text into the detector above. CleanGlyph scans every character and shows hidden ones as labelled tokens such as [ZWSP] or [NBSP], together with their code point and exact position.
How do I remove zero-width characters?
After pasting your text, check the cleaning options, then choose “Clean text”. Zero-width spaces, word joiners and byte order marks are selected for removal by default. Copy the cleaned result when you are done.
What is U+200B?
U+200B is the code point of the Zero Width Space. “U+” means Unicode and 200B is the hexadecimal number that identifies the character.
What is a non-breaking space?
A non-breaking space (U+00A0) looks identical to a normal space but prevents a line break between the words it separates. It is common in text copied from web pages and word processors and often causes failed matches in spreadsheets and code. CleanGlyph replaces it with a normal space by default.
Why does copied text contain hidden characters?
Websites, PDFs, word processors, chat apps and text-generation tools use Unicode formatting characters for layout, hyphenation, emoji and right-to-left text. When you copy text, those characters are often copied too, even though you cannot see them.
Are invisible Unicode characters dangerous?
Most are harmless but annoying. They can, however, break data matching, create look-alike usernames, or hide content. Bidirectional override characters can make source code or file names appear different from what they really are, so they are worth removing from code and identifiers.
Does CleanGlyph upload my text?
No. All detection and cleaning runs in your browser using JavaScript on your device. Your text is never sent to a server, saved or logged by CleanGlyph.
Should I remove ZWJ and ZWNJ characters?
Not automatically. The Zero Width Joiner builds emoji such as 👩💻 and the Zero Width Non-Joiner is required for correct spelling in Persian and several other languages. Remove them from identifiers, data and code, but keep them in natural-language text that needs them.