Invisible Character Remover
β‘ Runs locallyπ No uploadsπ Free, no sign-up
You copy a paragraph from a website, paste it into a search box, and get zero results. Or a snippet that looks perfectly fine breaks your code editor with a cryptic error, and no amount of staring reveals anything wrong. The culprit is usually invisible characters β Unicode includes a family of glyph-less, zero-width characters such as the zero-width space (U+200B), the byte order mark (U+FEFF), and the soft hyphen (U+00AD). They render as nothing on screen, yet they occupy real positions in the string and take part in comparison, counting, and parsing.
These characters arrive through two main routes. The first is accidental: copying text from web pages, PDFs, or Word documents drags along soft hyphens and bidirectional controls used for layout, and Windows Notepad prepends a BOM to UTF-8 files, which becomes an illegal character the moment it lands inside code. The second is deliberate: many content platforms silently inject zero-width characters as a watermark when you copy text, encoding a unique fingerprint into the arrangement of invisible characters to track where their content gets republished. If your article was plagiarized, the original author may have found you through exactly these hidden watermarks.
There are three typical use cases. First, sanitizing copied text: when you lift long passages from the web for a paper, a marketing draft, or research notes, run them through the remover first β it strips invisible characters that would pollute your document and stops you from unknowingly carrying someone else's tracking watermark into your own content. Second, debugging bizarre failures: JSON that refuses to parse, a password that is 'definitely correct' yet rejected at login, a CSV import with shifted columns β paste the suspicious text here and an invisible character is the culprit more often than you would expect. Third, pre-publication self-checks: scan your draft before publishing to confirm it carries no platform-injected tracking watermark, keeping your content untraceable.
A few edge cases deserve attention. The zero-width joiner (U+200D) is not purely malicious: it is a required shaping control in scripts like Arabic and Devanagari, and more visibly, compound emoji such as families and professions (e.g. the family emoji) are built by joining individual pictographs with zero-width joiners β remove them and the compound emoji shatters into separate figures. The tool selects all categories by default, but you can toggle each category individually; remember to uncheck 'Zero-width joiner' when cleaning text that contains emoji. Similarly, a BOM (U+FEFF) at the very start of a file is a legitimate encoding signature and only anomalous mid-text, and soft hyphens are legitimate line-breaking hints in languages like German. The tool performs mechanical find-and-delete without judging semantics, so back up important text before cleaning.
All detection and cleaning happens locally in your browser β your text is never uploaded to any server, so you can safely paste content containing sensitive information. The detection report lists the count and code points of every invisible character category in real time, and after cleaning it shows exactly how many characters were removed, so the whole process is transparent and verifiable.
How to use
- Paste or type the text you want to inspect. The detection report updates live with per-category counts.
- Under 'Choose categories to remove', tick the character types to delete (all selected by default; uncheck 'Zero-width joiner' for text containing emoji).
- Click 'Clean now'. The cleaned text appears in the box below, ready to copy.
FAQ
Can cleaning delete my normal text by accident?
No. The tool only removes eight well-defined categories of invisible format characters (zero-width spaces, joiners/non-joiners, BOM, soft hyphen, word joiner, bidirectional controls, and similar). All visible characters, spaces, line breaks, and punctuation are preserved untouched. The one exception is emoji sequences built on zero-width joiners β deselect that category before cleaning such text.
Why does text copied from a website fail to match in search?
Most likely it picked up zero-width spaces or a platform-injected invisible watermark during copying. These characters are invisible to the eye, but search engines and string comparison treat them as real characters, so 'looks identical' becomes 'not equal' to the machine. Detect and clean the text with this tool, then search again.
What is a zero-width character watermark?
A stealth tracking technique: different arrangements of zero-width spaces, joiners, and similar characters encode a fingerprint into the text that is completely invisible to readers but machine-readable. Content platforms use it to trace where copied articles get republished. Paste your text here β a detection report full of zero-width characters means it carries a watermark.
What is a BOM and why does it cause errors?
The byte order mark (U+FEFF) is a legitimate encoding signature at the start of a Unicode file. But when it appears mid-text β for example after copying content out of Notepad into code or JSON β parsers treat it as an illegal character, breaking JSON parsing or scripts. Tick 'Byte order mark (BOM)' when cleaning to remove it.
Can I undo the cleaning?
Removal is irreversible: deleted invisible characters cannot be recovered from the cleaned result. Keep a copy of the original before cleaning. The tool never stores your input β everything disappears when you refresh the page.