Unicode text repair — broken characters and invisible ones
Text breaks for a small number of specific reasons: the encoding was read wrongly, the same character was stored two different ways, or something invisible crept in. Each has a different fix, so the first job is working out which one you have.
This tool handles three problems. Text read with the wrong encoding, where 안녕하세요 arrives as ìëíì¸ì. Text that looks identical on screen but is stored differently. And invisible or unusual spaces mixed into otherwise clean text.
None of the three can be diagnosed by eye, so the tool shows what it changed and why rather than silently rewriting.
Mojibake can be undone
In UTF-8 a Korean or Japanese character is three bytes. Read those bytes one at a time as Western European characters and you get a string of letters and symbols — the effect known as mojibake.
Nothing has been deleted; the bytes were merely read wrongly. Turning them back into bytes and decoding again recovers the original. A CSV opened in a spreadsheet, filenames inside a zip archive, and database exports are the usual sources.
The catch is that this must not be applied indiscriminately. café and Grüße are correct text that would be destroyed by the same operation. This tool applies the repair only when the result contains East Asian characters that were not there before; otherwise it leaves the text alone.
What comes back
| Broken | Repaired |
|---|---|
| ìëíì¸ì | 안녕하세요 |
| ìì¸ ë§ì§ | 서울 맛집 |
| ããã«ã¡ã¯ | こんにちは |
| café | unchanged — already correct |
| Grüße | unchanged — already correct |
Characters that look the same but are not
The same word typed on a Mac and on Windows can be stored as six code points or as two, because macOS writes text decomposed into its parts. On screen they are identical; in a search box or a filename comparison they are different strings.
There are related traps in individual letters. A Korean consonant typed on its own (U+3131) is a different character from the one used as a building block inside a syllable (U+1100), even though they draw the same shape.
This tool normalizes text to a single standard form, so that text from any machine compares as equal when it reads as equal.
Invisible characters
Text copied from the web brings characters with no visible mark: a zero-width space (U+200B), a non-breaking space (U+00A0), and the full-width space used in East Asian typesetting (U+3000).
They look like an ordinary space or like nothing at all, but they count when searching. Copying a sentence and finding no results is usually one of these. In code they cause errors, and in a spreadsheet they turn numbers into text.
When to reach for this
- •A CSV opened in a spreadsheet shows scrambled characters.
- •Filenames inside an extracted archive are unreadable.
- •A filename from a colleague's Mac cannot be found by search.
- •A phrase copied from a website returns no search results.
- •Pasted numbers are treated as text by a spreadsheet.
Frequently asked questions
Is repair always possible?
No. If the faulty read replaced some bytes with question marks or replacement characters before saving, that information is gone and cannot be recovered. The original file is the only way back.
Why leave some text untouched?
Because it is correct. German and French text resembles mojibake without being it. Requiring that East Asian characters appear in the result is what separates the two.
Does normalizing change the characters?
Not what you see — only how it is stored. Note that a filename normalized here may be decomposed again by a system that prefers the other form.
Is the text stored anywhere?
No. Everything runs in your browser.
