Unicode Encoder / Decoder
Text ⇄ \uXXXX escapes, U+ code points and HTML entities — surrogate-pair aware, with mixed-format decoding.
How to use
- Paste text or encoded tokens into the input box.
- Pick Encode or Decode; when encoding, choose the target notation.
- Copy the result — mixed notations decode in one pass.
Frequently asked questions
Which notations can I convert to?
Four: \uXXXX escapes as JavaScript writes them (lowercase hex, surrogate pairs for emoji), U+ code-point notation as used in Unicode charts, and HTML entities in both hexadecimal (😀) and decimal (😀) form.
Why does one emoji become two \u escapes?
JavaScript strings are UTF-16, and characters outside the Basic Multilingual Plane are stored as a surrogate pair of two 16-bit units — 😀 is \ud83d\ude00. The U+ and HTML formats work with single code points, so the same emoji is one token there.
How does decoding know which format I pasted?
It recognizes all four notations at once and can mix them in one string — "A\u4e2d U+1F600 A" decodes in a single pass. Whitespace between U+ tokens is treated as a separator, and anything that is not an escape is kept as-is.
What breaks the decoder?
Code points above U+10FFFF or a bare surrogate written in U+ form — neither is a legal character. Those show an error rather than a wrong result.
Is my text stored anywhere?
No. The conversion is done with local string operations only; nothing leaves the page.