Unicode Encode / Decode Online — \uXXXX format ↔ Chinese / Emoji
Text to \uXXXX / \uXXXX to text / supports Emoji supplementary characters
About Unicode Encoding Conversion
Unicode is the international character-encoding standard (Unicode 15.1 encodes about 150,000 characters), assigning a unique code point to each character. This tool converts both ways between text and Unicode escape sequences: BMP characters (U+0000–U+FFFF, including Chinese, English and punctuation) become \uXXXX (4 hex digits), while supplementary characters (above U+10000, including Emoji and CJK Extension characters) become \UXXXXXXXX (8 hex digits). It uses codePointAt() internally to handle surrogate pairs correctly, so Emoji are not split into two \u escapes. For URL percent-encoding, use the URL encoder/decoder.
Use Cases
- JavaScript string handling: embed Unicode escapes in JS code (e.g.
'\u4F60\u597D'outputs "你好") to handle multilingual string literals. - CSS icons and iconfont: custom icon fonts use codes such as
content: '\e600'in CSS; this tool quickly finds the code for a character. - JSON data transfer: the JSON standard allows
\uXXXXescapes, ensuring non-ASCII characters such as Chinese transfer correctly between systems. - Emoji encoding: most Emoji are supplementary characters (e.g. 😀 = U+1F600) requiring 4-byte UTF-16 encoding; this tool handles them with the
\UXXXXXXXXformat.
FAQ
What is the difference between \uXXXX and \UXXXXXXXX?
\uXXXX is for BMP (Basic Multilingual Plane) characters, code points U+0000–U+FFFF, written as 4 hex digits and covering common characters such as Chinese, English and digits. \UXXXXXXXX is for supplementary characters (U+10000 and above), written as 8 hex digits, mainly used for Emoji (e.g. 😀 = U+1F600). Both are Unicode escape sequences and differ only in digit count.
Why do some Emoji become two \u escapes?
This is caused by surrogate pairs in UTF-16 encoding. If processed character by character with charAt(), an Emoji is split into a high surrogate (0xD800–0xDBFF) and a low surrogate (0xDC00–0xDFFF) — two code units. This tool uses codePointAt() to obtain the full code point, so each Emoji outputs a single \UXXXXXXXX instead of two incomplete \uXXXX.
How are Unicode and UTF-8 related?
Unicode is the character-set standard — it assigns each character a unique number (code point); "中", for example, is U+4E2D. UTF-8 is an encoding scheme — it defines how code points are turned into byte sequences for storage or transfer. The same Unicode code point can be represented by different encodings (UTF-8, UTF-16, UTF-32). This tool works on JavaScript strings (UTF-16 encoded); the \uXXXX it outputs is the code point value, not UTF-8 bytes.
What if decoding \u4F60\u597D fails?
Check that the format is strictly correct: ① the \u or \U prefix must be lowercase (\u cannot be written as \U or in mixed case); ② the number of hex digits must be right (\u takes 4, \U takes 5–8); ③ letters A-F may be upper or lower case. You can also cross-check the escapes with the JSON formatter.