A code point is the abstract number assigned to a character (e.g. U+00E9 for 'é'). UTF-8 is a specific way of encoding that number as bytes — non-ASCII characters typically expand to 2-4 bytes.
Convert text to Unicode code points, UTF-8, UTF-16, UTF-32, percent-encoding, Base64, and decimal — or decode any of those formats back to text. Correctly handles emoji, CJK, and other multi-byte characters.
Encode text to all formats, or decode one format back to text.
Type or paste text, or paste an encoded string with its format selected.
See every format simultaneously, or the decoded plain text.
Copy any individual format's output with one click.
Use our Invisible Character Viewer to find hidden zero-width spaces, BOM, and smart-quote look-alikes.
Open Invisible Character Viewer →| Format | Example (for "A") | What it represents |
|---|---|---|
| Unicode | U+0041 | The character's Unicode code point — a unique number assigned to every character in the Unicode standard. |
| UTF-8 | 41 | The character encoded as UTF-8 bytes. ASCII characters are 1 byte; most other characters use 2-4 bytes. |
| UTF-16 | 0041 | The character encoded as UTF-16 code units. Characters outside the Basic Multilingual Plane (like most emoji) use two 16-bit "surrogate pair" units. |
| UTF-32 | 00000041 | The character's code point padded to exactly 4 bytes (32 bits) — fixed-width, simple but memory-inefficient. |
| Percent | %41 | URL/percent-encoding of the UTF-8 bytes — used in URLs, query strings, and form submissions. |
| Base64 | QQ== | The UTF-8 bytes encoded as Base64 — used for embedding binary-safe text in JSON, email, and data URIs. |
| Decimal (code point) | 65 | The Unicode code point expressed as a decimal number instead of hex. |
| Decimal (UTF-8 byte) | 65 | Each UTF-8 byte expressed as a decimal number — differs from the code point for non-ASCII characters. |
For ASCII characters (code points 0-127), the UTF-8 byte value and the decimal code point are identical. For any character above U+007F, UTF-8 splits the code point across 2-4 bytes, so the decimal byte sequence no longer matches the single code point number.
UTF-16 can only directly represent code points up to U+FFFF. Characters beyond that — including most emoji (U+1F300 and above) — are represented as a pair of 16-bit "surrogate" code units that combine to form the full code point. This tool handles surrogate pairs correctly in both directions.
A code point is the abstract number assigned to a character (e.g. U+00E9 for 'é'). UTF-8 is a specific way of encoding that number as bytes — non-ASCII characters typically expand to 2-4 bytes.
Most emoji require UTF-16 surrogate pairs. This tool iterates by Unicode code point rather than raw character index, so multi-unit characters are never split incorrectly.
No. All conversion happens entirely in your browser. Nothing is sent to any server.