Unicode Inspector
See every code point in your text with its Unicode name, category, script, UTF-8 and UTF-16 encoding, normalization and exact position — including invisible characters and emoji sequences.
Your text never leaves your browser.
Inspector
Paste text to inspect every Unicode code point.
You will see names, categories, scripts, encodings, normalization details and the exact position of each character.
What is a Unicode code point?
Every character in Unicode has a number called a code point, written as U+ followed by hexadecimal digits. The inspector shows that number together with the character's official name, general category and script, so you can tell exactly what is in your text instead of guessing from how it looks.
Code points, graphemes and UTF-16 units
What you see as one character is a grapheme. It may contain several code points: e plus a combining accent, a Devanagari conjunct such as क्ष, or an emoji joined with zero width joiners. Click a character in the inspector, then select each of its code points.
JavaScript and many editors count UTF-16 code units instead. Characters above U+FFFF, including most emoji, take two units, which is why positions are shown both as code-point positions and as UTF-16 offsets.
UTF-8 vs UTF-16
UTF-8 stores a code point in one to four bytes and is used by most files, APIs and web pages. UTF-16 uses one or two 16-bit units. The inspector shows both encodings and the JavaScript and HTML escapes for each code point, ready to copy.
Normalization: NFC and NFD
The same accented letter can be written precomposed (NFC, e.g. U+00E9) or decomposed (NFD, U+0065 U+0301). They render identically but compare as different strings, which breaks searches and deduplication. When a character has more than one canonical form, the inspector shows both sequences and whether your text is already in NFC.
Emoji sequences
Many emoji are built from several code points: skin-tone modifiers, pairs of regional indicators for flags, variation selectors and zero width joiners. These are legitimate. The inspector keeps the emoji intact in the text view and lists each part separately.
Invisible and format characters
Zero-width spaces, byte order marks, bidirectional controls and other format characters have no visible glyph. The inspector labels them inline. Some, like the zero width non-joiner in Persian or the joiner in emoji, are required in context and are marked for review rather than removal. To clean text, use the Invisible Character Detector or the Zero Width Space Remover. To find where two similar texts differ, use the Text Difference Checker.
Common examples
- U+0041LATIN CAPITAL LETTER APlain ASCII letter: one byte in UTF-8.
- U+00E9LATIN SMALL LETTER E WITH ACUTEPrecomposed é, the NFC form.
- U+0301COMBINING ACUTE ACCENTAttaches to the previous letter, as in NFD é.
- U+00A0NO-BREAK SPACELooks like a space but prevents line breaks.
- U+200BZERO WIDTH SPACEInvisible break opportunity, common in copied text.
- U+200CZERO WIDTH NON-JOINERRequired in Persian and other scripts.
- U+200DZERO WIDTH JOINERJoins emoji into sequences.
- U+1F600GRINNING FACEOutside the BMP: a UTF-16 surrogate pair.
| Code point | Name | Why it matters |
|---|---|---|
| U+0041 | LATIN CAPITAL LETTER A | Plain ASCII letter: one byte in UTF-8. |
| U+00E9 | LATIN SMALL LETTER E WITH ACUTE | Precomposed é, the NFC form. |
| U+0301 | COMBINING ACUTE ACCENT | Attaches to the previous letter, as in NFD é. |
| U+00A0 | NO-BREAK SPACE | Looks like a space but prevents line breaks. |
| U+200B | ZERO WIDTH SPACE | Invisible break opportunity, common in copied text. |
| U+200C | ZERO WIDTH NON-JOINER | Required in Persian and other scripts. |
| U+200D | ZERO WIDTH JOINER | Joins emoji into sequences. |
| U+1F600 | GRINNING FACE | Outside the BMP: a UTF-16 surrogate pair. |
Frequently asked questions
What is a Unicode code point?
A code point is the number Unicode assigns to a character, written as U+ followed by hexadecimal digits. For example, U+0041 is the Latin capital letter A and U+200B is the zero width space.
What is the difference between a code point and a grapheme?
A grapheme is what a reader sees as one character. It can be made of several code points: an accented letter written as e plus a combining accent, a Devanagari conjunct, or an emoji sequence joined with zero width joiners.
Why does one emoji contain several code points?
Many emoji are sequences. Skin tones are added with a modifier code point, flags use two regional indicators, and emoji such as a woman technologist combine several emoji with U+200D zero width joiners. The inspector lists every part.
What is the difference between UTF-8 and UTF-16?
They are two ways to store the same code points as bytes. UTF-8 uses one to four bytes per code point and is common in files and on the web. UTF-16 uses one or two 16-bit code units and is what JavaScript strings use internally.
What is a surrogate pair?
Code points above U+FFFF do not fit in one UTF-16 code unit, so they are stored as two: a high surrogate followed by a low surrogate. That is why an emoji has a string length of 2 in JavaScript.
What are NFC and NFD?
They are Unicode normalization forms. NFC prefers precomposed characters such as U+00E9 for é, while NFD splits them into a base letter and combining marks. Both look identical, but they compare as different text unless normalized.
Why do two identical-looking characters compare differently?
They can be different code points that look alike, such as Latin a and Cyrillic а, or the same letter in different normalization forms. Inspecting each code point shows exactly which characters are present.
Is my text uploaded?
No. Every character is analysed locally in your browser. Your text is never sent to a server or stored.