Localization engineer
See why an accented glyph or emoji has multiple components.
Encode emoji plus decomposed e-accent and compare scalar, UTF-16 and UTF-8 columns.
Combining marks and offsets are visible without normalization.
Inspect Unicode scalars, UTF-16 units and UTF-8 bytes. Encode text as U+ tokens or Unicode escapes, preserve hidden characters, and export complete notation.
A quick decision brief for this specific tool
Text source
Encoded notation
A dedicated interactive panel for this data format or operation
Choose your path
Inspect exact Unicode scalars and encode them as U+ notation, braced escapes or paired UTF-16 escapes without conflating glyphs, units and bytes.
See why an accented glyph or emoji has multiple components.
Encode emoji plus decomposed e-accent and compare scalar, UTF-16 and UTF-8 columns.
Combining marks and offsets are visible without normalization.
Choose the notation expected by a receiving system.
Compare U+ tokens, braced escapes and fixed surrogate pairs; choose case and separators.
Astral characters remain valid pairs or scalar escapes.
Find hidden characters without losing the complete source.
Import BOM/CRLF text, inspect labeled controls, page many scalars and restore a prior run.
Invisible data is labeled and complete output remains recoverable.
Outputs and checklists are planning aids. Review the linked current authorities and the records, terms, instructions, and requirements that apply to your exact situation before a consequential decision.
Tools you might need next
Decode complete U+ tokens or Unicode escapes into text. Validate scalars and surrogate pairs, inspect hidden characters, and refuse silently discarded input.
Decode binary bytes with strict UTF-8 or ASCII validation. Choose token grouping and BOM handling, inspect every byte, and export the complete decoded text.
Decode complete Morse tokens with explicit slash or line-break word boundaries. Refuse unknown signals, inspect character rows, and export the full transcript.
Validate paired UTF-16 source, iterate complete code points and retain combining marks, format characters and line endings.
U+ uses 4–6 hexadecimal digits. Braced escapes use scalar values; fixed escapes emit adjacent surrogate pairs for astral characters. Hex case and token separators are explicit.
Each review row is a scalar, not a grapheme cluster. It shows a visible label, U+ value, UTF-16 units, UTF-8 bytes and original UTF-16 offset.
Updated: August 2026
See why an accented glyph or emoji has multiple components. Encode emoji plus decomposed e-accent and compare scalar, UTF-16 and UTF-8 columns.
Choose the notation expected by a receiving system. Compare U+ tokens, braced escapes and fixed surrogate pairs; choose case and separators.
Find hidden characters without losing the complete source. Import BOM/CRLF text, inspect labeled controls, page many scalars and restore a prior run.
Check the individual components; a grapheme can contain several code points.
Choose the receiver's accepted escape format and add any required string-literal framing separately.
Make the structure of multilingual text visible. Select a notation and compare scalars, UTF-16 offsets and UTF-8 bytes without normalizing or trimming your source.