Localization engineer
Distinguish emoji and combining sequences from UTF-16 units.
Compare two emoji, decomposed accents and ZWJ sequences in scalar and grapheme modes.
Counts and percentages share the selected symbol denominator.
Analyze Unicode scalar or grapheme frequencies with explicit text policies, correct percentages, invisible-character review and complete JSON reports.
A quick decision brief for this specific tool
Text source
Complete JSON report
A dedicated interactive panel for this data format or operation
Choose your path
Count Unicode scalars or grapheme clusters with explicit normalization, case and whitespace policies; inspect every frequency and export the full distribution.
Distinguish emoji and combining sequences from UTF-16 units.
Compare two emoji, decomposed accents and ZWJ sequences in scalar and grapheme modes.
Counts and percentages share the selected symbol denominator.
Know whether normalization and whitespace change a comparison.
Compare NFC, lowercase and whitespace settings while retaining original text.
Every transformation is explicit and no invisible format character silently disappears.
Review many distinct symbols without losing the complete distribution.
Import UTF-8 with BOM/CRLF, page 220 unique symbols, copy/export and undo replacement.
Preview limits do not truncate the report or destroy the source.
Outputs and checklists are planning aids. Review the linked current authorities and the records, terms, instructions, and requirements that apply to your exact situation before a consequential decision.
Tools you might need next
Compare UTF-16 code units, Unicode code points, grapheme clusters, and non-whitespace units for exact text, with explicit platform limits.
Count top NFC and case-normalized word-like tokens with total, unique, deterministic tie, displayed-row, multilingual segmentation, and semantic limitations.
Find duplicate line groups with exact source positions and original values. Choose line, case and whitespace policies and export a complete review report.
Choose Unicode scalars or runtime grapheme clusters. UTF-16 source length is separately labeled; frequency percentages divide by the selected symbol total.
Optional NFC and default lowercase precede whitespace filtering. These are analysis choices, not edits to the source or full linguistic case folding.
Review up to 200 rows, 20 per page. The JSON report includes every result row, settings and methods. Code points identify hidden characters.
Updated: August 2026
Distinguish emoji and combining sequences from UTF-16 units. Compare two emoji, decomposed accents and ZWJ sequences in scalar and grapheme modes.
Know whether normalization and whitespace change a comparison. Compare NFC, lowercase and whitespace settings while retaining original text.
Review many distinct symbols without losing the complete distribution. Import UTF-8 with BOM/CRLF, page 220 unique symbols, copy/export and undo replacement.
Use the selected symbol total and compare separately labeled source units.
Keep exact source data and choose normalization only when the comparison calls for it.
Count Unicode scalars or grapheme clusters with explicit normalization, case and whitespace policies; inspect every frequency and export the full distribution.