Unicode normalization maps canonically equivalent character sequences to a selected normal form; it does not define a complete user-name or security policy.
Java text normalization: equal labels and explicit comparison policy
Java 8+. This is a complete program using JDK classes.
Separate stored text from comparison keys
A customer name can arrive as a precomposed accented character or as a base character followed by a combining mark. String.equals compares their UTF-16 sequences. NFC normalization gives the application a common representation for this particular equivalence before it uses a value as a map key.
Keep the display value. Collapsing all case, punctuation, or accents can destroy distinctions a person expects to see. Canonical normalization is different from removing accents, and compatibility forms such as NFKC can fold extra character distinctions. Choose the form for the field, document it, and apply it on every write and lookup.
The program uses codePointCount only to report code points. That count is not the number of displayed characters in every script. Emoji sequences and combining marks can make one displayed unit contain several code points. A UI truncation rule needs a segmentation policy rather than a substring index guessed from the count.
Do not hide locale decisions
A normalized spelling can still differ in case. Locale-sensitive sorting belongs to a separate comparison operation. Identifiers in a machine protocol usually need a restricted alphabet and a fixed comparison rule; a human directory may instead require language-specific collation. Neither follows automatically from using UTF-8.
Working program
import java.text.Normalizer;
public class CustomerLabelNormalization {
public static void main(String[] args) {
String composed = "Caf\u00e9";
String decomposed = "Cafe\u0301";
System.out.println(composed.equals(decomposed));
String normalized = Normalizer.normalize(decomposed, Normalizer.Form.NFC);
System.out.println(composed.equals(normalized));
System.out.println(decomposed.codePointCount(0, decomposed.length()));
System.out.println(normalized.codePointCount(0, normalized.length()));
}
}Output
false
true
5
4Costs and boundaries
Normalization must inspect the input and creates an output representation. Bound accepted input size and account for allocation in bulk imports. This example verifies one canonical equivalence; it is not a comprehensive grapheme segmenter or a confusable-character detector.
Common Mistakes
- Do not describe code points as bytes or displayed characters.
- Apply the same key policy at insertion and lookup.
- Normalization does not sanitize a name for HTML, SQL or a file path.
