Unicode normalization maps equivalent code-point sequences into a selected normalization form; Python string equality otherwise compares their code points.
Python Unicode normalization: equality is not visual identity
Operation contract
The display label uses one precomposed letter in one record and a letter plus combining accent in another. NFC normalizes both into the same canonical form for this comparison. The program keeps display text separate from an identifier policy: deciding that two labels compare equally is not permission to merge two customer records.
Failure and ownership boundary
Normalization does not remove every visually confusable character and does not count grapheme clusters. A normalized string can still include mixed scripts, controls or misleading labels. Use an explicit identifier alphabet when the protocol allows one, and enforce length after transformations as well as before them. Python regular expressions: use fullmatch for a complete field contract and Python strings and bytes: reject decoding errors before parsing records address different layers.
Working program
import unicodedata
stored_label = "Caf\u00e9"
imported_label = "Cafe\u0301"
print(stored_label == imported_label)
print(len(stored_label), len(imported_label))
print(unicodedata.normalize("NFC", stored_label) ==
unicodedata.normalize("NFC", imported_label))Output
False
4 5
TrueCosts and limits
Normalization scans input and allocates transformed text; work and storage depend on code points and combining sequences. len(str) counts code points, not visible symbols or UTF-8 bytes.
Common Mistakes
- Normalization is not a confusable-character detector.
- Never silently merge accounts merely because their display labels normalize equally.
Connected lessons
Python strings and bytes: reject decoding errors before parsing records, Python regular expressions: use fullmatch for a complete field contract, Pandas joins: cardinality validation and missing lookup keys.
