Strict UTF-8 decoding raises UnicodeDecodeError when an incoming byte sequence is malformed.
Python UTF-8 imports: reject malformed bytes and unexpected BOMs
Operation contract
The receipt boundary caps the raw payload before decoding, rejects a leading UTF-8 byte-order mark, and decodes with the default strict error handler. A normal accented customer name survives unchanged. A malformed prefix is rejected instead of being replaced with a character that might change an identifier or signature input.
Failure and ownership boundary
Using errors="ignore" or "replace" changes received data. If a format explicitly permits a BOM, declare that protocol and decode with an intentional policy; do not guess encodings. Normalization and text identity remain separate. Python strings and bytes: reject decoding errors before parsing records, Python Unicode normalization: equality is not visual identity and Python JSON duplicate keys: reject ambiguous object fields connect the layers.
Working program
def decode_receipt(raw):
if type(raw) is not bytes or len(raw) > 80:
raise ValueError("byte budget")
if raw.startswith(b"\xef\xbb\xbf"):
raise ValueError("BOM not permitted")
return raw.decode("utf-8")
print(decode_receipt("Café R41".encode("utf-8")))
for payload in (b"\xc3(", b"\xef\xbb\xbfR41"):
try:
decode_receipt(payload)
except (UnicodeDecodeError, ValueError):
print("bytes rejected")Output
Café R41
bytes rejected
bytes rejectedCosts and limits
Validation and decoding take O(n) time and allocate O(n) text for n bytes. The limit is applied to already received bytes; an HTTP server must cap reception upstream.
Common Mistakes
- A replacement character is a data change, not validation.
- Do not confuse an encoding policy with Unicode normalization.
Connected lessons
Python strings and bytes: reject decoding errors before parsing records, Python Unicode normalization: equality is not visual identity, Python JSON duplicate keys: reject ambiguous object fields.
