Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Python UTF-8 imports: reject malformed bytes and unexpected BOMs

Last updated: 30 Sept 20264 min read
tutorial
IntermediateBy AITrove Editorial

Strict UTF-8 decoding raises UnicodeDecodeError when an incoming byte sequence is malformed.

Download Python source kit

Operation contract

The receipt boundary caps the raw payload before decoding, rejects a leading UTF-8 byte-order mark, and decodes with the default strict error handler. A normal accented customer name survives unchanged. A malformed prefix is rejected instead of being replaced with a character that might change an identifier or signature input.

Failure and ownership boundary

Using errors="ignore" or "replace" changes received data. If a format explicitly permits a BOM, declare that protocol and decode with an intentional policy; do not guess encodings. Normalization and text identity remain separate. Python strings and bytes: reject decoding errors before parsing records, Python Unicode normalization: equality is not visual identity and Python JSON duplicate keys: reject ambiguous object fields connect the layers.

Working program

python
def decode_receipt(raw):
    if type(raw) is not bytes or len(raw) > 80:
        raise ValueError("byte budget")
    if raw.startswith(b"\xef\xbb\xbf"):
        raise ValueError("BOM not permitted")
    return raw.decode("utf-8")

print(decode_receipt("Café R41".encode("utf-8")))
for payload in (b"\xc3(", b"\xef\xbb\xbfR41"):
    try:
        decode_receipt(payload)
    except (UnicodeDecodeError, ValueError):
        print("bytes rejected")

Output

Output
Café R41
bytes rejected
bytes rejected

Costs and limits

Validation and decoding take O(n) time and allocate O(n) text for n bytes. The limit is applied to already received bytes; an HTTP server must cap reception upstream.

Common Mistakes

  • A replacement character is a data change, not validation.
  • Do not confuse an encoding policy with Unicode normalization.

Connected lessons

Python strings and bytes: reject decoding errors before parsing records, Python Unicode normalization: equality is not visual identity, Python JSON duplicate keys: reject ambiguous object fields.

python
strict-utf8-import
Storage details