A str contains Unicode text, while bytes contains integer byte values that require an encoding contract before text parsing.
Python strings and bytes: reject decoding errors before parsing records
Operation contract
The import decodes UTF-8 with strict error handling and prints a count of code points for the decoded amount marker. A truncated byte sequence is rejected before field parsing. Character count and encoded byte count differ, so a byte-size limit must be enforced before accepting a large transmitted payload.
Failure and ownership boundary
Replacement decoding can conceal a damaged record by inserting replacement characters. Code-point indexes also do not identify every user-visible grapheme cluster; combining marks can occupy more than one index. Python JSON validation: reject duplicate members and non-integer amounts follows decoding, and Java grapheme boundaries: truncate display text without splitting a cluster describe a separate concern.
Working program
def decode_record(payload):
if len(payload) > 4096:
raise ValueError("record exceeds byte limit")
return payload.decode("utf-8", errors="strict")
record = decode_record("₹125".encode("utf-8"))
print(record)
print(len(record))
try:
decode_record(b"\xe2\x82")
except UnicodeDecodeError:
print("malformed UTF-8 rejected")Output
₹125
4
malformed UTF-8 rejectedCosts and limits
Decoding b bounded bytes takes O(b) time and text storage in the worst case. The 4096-byte cap is this record contract, not a universal HTTP body limit.
Common Mistakes
- Specify the encoding instead of relying on the machine default.
- Byte counts and visible-character counts are different.
Connected lessons
Python JSON validation: reject duplicate members and non-integer amounts, Python pathlib files: specify encoding and close the resource owner, Java UTF-8 decoding: reject malformed bytes before parsing.
Apply this boundary
Python Unicode normalization: equality is not visual identity, Python prefix-function search: overlapping matches without rescanning.
Related Python operation checks
Python memoryview: shared buffers and explicit release, Django templates: escape untrusted text at HTML output.
Follow the service contract
Python struct: a fixed-endian receipt record with exact byte length, Python strict Base64: reject alternate spellings before decoding a record.
Follow the integrity boundary
Python UTF-8 imports: reject malformed bytes and unexpected BOMs.
