On POSIX, fsdecode and fsencode can preserve filename bytes that do not form valid UTF-8 text.
Python POSIX filename bytes: round-trip an undecodable path
Operation contract
A synthetic receipt filename contains one byte outside valid UTF-8. The filesystem decoder represents that byte with a surrogate code point, and the matching encoder recovers the original byte sequence. The program does not create a real file or print the surrogate character into a terminal.
Failure boundary
This is a POSIX filename boundary, not a rule for application text or network JSON. Never store surrogate-containing paths in a format that requires ordinary Unicode scalar values without a deliberate encoding strategy. Windows path handling differs. Display names should have a separate policy from byte-exact filesystem identity.
Working program
import os
raw_filename = b"receipt-\xff"
decoded_filename = os.fsdecode(raw_filename)
print("roundtrip", os.fsencode(decoded_filename) == raw_filename)
print("has_surrogate", any(0xDC80 <= ord(character) <= 0xDCFF
for character in decoded_filename))Output
roundtrip True
has_surrogate TrueCosts and limits
Encoding and decoding each scan the path once and allocate a representation proportional to path length. The fixture keeps only one short name.
Common Mistakes
- Do not treat a surrogate-containing filename as normal user text.
- Do not assume the same byte path behavior on Windows.
- A display-safe label is separate from the path bytes used by the filesystem.
