Gzip decoding expands a compressed byte stream and verifies its format and checksum while reading to its end.
Python bounded gzip decoding: cap expanded output before publishing it
Operation contract
The decoder accepts at most 512 compressed bytes and reads at most 129 expanded bytes for a 128-byte application limit. An oversized expansion is rejected before it is returned. A small accepted expansion reaches end-of-stream so a damaged trailer is not silently ignored. The fixture accepts concatenated gzip members as one bounded output because GzipFile supports that interpretation; a single-member protocol needs an additional framing rule.
Failure and ownership boundary
The output cap limits what this function returns, not a hard bound on decompressor internals, CPU time or already received compressed storage. It never writes a partially decoded file. Python JSON Lines: cap line bytes and validate the whole batch before returning it, Python ZIP extraction: validate names and decompressed budgets before owned writes and Python file project: stage a report before replacing the visible file each add a different boundary after or before decoding.
Working program
import gzip
import io
import zlib
def bounded_gzip(payload):
if type(payload) is not bytes or len(payload) > 512:
raise ValueError("compressed byte budget exceeded")
try:
with gzip.GzipFile(fileobj=io.BytesIO(payload), mode="rb") as decoder:
expanded = decoder.read(129)
except (OSError, EOFError, zlib.error) as failure:
raise ValueError("invalid gzip stream") from failure
if len(expanded) > 128:
raise ValueError("expanded byte budget exceeded")
return expanded
encoded = gzip.compress(b"receipt:125", mtime=0)
print("decoded:", bounded_gzip(encoded).decode("ascii"))
print("members:", bounded_gzip(gzip.compress(b"A", mtime=0) + gzip.compress(b"B", mtime=0)))
for invalid in (gzip.compress(b"x" * 129, mtime=0), encoded[:-1]):
try:
bounded_gzip(invalid)
except ValueError:
print("stream rejected")Output
decoded: receipt:125
members: b'AB'
stream rejected
stream rejectedCosts and limits
Returned output and accepted compressed input have explicit byte bounds. Decoder buffering and decompression work are implementation costs outside the returned-output limit. For hostile workloads, apply transport budgets and execution isolation rather than advertising this helper as a resource sandbox.
Common Mistakes
- Small compressed input can expand beyond an application limit.
- Stopping after a valid prefix may leave trailer corruption unobserved.
Connected lessons
Python JSON Lines: cap line bytes and validate the whole batch before returning it, Python ZIP extraction: validate names and decompressed budgets before owned writes, Python strings and bytes: reject decoding errors before parsing records.
