ZIP member validation separates archive metadata and decompressed byte limits from the filesystem paths used to write extracted files.
Python ZIP extraction: validate names and decompressed budgets before owned writes
Operation contract
The fixture accepts at most eight flat ASCII .txt filenames, rejects duplicate names and nonregular file modes, and caps each declared member at 256 bytes and the batch at 512. It reads each member with a second byte cap before any output file is created. Only then does it write into a freshly created owned temporary directory using exclusive file creation.
Failure and ownership boundary
This intentionally rejects directories, nested paths, backslashes and symlinks instead of normalizing received paths into something accepted. Metadata alone is not a decompression guarantee; actual reads enforce the second budget and let ZIP integrity failures propagate. The archive byte buffer is already bounded/owned by the caller. This is not a race-safe extractor for a shared deployment directory or crash-durable publication. Python pathlib files: specify encoding and close the resource owner and Python file project: stage a report before replacing the visible file cover different limits.
Working program
import io
from pathlib import Path
import re
import stat
import tempfile
import zipfile
def staged_zip_members(payload):
if not isinstance(payload, bytes) or len(payload) > 2048:
raise ValueError("archive byte budget rejected")
staged = {}; total = 0
with zipfile.ZipFile(io.BytesIO(payload)) as archive:
members = archive.infolist()
if len(members) > 8: raise ValueError("member budget rejected")
for member in members:
kind = stat.S_IFMT(member.external_attr >> 16)
if re.fullmatch(r"[a-z][a-z0-9_-]{0,15}\.txt", member.filename) is None or member.filename in staged or kind not in (0, stat.S_IFREG):
raise ValueError("member name or mode rejected")
if member.file_size > 256: raise ValueError("member bytes rejected")
with archive.open(member) as stream:
data = stream.read(257)
total += len(data)
if len(data) > 256 or total > 512: raise ValueError("decompressed budget rejected")
staged[member.filename] = data
return staged
def owned_archive(name, content):
buffer = io.BytesIO()
with zipfile.ZipFile(buffer, "w", zipfile.ZIP_DEFLATED) as archive:
archive.writestr(name, content)
return buffer.getvalue()
members = staged_zip_members(owned_archive("receipts.txt", b"R-0041\n"))
with tempfile.TemporaryDirectory() as directory:
for name, data in members.items():
with (Path(directory) / name).open("xb") as destination: destination.write(data)
print("owned files:", sorted(path.name for path in Path(directory).iterdir()))
for name, content in (("../escape.txt", b"rejected"), ("receipts.txt", b"x" * 257)):
try: staged_zip_members(owned_archive(name, content))
except ValueError: print("archive rejected before writes")Output
owned files: ['receipts.txt']
archive rejected before writes
archive rejected before writesCosts and limits
Archive and decompressed budgets bound this fixture’s retained buffers. Compression processing still has format-dependent work. Entire accepted contents are staged in memory before writes; a larger importer needs streaming staging and a separate publication contract.
Common Mistakes
- Validate actual decompressed bytes as well as advertised sizes.
- A fresh owned directory is different from a shared directory with concurrent symlink changes.
Connected lessons
Python pathlib files: specify encoding and close the resource owner, Python file project: stage a report before replacing the visible file, Python JSON Lines: cap line bytes and validate the whole batch before returning it.
Follow the service contract
Python bounded gzip decoding: cap expanded output before publishing it.
