Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Python tar extraction: reject members before exceeding a byte quota

Last updated: 1 Oct 20265 min read
tutorial
IntermediateBy AITrove Editorial

An archive importer needs a limit on files and uncompressed bytes before writing each member.

Download Python source kit

Operation contract

Two regular report members total 54 bytes, while this import accepts at most 50. The first file is written to a fresh directory; the second is rejected before opening a destination. Each member must stay beneath the reports folder, contain no parent component, and be a regular file. The code copies through a small chunk buffer rather than calling extractall.

Failure and ownership boundary

The limit uses declared member sizes; a production importer should also count bytes actually read and written, reject truncated content, cap member count and name length, and remove the temporary directory on any failure. Concurrent changes to the source or destination need stronger isolation. This local archive is fixed and trusted for test construction.

Working program

python
import io
import tarfile
import tempfile
from pathlib import Path, PurePosixPath

archive_bytes = io.BytesIO()
with tarfile.open(fileobj=archive_bytes, mode="w") as archive:
    for name, payload in (
        ("reports/receipt-47.txt", b"R" * 47),
        ("reports/receipt-48.txt", b"S" * 7),
    ):
        member = tarfile.TarInfo(name)
        member.size = len(payload)
        archive.addfile(member, io.BytesIO(payload))

archive_bytes.seek(0)
accepted = []
rejected = 0
total_bytes = 0
with tempfile.TemporaryDirectory() as directory:
    with tarfile.open(fileobj=archive_bytes, mode="r:") as archive:
        for member in archive:
            name = PurePosixPath(member.name)
            if (not member.isfile() or not name.parts or name.is_absolute()
                    or ".." in name.parts or name.parts[0] != "reports"
                    or total_bytes + member.size > 50):
                rejected += 1
                continue
            destination = Path(directory).joinpath(*name.parts)
            destination.parent.mkdir(parents=True, exist_ok=True)
            with archive.extractfile(member) as source, destination.open("xb") as target:
                remaining = member.size
                while remaining:
                    chunk = source.read(min(16, remaining))
                    if not chunk:
                        raise EOFError("truncated archive member")
                    target.write(chunk)
                    remaining -= len(chunk)
            total_bytes += member.size
            accepted.append(member.name)
print("accepted", accepted)
print("rejected", rejected)
print("bytes", total_bytes)

Output

Output
accepted ['reports/receipt-47.txt']
rejected 1
bytes 47

Costs and limits

The importer scans O(member count + extracted bytes), keeps a 16-byte application copy buffer, and writes up to the stated quota in this fixture.

Common Mistakes

  • The tar header's size is not a complete resource limit by itself.
  • A safe filename check does not authorize links or special files.
  • Duplicate names and case-insensitive collisions need an explicit policy.

Connected lessons

Test this boundary.

python
tar-extraction-byte-quota
Storage details