A pipe reader can reject output as soon as it consumes one byte past the allowed limit.
Python subprocess pipe: stop after the first byte beyond budget
Operation contract
A fixed child writes 47,000 bytes. The parent reads only 4,097 bytes against a 4,096-byte contract, then kills and reaps the child. Unlike a post-run spool check, it does not wait for a full output file before rejecting. The one extra byte distinguishes exactly-at-limit from over-limit output.
Failure and ownership boundary
This is one boundary, not a full untrusted-process runner. A child that emits fewer bytes and stalls can block read, so a separate wall-clock deadline is required. stderr is discarded here; a production runner needs a separate bounded diagnostic channel. The direct child is stopped, but a process tree requires an OS-specific group policy.
Working program
import subprocess
import sys
child = subprocess.Popen(
[sys.executable, "-c", "import sys; sys.stdout.buffer.write(b'R' * 47000)"],
stdout=subprocess.PIPE,
stderr=subprocess.DEVNULL,
)
prefix = child.stdout.read(4097)
too_large = len(prefix) > 4096
if too_large:
child.kill()
child.stdout.close()
child.wait(timeout=5)
print("too_large", too_large)
print("reaped", child.returncode is not None)Output
too_large True
reaped TrueCosts and limits
The decision retains O(limit) bytes and stops reading after limit plus one. Kernel pipe and buffered-reader capacity add bounded buffering; a separate timeout is still required.
Common Mistakes
- A byte cutoff is not a wall-clock cutoff.
- Closing stdout without reaping leaves process ownership unresolved.
- A discarded stderr pipe cannot later supply diagnostics.
