A ZIP entry name is untrusted input; before writing it, an extractor must establish that its destination remains inside the owned output directory.
Java ZIP inputs: bounded staging and rejected path traversal
This program targets Java 8 without preview flags. Use a full JDK for compiler-based examples.
Normalize a destination, not just a string prefix
A receipt archive can contain a name such as ../outside.txt. Concatenating that name with an output directory can escape the intended location. Resolve the entry against the root, normalize the resulting Path, and compare path components with startsWith; a string prefix comparison can confuse sibling directories.
This fixture rejects absolute names, backslashes, colons, empty names, root-only destinations and paths that normalize outside the root. It stages bytes in memory and does not write them to the file system. Restricting the accepted name grammar makes the lexical policy reproducible across the operating systems used by this example.
Compressed size is not the work budget
The header size may be absent or misleading. Count bytes as they are read and stop when the total uncompressed budget is exceeded. The program permits at most eight files and 64 KiB total output, and rejects duplicate normalized destinations before publication.
A lexical check does not settle symbolic-link races in a pre-existing directory. A production extractor needs an owned private staging directory, a policy for links and permissions, and a publication rule that cannot redirect writes between validation and creation. No general-purpose secure extraction guarantee is made by this in-memory example.
Working program
import java.io.ByteArrayInputStream;
import java.io.ByteArrayOutputStream;
import java.nio.charset.StandardCharsets;
import java.nio.file.Path;
import java.nio.file.Paths;
import java.util.LinkedHashMap;
import java.util.Map;
import java.util.zip.ZipEntry;
import java.util.zip.ZipInputStream;
import java.util.zip.ZipOutputStream;
public class ReceiptArchiveStaging {
static Map<Path, byte[]> stage(byte[] archive) throws Exception {
Path root = Paths.get("receipt-output").toAbsolutePath().normalize();
Map<Path, byte[]> staged = new LinkedHashMap<>(); int total=0;
try (ZipInputStream input = new ZipInputStream(new ByteArrayInputStream(archive))) {
ZipEntry entry;
while ((entry=input.getNextEntry()) != null) {
String name=entry.getName();
if (name.isEmpty() || name.startsWith("/") || name.contains("\\") || name.contains(":")) throw new IllegalArgumentException("Invalid name");
Path target=root.resolve(name).normalize();
if (!target.startsWith(root) || target.equals(root)) throw new IllegalArgumentException("Outside root");
if (entry.isDirectory()) { input.closeEntry(); continue; }
if (staged.size()==8 || staged.containsKey(target)) throw new IllegalArgumentException("File count or duplicate");
ByteArrayOutputStream content=new ByteArrayOutputStream(); byte[] buffer=new byte[1024]; int count;
while ((count=input.read(buffer)) != -1) {
if (count>65536-total) throw new IllegalArgumentException("Uncompressed budget");
total+=count; content.write(buffer,0,count);
}
staged.put(target,content.toByteArray()); input.closeEntry();
}
}
return staged;
}
static byte[] archive(String name) throws Exception {
ByteArrayOutputStream bytes=new ByteArrayOutputStream();
try (ZipOutputStream output=new ZipOutputStream(bytes)) {
output.putNextEntry(new ZipEntry(name)); output.write("R-51".getBytes(StandardCharsets.UTF_8)); output.closeEntry();
}
return bytes.toByteArray();
}
public static void main(String[] args) throws Exception {
System.out.println(stage(archive("reports/receipt.txt")).size());
try { stage(archive("../outside.txt")); }
catch (IllegalArgumentException rejected) { System.out.println("traversal rejected"); }
}
}Output
1
traversal rejectedCosts and boundaries
Decompression work depends on input and output bytes; staging retains at most 64 KiB of payload plus bounded path and buffer overhead. The caller must also bound compressed input before passing a byte array. This sample does not model encrypted archives, filesystem links, concurrent directory changes or crash-durable publication.
Common Mistakes
- Do not validate destinations using string prefixes.
- Do not trust ZIP header size as the decompression budget.
- Do not publish partial files before validating the complete archive.
Read next
Path resolution, Validated file publication, Staged import ownership.
