A bounded CSV import validates its byte budget, row count and field schema before publishing accepted records.
Python CSV ingestion: cap bytes, rows and fields before publication
Operation contract
The importer receives a byte buffer no larger than 512 bytes, decodes strict UTF-8 and requires a receipt_id,amount_minor header. It stages at most three records, rejects duplicate IDs and accepts only bounded ASCII decimal fields. The return value appears only after the complete import passes, so callers cannot mistake a yielded prefix for an accepted batch.
Failure and ownership boundary
The byte input is already materialized; a listening server must enforce the byte limit while receiving it. This fixture does not handle spreadsheet formula export, uploaded file storage or authorization. CSV parsing alone does not establish numeric units or identity. Python CSV imports: parse quoted fields before validating rows and Python SQLite project: transactional batches, duplicate IDs and reopen checks separate format handling from persistence.
Working program
import csv
import io
def import_receipts(payload):
if type(payload) is not bytes or len(payload) > 512:
raise ValueError("byte budget")
reader = csv.reader(io.StringIO(payload.decode("utf-8", errors="strict")), strict=True)
if next(reader, None) != ["receipt_id", "amount_minor"]:
raise ValueError("header mismatch")
staged = {}
for row in reader:
if len(staged) == 3 or len(row) != 2:
raise ValueError("row budget or shape")
if any(not field or len(field) > 6 or not field.isascii() or not field.isdecimal() for field in row):
raise ValueError("bounded decimal fields required")
receipt_id, amount = map(int, row)
if receipt_id <= 0 or amount <= 0 or receipt_id in staged:
raise ValueError("invalid or duplicate receipt")
staged[receipt_id] = amount
return staged
print(import_receipts(b"receipt_id,amount_minor\n41,125\n42,75\n"))
try:
import_receipts(b"receipt_id,amount_minor\n41,125\n41,75\n")
except ValueError:
print("duplicate batch rejected")Output
{41: 125, 42: 75}
duplicate batch rejectedCosts and limits
Parsing costs O(b) for b bounded input bytes, with staged storage O(r) for r accepted rows. The input buffer and decoded text also consume memory; row limits alone do not bound upload memory.
Common Mistakes
- Enforce the byte budget during reception as well as at this function boundary.
- Do not insert records as they are parsed if the whole batch must pass first.
Connected lessons
Python CSV imports: parse quoted fields before validating rows, Python regular expressions: use fullmatch for a complete field contract, Python SQLite project: transactional batches, duplicate IDs and reopen checks.
Related Python operation checks
Pandas chunked CSV aggregation: validate every chunk before returning totals.
Check this related boundary
Python assignment expressions: capture a checked value once.
Follow the integrity boundary
Python CSV headers: reject duplicate names before row mapping.
