Chunked CSV reading yields batches so aggregation need not retain every parsed row at once.
Pandas chunked CSV aggregation: validate every chunk before returning totals
Operation contract
The import accepts exactly region and amount columns, known region labels and nonnegative integer minor units. Each chunk is validated before its group totals enter a private accumulator. A failure in a later chunk raises without returning a partial result. The fixture bounds source text and row count as well as chunk size.
Failure and ownership boundary
Chunking bounds a batch, not the entire input or number of groups. Real files require a pre-read byte policy and parser limits; this fixture already owns a small string. It does not write results externally, so no transaction is needed to hide partial internal accumulation. Python CSV ingestion: cap bytes, rows and fields before publication and Python SQLite project: transactional batches, duplicate IDs and reopen checks cover those different stages.
Tested environment
Dependency check: this program was executed on CPython 3.14.6 with pandas==3.0.6. Install these versions in a separate virtual environment. The download includes the recorded environment snapshot; no third-party package is part of the website runtime.
Working program
from io import StringIO
import pandas as pd
def region_totals(text):
if not isinstance(text, str) or len(text) > 1024:
raise ValueError("CSV budget exceeded")
totals = {}
rows = 0
with pd.read_csv(StringIO(text), chunksize=2,
dtype={"region": "string", "amount": "Int64"}) as chunks:
for chunk in chunks:
rows += len(chunk)
if list(chunk.columns) != ["region", "amount"] or rows > 100:
raise ValueError("CSV shape rejected")
if chunk.isna().any().any() or not chunk["region"].isin(["DEL", "BOM"]).all():
raise ValueError("CSV fields rejected")
if (chunk["amount"] < 0).any() or (chunk["amount"] > 1000000).any():
raise ValueError("amount range rejected")
for region, amount in chunk.groupby("region")["amount"].sum().items():
totals[region] = totals.get(region, 0) + int(amount)
return totals
print(region_totals("region,amount\nDEL,125\nBOM,75\nDEL,250\n"))
try:
region_totals("region,amount\nDEL,125\nBOM,75\nDEL,-1\n")
except ValueError:
print("later chunk rejected; no result returned")Output
{'BOM': 75, 'DEL': 375}
later chunk rejected; no result returnedCosts and limits
For n accepted rows, aggregation performs row-dependent work and retains one chunk plus group accumulators. With c chunk rows and g group keys, parsed working storage is approximately O(c+g), excluding the owned source string and parser buffers. Integer overflow policy matters for large vectorized sums; per-row and row-count bounds keep this fixture below that limit.
Common Mistakes
- A chunk size is not a file-size limit.
- Do not expose partial totals before the complete import succeeds.
Connected lessons
Python CSV ingestion: cap bytes, rows and fields before publication, Pandas groupby: retain missing keys and define all-null totals, Python SQLite project: transactional batches, duplicate IDs and reopen checks.
Follow the related contract
Python aggregation exercise: validate records before returning group totals.
