Validated aggregation combines accepted records into per-key totals without treating malformed rows as ordinary zero values.
Python aggregation exercise: validate records before returning group totals
Operation contract
The region exercise accepts an exact pair of fields, one known ASCII region and a bounded nonnegative integer amount. It accumulates into a private dictionary and returns it only after the complete batch succeeds. A rejected later record therefore cannot become a partially accepted public answer. The returned totals do not alias the original row dictionaries.
Failure and ownership boundary
The batch is already an owned list, so its row limit is not a network reception limit. Publishing totals to a database or file needs a separate transaction or staging contract. Missing keys, booleans and unknown regions are deliberately rejected instead of guessed. Pandas chunked CSV aggregation: validate every chunk before returning totals and Python defaultdict: missing-key reads can create state show adjacent implementation choices.
Working program
def accepted_totals(records):
if not isinstance(records, list) or len(records) > 64:
raise ValueError("record batch rejected")
totals = {}
for record in records:
if not isinstance(record, dict) or set(record) != {"region", "amount"}:
raise ValueError("record fields rejected")
if not isinstance(record["region"], str) or record["region"] not in {"DEL", "BOM"}:
raise ValueError("region rejected")
amount = record["amount"]
if type(amount) is not int or not 0 <= amount <= 1000000:
raise ValueError("amount rejected")
region = record["region"]
totals[region] = totals.get(region, 0) + amount
return totals
print(accepted_totals([{"region": "DEL", "amount": 125}, {"region": "DEL", "amount": 250}]))
try:
accepted_totals([{"region": "DEL", "amount": 125}, {"region": "BOM", "amount": True}])
except ValueError:
print("later record rejected; no totals returned")Output
{'DEL': 375}
later record rejected; no totals returnedCosts and limits
For n accepted rows and g regions, expected mapping work is O(n) and result storage is O(g), excluding numeric digit costs. This schema has only two permitted groups and a row/amount budget, so aggregate growth is bounded.
Common Mistakes
- A boolean is not an accepted amount even though bool subclasses int.
- Do not return or publish partial totals after a rejected row.
Connected lessons
Pandas chunked CSV aggregation: validate every chunk before returning totals, Python defaultdict: missing-key reads can create state, Python SQLite project: transactional batches, duplicate IDs and reopen checks.
Check the next state boundary
Python reconciliation exercise: reject duplicate IDs before comparing ledgers.
