itertools.groupby starts a new group whenever the key changes in input order.
Make this comfortable
Python itertools.groupby: contiguous runs are not global groups
Observed contract
Three receipts arrive for north, south, then north again. Streaming groupby reports three runs. Sorting a copy by region first lets the same operation produce one total per region without changing the original arrival list.
Boundary
The group iterator shares the underlying input iterator and must be consumed before advancing to the next group. Sorting creates a full materialized copy and loses arrival order; choose a dictionary accumulator when global totals matter more than sorted output.
Executable case
from itertools import groupby
receipts = [("north", 47), ("south", 26), ("north", 11)]
def region_totals(ordered_receipts):
return [(region, sum(amount for _, amount in run))
for region, run in groupby(ordered_receipts, key=lambda receipt: receipt[0])]
print("arrival_runs", region_totals(receipts))
sorted_receipts = sorted(receipts, key=lambda receipt: receipt[0])
print("region_totals", region_totals(sorted_receipts))
print("original_order", [region for region, _ in receipts])Output
arrival_runs [('north', 47), ('south', 26), ('north', 11)]
region_totals [('north', 58), ('south', 26)]
original_order ['north', 'south', 'north']Cost
Grouping already ordered data scans it once with constant iterator overhead beyond each active run. Sorting an unsorted collection costs O(n log n) time and O(n) storage for the new list.
Common Mistakes
- Do not treat separated equal keys as one group.
- Consume each run before requesting the next group.
- Sorting is a semantic choice when arrival order carries meaning.
Connected lessons
python
groupby-consecutive-keys
