Itertools composes iterator operations whose work and retained state depend on when consumers request values.
Python itertools: adjacent groups and shared iterator consumption
Operation contract
The receipt stream has two separated runs for one region. Groupby emits three adjacent groups; it does not combine every equal key across the stream. Chain joins two batches without constructing one combined list, and islice consumes only the selected prefix. Zip with strict=True rejects unequal lengths rather than silently discarding unmatched records.
Failure and ownership boundary
A group iterator shares the source with the outer group iterator, so consume or materialize a group before advancing to the next one. Sorting first can produce global key groups but changes order and adds storage/work. Tee can retain a growing backlog when consumers move at different speeds. Python generators: lazy iteration does not make retained output free and Pandas groupby: retain missing keys and define all-null totals use different ownership models.
Working program
from itertools import chain, groupby, islice
receipts = [("DEL", 125), ("BOM", 75), ("DEL", 250)]
runs = [(region, sum(amount for _, amount in group)) for region, group in groupby(receipts, key=lambda row: row[0])]
print("adjacent groups:", runs)
print("bounded prefix:", list(islice(chain([125, 250], [75, 40]), 3)))
try:
list(zip(["R-0041", "R-0042"], [125], strict=True))
except ValueError:
print("unequal batches rejected")Output
adjacent groups: [('DEL', 125), ('BOM', 75), ('DEL', 250)]
bounded prefix: [125, 250, 75]
unequal batches rejectedCosts and limits
The consumed prefix costs O(k) iteration work. Materialized group totals retain output per run, while groupby itself need not retain all input. A bounded islice caps consumption only when each upstream next call itself terminates.
Common Mistakes
- Groupby groups adjacent equal keys, not all equal keys everywhere.
- Strict zip can reject after earlier pairs have already been consumed.
Connected lessons
Python generators: lazy iteration does not make retained output free, Python map, filter and reduce: iterator timing and reduction identity, Pandas groupby: retain missing keys and define all-null totals.
