The buffersize argument limits submitted map work whose results have not yet been yielded.
Python Executor.map buffersize: pause submission when results wait
Operation contract
An input generator counts how many receipt IDs have been pulled. With buffersize two, creating the map iterator consumes the first two IDs; consuming results allows later IDs to be pulled. Workers return deterministic labels so scheduling order does not affect displayed output.
Failure boundary
A bounded submission buffer is not a timeout, output-size quota, or cancellation policy. A slow early result can hold ordered map iteration while later work completes. Thread workers share a process, so this pattern does not isolate untrusted code or speed up CPU-heavy Python code on a conventional GIL build.
Working program
from concurrent.futures import ThreadPoolExecutor
submitted = []
def receipt_ids():
for number in range(47, 53):
submitted.append(number)
yield number
def normalize_receipt(number):
return f"receipt-{number}"
with ThreadPoolExecutor(max_workers=2) as pool:
results = pool.map(normalize_receipt, receipt_ids(), buffersize=2)
print("pulled_before_next", len(submitted))
print("first", next(results))
print("all", len([*results]) + 1)Output
pulled_before_next 2
first receipt-47
all 6Costs and limits
At most the configured number of submitted, unyielded map items is buffered by this API; the example's submitted list intentionally retains all IDs for observation. Result collection has O(n) space here.
Common Mistakes
- buffersize does not bound task runtime.
- Ordered results can wait behind a slow first task.
- Do not equate a thread pool with process isolation.
