Bootstrap resampling draws with replacement from an observed sample to examine variation in a chosen statistic under that empirical sampling model.
NumPy bootstrap means: seeded resampling does not repair biased input
Operation contract
The fixture accepts two to 64 finite numeric observations and one to 4096 resamples. It draws an owned index matrix with a seeded Generator, computes each resampled mean, and reports the fifth and ninety-fifth percentiles. Repeating the same seed under the tested package version reproduces the interval. The sampling rows do not alias the original observations.
Failure and ownership boundary
This interval assumes the observed rows are suitable units to resample independently. Repeated measurements, clustering, selection bias and distribution shift can invalidate that interpretation. A fixed seed aids reproducibility; it supplies no scientific or business validity. The percentile range is not a guaranteed coverage statement for every underlying population. Scikit-learn pipelines: fit preprocessing on training data only and NumPy floating-point checks: finite values and declared tolerances remain separate checks.
Tested environment
Dependency check: this program was executed on CPython 3.14.6 with numpy==2.5.3. Install these versions in a separate virtual environment. The download includes the recorded environment snapshot; no third-party package is part of the website runtime.
Working program
import numpy as np
def resampled_mean_interval(values, draws, seed):
if not isinstance(values, list) or not 2 <= len(values) <= 64 or any(type(value) not in (int, float) or not np.isfinite(value) or not 0 <= value <= 1000000 for value in values):
raise ValueError("observations rejected")
if type(draws) is not int or not 1 <= draws <= 4096 or type(seed) is not int or not 0 <= seed <= 1000000:
raise ValueError("resampling budget rejected")
observed = np.array(values, dtype=float)
generator = np.random.default_rng(seed)
indexes = generator.integers(0, len(values), size=(draws, len(values)))
means = observed[indexes].mean(axis=1)
return np.quantile(means, [.05, .95])
values = [100, 200, 300, 400]
interval = resampled_mean_interval(values, 2000, 41)
print("percentile interval:", np.round(interval, 1).tolist())
print("seed reproduced:", bool(np.array_equal(interval, resampled_mean_interval(values, 2000, 41))))
print("input unchanged:", values == [100, 200, 300, 400])Output
percentile interval: [150.0, 350.0]
seed reproduced: True
input unchanged: TrueCosts and limits
For b resamples of n observations, the index matrix and sampled values require O(bn) work and memory before percentile selection. Explicit budgets cap that retained matrix here. Seeded results can change across algorithms or package versions.
Common Mistakes
- A seed addresses reproducibility, not sample representativeness.
- Resampling independent rows is not a policy for clustered observations.
Connected lessons
NumPy floating-point checks: finite values and declared tolerances, Scikit-learn pipelines: fit preprocessing on training data only, Python classification metrics: fix label order before counting errors.
Follow the service contract
Python quantiles: choose an interpolation policy before reporting a threshold.
