Skip to content
AITroveRead. Build. Understand.
Make this comfortable

NumPy bootstrap means: seeded resampling does not repair biased input

Last updated: 30 Sept 20264 min read
tutorial
IntermediateBy AITrove Editorial

Bootstrap resampling draws with replacement from an observed sample to examine variation in a chosen statistic under that empirical sampling model.

Download Python source kit

Operation contract

The fixture accepts two to 64 finite numeric observations and one to 4096 resamples. It draws an owned index matrix with a seeded Generator, computes each resampled mean, and reports the fifth and ninety-fifth percentiles. Repeating the same seed under the tested package version reproduces the interval. The sampling rows do not alias the original observations.

Failure and ownership boundary

This interval assumes the observed rows are suitable units to resample independently. Repeated measurements, clustering, selection bias and distribution shift can invalidate that interpretation. A fixed seed aids reproducibility; it supplies no scientific or business validity. The percentile range is not a guaranteed coverage statement for every underlying population. Scikit-learn pipelines: fit preprocessing on training data only and NumPy floating-point checks: finite values and declared tolerances remain separate checks.

Tested environment

Dependency check: this program was executed on CPython 3.14.6 with numpy==2.5.3. Install these versions in a separate virtual environment. The download includes the recorded environment snapshot; no third-party package is part of the website runtime.

Working program

python
import numpy as np

def resampled_mean_interval(values, draws, seed):
    if not isinstance(values, list) or not 2 <= len(values) <= 64 or any(type(value) not in (int, float) or not np.isfinite(value) or not 0 <= value <= 1000000 for value in values):
        raise ValueError("observations rejected")
    if type(draws) is not int or not 1 <= draws <= 4096 or type(seed) is not int or not 0 <= seed <= 1000000:
        raise ValueError("resampling budget rejected")
    observed = np.array(values, dtype=float)
    generator = np.random.default_rng(seed)
    indexes = generator.integers(0, len(values), size=(draws, len(values)))
    means = observed[indexes].mean(axis=1)
    return np.quantile(means, [.05, .95])

values = [100, 200, 300, 400]
interval = resampled_mean_interval(values, 2000, 41)
print("percentile interval:", np.round(interval, 1).tolist())
print("seed reproduced:", bool(np.array_equal(interval, resampled_mean_interval(values, 2000, 41))))
print("input unchanged:", values == [100, 200, 300, 400])

Output

Output
percentile interval: [150.0, 350.0]
seed reproduced: True
input unchanged: True

Costs and limits

For b resamples of n observations, the index matrix and sampled values require O(bn) work and memory before percentile selection. Explicit budgets cap that retained matrix here. Seeded results can change across algorithms or package versions.

Common Mistakes

  • A seed addresses reproducibility, not sample representativeness.
  • Resampling independent rows is not a policy for clustered observations.

Connected lessons

NumPy floating-point checks: finite values and declared tolerances, Scikit-learn pipelines: fit preprocessing on training data only, Python classification metrics: fix label order before counting errors.

Follow the service contract

Python quantiles: choose an interpolation policy before reporting a threshold.

python
bootstrap-means
Storage details