Covariance measures joint variation of paired observations around their respective means, with a declared sample or population denominator.
Python covariance: state the denominator and preserve paired observations
Operation contract
The receipt fixture treats x and y as paired measurements. It rejects unequal lengths, fewer than two pairs, Boolean values, nonfinite values and magnitudes over one million. Centered products are summed with fsum and divided by n minus one for sample covariance or n for population covariance. The same three pairs therefore produce sample covariance four and population covariance eight thirds. Magnitude is checked before the finite-number conversion so an oversized integer is rejected by the declared boundary instead of overflowing during a float conversion. Pair membership stays fixed; sorting one column separately would change the question.
Failure and ownership boundary
This is descriptive calculation, not evidence that one variable causes the other. fsum reduces summation error but does not make input floats or centered products exact. Extreme ranges can still lose information, so bounds and a recorded unit remain relevant. Python quantiles: choose an interpolation policy before reporting a threshold, NumPy floating-point checks: finite values and declared tolerances and Pandas merge with null keys: quarantine missing identifiers before matching affect the interpretation.
Working program
import math
def paired_covariance(left, right, *, sample=True):
if type(sample) is not bool or type(left) is not list or type(right) is not list or not 2 <= len(left) == len(right) <= 64:
raise ValueError("paired lengths")
if any(type(value) not in (int, float) or abs(value) > 1000000 or not math.isfinite(value) for value in left + right):
raise ValueError("finite bounded observations")
count = len(left)
left_mean, right_mean = math.fsum(left) / count, math.fsum(right) / count
centered = math.fsum((x - left_mean) * (y - right_mean) for x, y in zip(left, right, strict=True))
return centered / (count - int(sample))
print("sample:", paired_covariance([1, 2, 3], [2, 4, 10]))
print("population:", round(paired_covariance([1, 2, 3], [2, 4, 10], sample=False), 6))
try:
paired_covariance([1, 2], [2, float("nan")])
except ValueError:
print("nonfinite pair rejected")Output
sample: 4.0
population: 2.666667
nonfinite pair rejectedCosts and limits
Two passes over n pairs give O(n) arithmetic work. The implementation concatenates validation lists and builds no matrix, so it retains O(n) temporary references. With <=64 bounded observations, product magnitude remains modest; larger data, streaming summaries and parallel merges need separate error analysis.
Common Mistakes
- Sorting or filtering one column alone destroys observation pairing.
- Sample and population denominators answer different declared questions.
Connected lessons
Python quantiles: choose an interpolation policy before reporting a threshold, NumPy floating-point checks: finite values and declared tolerances, NumPy bootstrap means: seeded resampling does not repair biased input.
