Least squares chooses coefficients that minimize squared residual error for a supplied linear design matrix and response vector.
NumPy least squares: inspect rank before trusting fitted coefficients
Operation contract
The fixture fits an intercept and one feature to an owned exact line, then checks the matrix rank and predictions. A second design with identical columns has deficient rank. The solver can still return coefficients there, but individual coefficient interpretation is not identified by the supplied data. The program reports that distinction instead of treating every numeric return as a valid domain model.
Failure and ownership boundary
The fixture has no generalization estimate, causality claim or noisy measurement model. Floating-point tolerance, feature scaling and ill conditioning need separate policies. Split data before fitting transformations when evaluating predictive behavior; Scikit-learn pipelines: fit preprocessing on training data only and NumPy floating-point checks: finite values and declared tolerances address those layers.
Tested environment
Dependency check: this program was executed on CPython 3.14.6 with numpy==2.5.3. Install these versions in a separate virtual environment. The download includes the recorded environment snapshot; no third-party package is part of the website runtime.
Working program
import numpy as np
feature = np.array([0, 1, 2, 3], dtype=float)
design = np.column_stack([np.ones_like(feature), feature])
observed = np.array([3, 5, 7, 9], dtype=float)
coefficients, residuals, rank, singular = np.linalg.lstsq(design, observed, rcond=None)
print("full rank:", int(rank) == design.shape[1])
print("coefficients:", np.round(coefficients, 6).tolist())
print("fit accepted:", bool(np.allclose(design @ coefficients, observed, rtol=1e-10, atol=1e-10)))
duplicate_columns = np.column_stack([feature, feature])
_, _, deficient_rank, _ = np.linalg.lstsq(duplicate_columns, observed, rcond=None)
print("deficient rank:", int(deficient_rank) < duplicate_columns.shape[1])Output
full rank: True
coefficients: [3.0, 2.0]
fit accepted: True
deficient rank: TrueCosts and limits
Dense least-squares cost depends on the matrix dimensions and decomposition used. Retaining an n-by-p design already costs O(np) values. Rank thresholds and conditioning matter; a small residual alone does not establish reliable coefficients or future predictions.
Common Mistakes
- A deficient design can return numbers without identifying individual effects.
- An exact training fit is not an evaluation of unseen records.
Connected lessons
NumPy floating-point checks: finite values and declared tolerances, Scikit-learn pipelines: fit preprocessing on training data only, NumPy broadcasting: declare axes and bound integer arithmetic.
