Skip to content
AITroveRead. Build. Understand.
Make this comfortable

NumPy least squares: inspect rank before trusting fitted coefficients

Last updated: 30 Sept 20264 min read
tutorial
IntermediateBy AITrove Editorial

Least squares chooses coefficients that minimize squared residual error for a supplied linear design matrix and response vector.

Download Python source kit

Operation contract

The fixture fits an intercept and one feature to an owned exact line, then checks the matrix rank and predictions. A second design with identical columns has deficient rank. The solver can still return coefficients there, but individual coefficient interpretation is not identified by the supplied data. The program reports that distinction instead of treating every numeric return as a valid domain model.

Failure and ownership boundary

The fixture has no generalization estimate, causality claim or noisy measurement model. Floating-point tolerance, feature scaling and ill conditioning need separate policies. Split data before fitting transformations when evaluating predictive behavior; Scikit-learn pipelines: fit preprocessing on training data only and NumPy floating-point checks: finite values and declared tolerances address those layers.

Tested environment

Dependency check: this program was executed on CPython 3.14.6 with numpy==2.5.3. Install these versions in a separate virtual environment. The download includes the recorded environment snapshot; no third-party package is part of the website runtime.

Working program

python
import numpy as np

feature = np.array([0, 1, 2, 3], dtype=float)
design = np.column_stack([np.ones_like(feature), feature])
observed = np.array([3, 5, 7, 9], dtype=float)
coefficients, residuals, rank, singular = np.linalg.lstsq(design, observed, rcond=None)
print("full rank:", int(rank) == design.shape[1])
print("coefficients:", np.round(coefficients, 6).tolist())
print("fit accepted:", bool(np.allclose(design @ coefficients, observed, rtol=1e-10, atol=1e-10)))
duplicate_columns = np.column_stack([feature, feature])
_, _, deficient_rank, _ = np.linalg.lstsq(duplicate_columns, observed, rcond=None)
print("deficient rank:", int(deficient_rank) < duplicate_columns.shape[1])

Output

Output
full rank: True
coefficients: [3.0, 2.0]
fit accepted: True
deficient rank: True

Costs and limits

Dense least-squares cost depends on the matrix dimensions and decomposition used. Retaining an n-by-p design already costs O(np) values. Rank thresholds and conditioning matter; a small residual alone does not establish reliable coefficients or future predictions.

Common Mistakes

  • A deficient design can return numbers without identifying individual effects.
  • An exact training fit is not an evaluation of unseen records.

Connected lessons

NumPy floating-point checks: finite values and declared tolerances, Scikit-learn pipelines: fit preprocessing on training data only, NumPy broadcasting: declare axes and bound integer arithmetic.

python
linear-least-squares
Storage details