A scikit-learn Pipeline chains preprocessing and an estimator so fitting and prediction apply the same ordered transformations.
Scikit-learn pipelines: fit preprocessing on training data only
Operation contract
The fixture has an explicit training partition and an independent evaluation partition. StandardScaler learns its mean from training rows during fit. Prediction transforms the evaluation rows with that fitted state without refitting the scaler. The program inspects the fitted mean before and after prediction to test that boundary rather than advertising an accuracy score on a tiny synthetic dataset.
Failure and ownership boundary
A pipeline cannot correct a split that already leaked related records, future observations or duplicate customers across partitions. Hyperparameter selection also needs a separate validation or cross-validation policy. This classifier has no calibrated business guarantee, fairness evaluation or deployment interface. Pandas rolling windows: minimum observations and causal boundaries and Python classification metrics: fix label order before counting errors cover two common evaluation failures.
Tested environment
Dependency check: this program was executed on CPython 3.14.6 with numpy==2.5.3, scikit-learn==1.9.1. Install these versions in a separate virtual environment. The download includes the recorded environment snapshot; no third-party package is part of the website runtime.
Working program
import numpy as np
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
training = np.array([[1, 10], [2, 20], [3, 30], [4, 40]], dtype=float)
labels = np.array([0, 0, 1, 1])
evaluation = np.array([[2.5, 25], [100, 1000]], dtype=float)
model = Pipeline([("scale", StandardScaler()),
("classify", LogisticRegression(random_state=41, max_iter=200))])
model.fit(training, labels)
owned_mean = model.named_steps["scale"].mean_.copy()
print("training mean:", owned_mean.tolist())
predictions = model.predict(evaluation)
print("prediction count:", len(predictions))
print("mean unchanged:", bool(np.array_equal(owned_mean, model.named_steps["scale"].mean_)))Output
training mean: [2.5, 25.0]
prediction count: 2
mean unchanged: TrueCosts and limits
Scaling n rows and p dense features costs O(np) work and retains p fitted statistics. Logistic regression training cost depends on solver, iterations, features and data. There is no general constant-time fit claim, and this fixture does not benchmark it.
Common Mistakes
- Never fit preprocessing on held-out evaluation rows.
- A pipeline does not choose an appropriate time, group or identity split for you.
Connected lessons
Pandas rolling windows: minimum observations and causal boundaries, Python classification metrics: fix label order before counting errors, NumPy floating-point checks: finite values and declared tolerances.
Follow the service contract
Python time-series validation: fit on past rows and leave a declared gap.
