Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Scikit-learn pipelines: fit preprocessing on training data only

Last updated: 30 Sept 20264 min read
tutorial
IntermediateBy AITrove Editorial

A scikit-learn Pipeline chains preprocessing and an estimator so fitting and prediction apply the same ordered transformations.

Download Python source kit

Operation contract

The fixture has an explicit training partition and an independent evaluation partition. StandardScaler learns its mean from training rows during fit. Prediction transforms the evaluation rows with that fitted state without refitting the scaler. The program inspects the fitted mean before and after prediction to test that boundary rather than advertising an accuracy score on a tiny synthetic dataset.

Failure and ownership boundary

A pipeline cannot correct a split that already leaked related records, future observations or duplicate customers across partitions. Hyperparameter selection also needs a separate validation or cross-validation policy. This classifier has no calibrated business guarantee, fairness evaluation or deployment interface. Pandas rolling windows: minimum observations and causal boundaries and Python classification metrics: fix label order before counting errors cover two common evaluation failures.

Tested environment

Dependency check: this program was executed on CPython 3.14.6 with numpy==2.5.3, scikit-learn==1.9.1. Install these versions in a separate virtual environment. The download includes the recorded environment snapshot; no third-party package is part of the website runtime.

Working program

python
import numpy as np
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression

training = np.array([[1, 10], [2, 20], [3, 30], [4, 40]], dtype=float)
labels = np.array([0, 0, 1, 1])
evaluation = np.array([[2.5, 25], [100, 1000]], dtype=float)
model = Pipeline([("scale", StandardScaler()),
                  ("classify", LogisticRegression(random_state=41, max_iter=200))])
model.fit(training, labels)
owned_mean = model.named_steps["scale"].mean_.copy()
print("training mean:", owned_mean.tolist())
predictions = model.predict(evaluation)
print("prediction count:", len(predictions))
print("mean unchanged:", bool(np.array_equal(owned_mean, model.named_steps["scale"].mean_)))

Output

Output
training mean: [2.5, 25.0]
prediction count: 2
mean unchanged: True

Costs and limits

Scaling n rows and p dense features costs O(np) work and retains p fitted statistics. Logistic regression training cost depends on solver, iterations, features and data. There is no general constant-time fit claim, and this fixture does not benchmark it.

Common Mistakes

  • Never fit preprocessing on held-out evaluation rows.
  • A pipeline does not choose an appropriate time, group or identity split for you.

Connected lessons

Pandas rolling windows: minimum observations and causal boundaries, Python classification metrics: fix label order before counting errors, NumPy floating-point checks: finite values and declared tolerances.

Follow the service contract

Python time-series validation: fit on past rows and leave a declared gap.

python
ml-pipeline
Storage details