Understand the idea
A pipeline fits each preparation only from the rows passed to fit. During predict, it reuses the learned preparation. Putting it inside CV gives each fold its own fitted statistics.
Python skill: Names the preprocessing step placed before the estimator.
Meet the syntax
Pipeline([('prepare', prepare), ('model', estimator)])
model.set_params(model=estimator)('prepare', prepare)- Names the preprocessing step placed before the estimator.
('model', estimator)- Names the final prediction step; pipeline fitting trains both steps in order.
Pipeline- Keeps fitting and later prediction on the same reusable preparation path.
model.set_params(model=estimator)- Replaces the named final pipeline step; fit the resulting recipe before predicting.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
prepare=ColumnTransformer([('numeric','passthrough',['distance','weight']),('categories',OneHotEncoder(handle_unknown='ignore',sparse_output=False,drop='first'),['service']),('flags','passthrough',['weekend'])])
from sklearn.pipeline import Pipeline
from sklearn.linear_model import LinearRegression
model=Pipeline([('prepare',prepare),('model',LinearRegression())])
model.fit(X_train,y_train)
answer=model.predict(X_test)
This practice: Read and run the Python. Next: Change · Keep preparation with the estimator.
Given data · MIX60
60 observations. One synthetic delivery observation generated for practice. The dataframe df is supplied afresh for each Run.
| distance | weight | service | weekend | duration |
|---|---|---|---|---|
| 15.7052 | 6.75035 | standard | 0 | 61.7704 |
| 9.33869 | 4.81674 | express | 1 | 35.1694 |
| 17.3134 | 5.73931 | economy | 0 | 59.0766 |
| 14.25 | 7.69699 | standard | 1 | 56.5785 |
| 2.78937 | 6.42024 | express | 0 | 19.5577 |
| 19.5368 | 5.62508 | economy | 1 | 77.0289 |
| 15.4617 | 5.68023 | standard | 0 | 56.8109 |
| 15.9352 | 3.17871 | express | 1 | 52.08 |
Column meanings and units
Distance, weight and duration use the fixture’s numeric units; no kilometres, kilograms, minutes or other physical units are specified. RMSE is reported in the same synthetic duration units as the target.
These deterministic teaching observations do not describe real deliveries. Service effects and the alternating weekend flag are built into the generated response; they do not establish real-world causal effects.
distancefloat64- Numeric delivery-distance inputUnit / values: Synthetic distance units; physical unit unspecified
weightfloat64- Numeric parcel-weight inputUnit / values: Synthetic weight units; physical unit unspecified
servicestr- Delivery-service categoryUnit / values: standard / express / economy
weekendint64- Binary weekend input, alternating in the fixtureUnit / values: 0 / 1 indicator
durationfloat64- Numeric delivery-duration targetUnit / values: Synthetic duration units; physical unit unspecified
Supplied setup · available if you need to inspect it
This code runs before your editor on every Run. These are the objects your exercise uses.
from sklearn.model_selection import train_test_split
X = df[['distance','weight','service','weekend']]
y = df['duration']
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42)
Your task · Follow
Build the mixed pipeline, fit training rows and predict test rows.
Hint 1 — Think
The same learned preparation must run before every later prediction.
Hint 2 — Tools
Pipeline with named prepare/model steps.
Hint 3 — Approach
Build mixed preparation, attach the line estimator, fit on training rows and predict test features through the whole pipeline.
Explained solution
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
prepare=ColumnTransformer([('numeric','passthrough',['distance','weight']),('categories',OneHotEncoder(handle_unknown='ignore',sparse_output=False,drop='first'),['service']),('flags','passthrough',['weekend'])])
from sklearn.pipeline import Pipeline
from sklearn.linear_model import LinearRegression
model=Pipeline([('prepare',prepare),('model',LinearRegression())])
model.fit(X_train,y_train)
answer=model.predict(X_test)
The linear pipeline preserves numeric units and learns a category schema with an omitted reference, matching the playground recipe. Prediction reuses that fitted preparation.
Helpful prior knowledge: Prepare different feature types together · A useful reference These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.