Understand the idea
ColumnTransformer combines separate operations by column meaning. Each branch is fitted on the same training population, then the transformed columns are combined.
Python skill: Combines different preparation steps for selected column groups.
Meet the syntax
ColumnTransformer([('numeric', StandardScaler(), numeric_columns), ...])
'passthrough'ColumnTransformer- Combines different preparation steps for selected column groups.
'numeric'- Names this transformer step for later inspection and parameter paths.
numeric_columns- Lists the columns sent to StandardScaler; other types need their own declared routes.
'passthrough'- Keeps declared binary flags unchanged while other branches transform their columns.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
prepare=ColumnTransformer([('numeric',StandardScaler(),['distance','weight']),('categories',OneHotEncoder(handle_unknown='ignore',sparse_output=False),['service']),('flags','passthrough',['weekend'])])
answer=prepare.fit_transform(X_train)
This practice: Read and run the Python. Next: Change · Prepare different feature types together.
Given data · MIX60
60 observations. One synthetic delivery observation generated for practice. The dataframe df is supplied afresh for each Run.
| distance | weight | service | weekend | duration |
|---|---|---|---|---|
| 15.7052 | 6.75035 | standard | 0 | 61.7704 |
| 9.33869 | 4.81674 | express | 1 | 35.1694 |
| 17.3134 | 5.73931 | economy | 0 | 59.0766 |
| 14.25 | 7.69699 | standard | 1 | 56.5785 |
| 2.78937 | 6.42024 | express | 0 | 19.5577 |
| 19.5368 | 5.62508 | economy | 1 | 77.0289 |
| 15.4617 | 5.68023 | standard | 0 | 56.8109 |
| 15.9352 | 3.17871 | express | 1 | 52.08 |
Column meanings and units
Distance, weight and duration use the fixture’s numeric units; no kilometres, kilograms, minutes or other physical units are specified. RMSE is reported in the same synthetic duration units as the target.
These deterministic teaching observations do not describe real deliveries. Service effects and the alternating weekend flag are built into the generated response; they do not establish real-world causal effects.
distancefloat64- Numeric delivery-distance inputUnit / values: Synthetic distance units; physical unit unspecified
weightfloat64- Numeric parcel-weight inputUnit / values: Synthetic weight units; physical unit unspecified
servicestr- Delivery-service categoryUnit / values: standard / express / economy
weekendint64- Binary weekend input, alternating in the fixtureUnit / values: 0 / 1 indicator
durationfloat64- Numeric delivery-duration targetUnit / values: Synthetic duration units; physical unit unspecified
Supplied setup · available if you need to inspect it
This code runs before your editor on every Run. These are the objects your exercise uses.
from sklearn.model_selection import train_test_split
X = df[['distance','weight','service','weekend']]
y = df['duration']
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42)
Your task · Follow
Fit the supplied mixed preparation recipe and transform the training inputs.
Hint 1 — Think
Each feature type should travel through the preparation appropriate to its meaning.
Hint 2 — Tools
ColumnTransformer, StandardScaler, OneHotEncoder and passthrough.
Hint 3 — Approach
Build the declared numeric, category and flag branches, then fit-transform training inputs.
Explained solution
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
prepare=ColumnTransformer([('numeric',StandardScaler(),['distance','weight']),('categories',OneHotEncoder(handle_unknown='ignore',sparse_output=False),['service']),('flags','passthrough',['weekend'])])
answer=prepare.fit_transform(X_train)
The transformer combines heterogeneous features without scaling named categories or changing existing binary flags.
Helpful prior knowledge: Encode categories without inventing order These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.