Understand the idea
One-hot encoding creates indicator columns. Unknown categories must not change the fitted schema. Binary flags, measured numbers and named categories carry different meanings.
Python skill: An unseen category produces zeros in that feature’s learned indicator columns instead of refitting the schema.
Meet the syntax
OneHotEncoder(handle_unknown='ignore', sparse_output=False)
encoder.get_feature_names_out()handle_unknown='ignore'- An unseen category produces zeros in that feature’s learned indicator columns instead of refitting the schema.
sparse_output=False- Returns a dense array for these small examples.
OneHotEncoder- Learns named category indicators without assigning a numeric ordering.
encoder.get_feature_names_out()- Reads indicator names in the learned output-column order.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
from sklearn.preprocessing import OneHotEncoder
encoder=OneHotEncoder(handle_unknown='ignore',sparse_output=False).fit(X_train[['service']])
answer=encoder.transform(X_train[['service']])
This practice: Read and run the Python. Next: Change · Encode categories without inventing order.
Given data · MIX60
60 observations. One synthetic delivery observation generated for practice. The dataframe df is supplied afresh for each Run.
| distance | weight | service | weekend | duration |
|---|---|---|---|---|
| 15.7052 | 6.75035 | standard | 0 | 61.7704 |
| 9.33869 | 4.81674 | express | 1 | 35.1694 |
| 17.3134 | 5.73931 | economy | 0 | 59.0766 |
| 14.25 | 7.69699 | standard | 1 | 56.5785 |
| 2.78937 | 6.42024 | express | 0 | 19.5577 |
| 19.5368 | 5.62508 | economy | 1 | 77.0289 |
| 15.4617 | 5.68023 | standard | 0 | 56.8109 |
| 15.9352 | 3.17871 | express | 1 | 52.08 |
Column meanings and units
Distance, weight and duration use the fixture’s numeric units; no kilometres, kilograms, minutes or other physical units are specified. RMSE is reported in the same synthetic duration units as the target.
These deterministic teaching observations do not describe real deliveries. Service effects and the alternating weekend flag are built into the generated response; they do not establish real-world causal effects.
distancefloat64- Numeric delivery-distance inputUnit / values: Synthetic distance units; physical unit unspecified
weightfloat64- Numeric parcel-weight inputUnit / values: Synthetic weight units; physical unit unspecified
servicestr- Delivery-service categoryUnit / values: standard / express / economy
weekendint64- Binary weekend input, alternating in the fixtureUnit / values: 0 / 1 indicator
durationfloat64- Numeric delivery-duration targetUnit / values: Synthetic duration units; physical unit unspecified
Supplied setup · available if you need to inspect it
This code runs before your editor on every Run. These are the objects your exercise uses.
from sklearn.model_selection import train_test_split
X = df[['distance','weight','service','weekend']]
y = df['duration']
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42)
Your task · Follow
Fit an encoder to training service values and transform them.
Hint 1 — Think
Named service categories need indicators, not invented distances between codes.
Hint 2 — Tools
OneHotEncoder with a fixed unknown-category policy.
Hint 3 — Approach
Fit the service schema from training values and inspect the dense transformed array.
Explained solution
from sklearn.preprocessing import OneHotEncoder
encoder=OneHotEncoder(handle_unknown='ignore',sparse_output=False).fit(X_train[['service']])
answer=encoder.transform(X_train[['service']])
The learned indicator columns represent membership without imposing an ordering on service names.
Helpful prior knowledge: Learn a scale from training rows These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.