Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Regression lessonsQUESTIONS · MODELS · EVIDENCE
Several predictors · ML-R05 · 12–18 MIN

Read coefficients after encoding

Align coefficients with encoded feature names and a reference category.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeAdapt the Python
  3. TransferExplain the Python result

Understand the idea

Dropping one category gives an explicit reference for linear coefficients with an intercept. Category coefficients compare with that reference, conditional on other inputs.

Align coefficients with encoded feature names and a reference category.Serviceexpressstandardeconomy00express10standard01Economy is the omitted reference in this example.
Schematic · Align coefficients with encoded feature names and a reference category.Scroll the diagram horizontally if needed.

Python skill: Omits one category indicator so coefficients compare other categories with that reference.

Meet the syntax

OneHotEncoder(drop='first')
prepare.get_feature_names_out()
drop='first'
Omits one category indicator so coefficients compare other categories with that reference.
prepare.get_feature_names_out()
Returns names in the transformed-column order so coefficients can be labelled correctly.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

encoder=model.named_steps['prepare'].named_transformers_['service']
answer=encoder.categories_[0][encoder.drop_idx_[0]]

This practice: Read and run the Python. Next: Change · Read coefficients after encoding.

Given data · MIX60

60 observations. One synthetic delivery observation generated for practice. The dataframe df is supplied afresh for each Run.

MIX60 · first 8 prepared rows
distanceweightserviceweekendduration
15.70526.75035standard061.7704
9.338694.81674express135.1694
17.31345.73931economy059.0766
14.257.69699standard156.5785
2.789376.42024express019.5577
19.53685.62508economy177.0289
15.46175.68023standard056.8109
15.93523.17871express152.08

Column meanings and units

Distance, weight and duration use the fixture’s numeric units; no kilometres, kilograms, minutes or other physical units are specified. RMSE is reported in the same synthetic duration units as the target.

These deterministic teaching observations do not describe real deliveries. Service effects and the alternating weekend flag are built into the generated response; they do not establish real-world causal effects.

distancefloat64
Numeric delivery-distance inputUnit / values: Synthetic distance units; physical unit unspecified
weightfloat64
Numeric parcel-weight inputUnit / values: Synthetic weight units; physical unit unspecified
servicestr
Delivery-service categoryUnit / values: standard / express / economy
weekendint64
Binary weekend input, alternating in the fixtureUnit / values: 0 / 1 indicator
durationfloat64
Numeric delivery-duration targetUnit / values: Synthetic duration units; physical unit unspecified
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.model_selection import train_test_split
X = df[['distance','weight','service','weekend']]
y = df['duration']
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42)
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder
from sklearn.pipeline import Pipeline
from sklearn.linear_model import LinearRegression
prepare=ColumnTransformer([('numeric','passthrough',['distance','weight']),('service',OneHotEncoder(drop='first',handle_unknown='ignore',sparse_output=False),['service'])])
model=Pipeline([('prepare',prepare),('model',LinearRegression())]).fit(X_train,y_train)

Your task · Follow

Identify the omitted service reference from the fitted encoder.

Hint 1 — Think

The omitted indicator establishes the category against which the encoded coefficient is compared.

Hint 2 — Tools

named_transformers_, categories_ and drop_idx_.

Hint 3 — Approach

Reach the fitted service encoder and use its dropped-position metadata to identify the reference.

Explained solution
encoder=model.named_steps['prepare'].named_transformers_['service']
answer=encoder.categories_[0][encoder.drop_idx_[0]]

Reading fitted metadata avoids assuming an alphabetical reference or guessing from input order.

Helpful prior knowledge: Several predictors, conditional comparisons · Encode categories without inventing order These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Identify the omitted service reference from the fitted encoder.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.