Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Classification lessonsQUESTIONS · MODELS · EVIDENCE
Probabilities and neighbours · ML-C06 · 12–18 MIN

A logistic workflow

Tune regularisation and investigate shortcuts using training evidence.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeAdapt the Python
  3. TransferAdapt the Python

Understand the idea

Larger C means weaker regularisation. Compare supported C values within a pipeline. A geographic-feature comparison must use matching training folds, not final-test results.

Tune regularisation and investigate shortcuts using training evidence.Fold 1VTTTTFold 2TVTTTFold 3TTVTTFold 4TTTVTFold 5TTTTVFinalT: fit on training rows V: validate Final: held away
Schematic · Tune regularisation and investigate shortcuts using training evidence.Scroll the diagram horizontally if needed.

Python skill: Addresses inverse regularisation strength in the pipeline’s model step.

Meet the syntax

GridSearchCV(model, {'model__C':[.1,1,10]}, cv=folds, scoring='f1_macro')
'model__C'
Addresses inverse regularisation strength in the pipeline’s model step.
[.1,1,10]
Declares stronger through weaker regularisation candidates.
scoring='f1_macro'
Chooses the candidate using equally weighted class F1 on training folds.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

search=GridSearchCV(model,{'model__C':[.1,1,10]},cv=folds,scoring='f1_macro').fit(X_train,y_train)
answer=search.cv_results_['mean_test_score']

This practice: Read and run the Python. Next: Change · A logistic workflow.

Given data · penguins

333 observations. One measured penguin. The dataframe df is supplied afresh for each Run.

Download source CSV · Source and original dictionary

penguins · first 8 prepared rows
islandbill_length_mmbill_depth_mmflipper_length_mmbody_mass_gsexyearspecies
Torgersen39.118.71813750male2007Adelie
Torgersen39.517.41863800female2007Adelie
Torgersen40.3181953250female2007Adelie
Torgersen36.719.31933450female2007Adelie
Torgersen39.320.61903650male2007Adelie
Torgersen38.917.81813625female2007Adelie
Torgersen39.219.61954675male2007Adelie
Torgersen41.117.61823200female2007Adelie

Column meanings and units

Bill length/depth and flipper length: mm. Body mass: grams. Year and island: sampling context.

ML uses 333 complete cases. Removing incomplete records may change the represented population. Geographic context may not generalise to new islands.

Input schema
ColumnStored type
islandstr
bill_length_mmfloat64
bill_depth_mmfloat64
flipper_length_mmint64
body_mass_gint64
sexstr
yearint64
speciesstr
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.model_selection import train_test_split
numeric=['bill_length_mm','bill_depth_mm','flipper_length_mm','body_mass_g']
X=df[numeric+['island','sex','year']]
y=df.species
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42,stratify=y)
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler,OneHotEncoder
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import GridSearchCV,StratifiedKFold,cross_validate
prepare=ColumnTransformer([('numeric',StandardScaler(),numeric),('category',OneHotEncoder(handle_unknown='ignore',sparse_output=False),['island','sex','year'])])
model=Pipeline([('prepare',prepare),('model',LogisticRegression(max_iter=2000,random_state=42))])
folds=StratifiedKFold(5,shuffle=True,random_state=42)

Your task · Follow

Search C for the mixed Penguin pipeline.

Hint 1 — Think

Regularisation strength is chosen using training-only class-balanced evidence.

Hint 2 — Tools

GridSearchCV, model__C and f1_macro.

Hint 3 — Approach

Search the declared C values with the supplied mixed pipeline and fixed folds.

Explained solution
search=GridSearchCV(model,{'model__C':[.1,1,10]},cv=folds,scoring='f1_macro').fit(X_train,y_train)
answer=search.cv_results_['mean_test_score']

Preparation is refitted within each fold, so the comparison evaluates the complete candidate recipe without test leakage.

Helpful prior knowledge: Logistic regression These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Search C for the mixed Penguin pipeline.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.