Understand the idea
Larger C means weaker regularisation. Compare supported C values within a pipeline. A geographic-feature comparison must use matching training folds, not final-test results.
Python skill: Addresses inverse regularisation strength in the pipeline’s model step.
Meet the syntax
GridSearchCV(model, {'model__C':[.1,1,10]}, cv=folds, scoring='f1_macro')'model__C'- Addresses inverse regularisation strength in the pipeline’s model step.
[.1,1,10]- Declares stronger through weaker regularisation candidates.
scoring='f1_macro'- Chooses the candidate using equally weighted class F1 on training folds.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
search=GridSearchCV(model,{'model__C':[.1,1,10]},cv=folds,scoring='f1_macro').fit(X_train,y_train)
answer=search.cv_results_['mean_test_score']
This practice: Read and run the Python. Next: Change · A logistic workflow.
Given data · penguins
333 observations. One measured penguin. The dataframe df is supplied afresh for each Run.
Download source CSV · Source and original dictionary
| island | bill_length_mm | bill_depth_mm | flipper_length_mm | body_mass_g | sex | year | species |
|---|---|---|---|---|---|---|---|
| Torgersen | 39.1 | 18.7 | 181 | 3750 | male | 2007 | Adelie |
| Torgersen | 39.5 | 17.4 | 186 | 3800 | female | 2007 | Adelie |
| Torgersen | 40.3 | 18 | 195 | 3250 | female | 2007 | Adelie |
| Torgersen | 36.7 | 19.3 | 193 | 3450 | female | 2007 | Adelie |
| Torgersen | 39.3 | 20.6 | 190 | 3650 | male | 2007 | Adelie |
| Torgersen | 38.9 | 17.8 | 181 | 3625 | female | 2007 | Adelie |
| Torgersen | 39.2 | 19.6 | 195 | 4675 | male | 2007 | Adelie |
| Torgersen | 41.1 | 17.6 | 182 | 3200 | female | 2007 | Adelie |
Column meanings and units
Bill length/depth and flipper length: mm. Body mass: grams. Year and island: sampling context.
ML uses 333 complete cases. Removing incomplete records may change the represented population. Geographic context may not generalise to new islands.
| Column | Stored type |
|---|---|
| island | str |
| bill_length_mm | float64 |
| bill_depth_mm | float64 |
| flipper_length_mm | int64 |
| body_mass_g | int64 |
| sex | str |
| year | int64 |
| species | str |
Supplied setup · available if you need to inspect it
This code runs before your editor on every Run. These are the objects your exercise uses.
from sklearn.model_selection import train_test_split
numeric=['bill_length_mm','bill_depth_mm','flipper_length_mm','body_mass_g']
X=df[numeric+['island','sex','year']]
y=df.species
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42,stratify=y)
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler,OneHotEncoder
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import GridSearchCV,StratifiedKFold,cross_validate
prepare=ColumnTransformer([('numeric',StandardScaler(),numeric),('category',OneHotEncoder(handle_unknown='ignore',sparse_output=False),['island','sex','year'])])
model=Pipeline([('prepare',prepare),('model',LogisticRegression(max_iter=2000,random_state=42))])
folds=StratifiedKFold(5,shuffle=True,random_state=42)
Your task · Follow
Search C for the mixed Penguin pipeline.
Hint 1 — Think
Regularisation strength is chosen using training-only class-balanced evidence.
Hint 2 — Tools
GridSearchCV, model__C and f1_macro.
Hint 3 — Approach
Search the declared C values with the supplied mixed pipeline and fixed folds.
Explained solution
search=GridSearchCV(model,{'model__C':[.1,1,10]},cv=folds,scoring='f1_macro').fit(X_train,y_train)
answer=search.cv_results_['mean_test_score']
Preparation is refitted within each fold, so the comparison evaluates the complete candidate recipe without test leakage.
Helpful prior knowledge: Logistic regression These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.