Complete a stratified mixed Penguin logistic workflow with macro F1, a reference, tuning and confusion evidence.
This practice: Complete a stratified mixed Penguin logistic workflow with macro F1, a reference, tuning and confusion evidence. Next: explore Regression or choose another model family.
Given data · penguins
333 observations. One measured penguin. The dataframe df is supplied afresh for each Run.
Download source CSV · Source and original dictionary
| island | bill_length_mm | bill_depth_mm | flipper_length_mm | body_mass_g | sex | year | species |
|---|---|---|---|---|---|---|---|
| Torgersen | 39.1 | 18.7 | 181 | 3750 | male | 2007 | Adelie |
| Torgersen | 39.5 | 17.4 | 186 | 3800 | female | 2007 | Adelie |
| Torgersen | 40.3 | 18 | 195 | 3250 | female | 2007 | Adelie |
| Torgersen | 36.7 | 19.3 | 193 | 3450 | female | 2007 | Adelie |
| Torgersen | 39.3 | 20.6 | 190 | 3650 | male | 2007 | Adelie |
| Torgersen | 38.9 | 17.8 | 181 | 3625 | female | 2007 | Adelie |
| Torgersen | 39.2 | 19.6 | 195 | 4675 | male | 2007 | Adelie |
| Torgersen | 41.1 | 17.6 | 182 | 3200 | female | 2007 | Adelie |
Column meanings and units
Bill length/depth and flipper length: mm. Body mass: grams. Year and island: sampling context.
ML uses 333 complete cases. Removing incomplete records may change the represented population. Geographic context may not generalise to new islands.
| Column | Stored type |
|---|---|
| island | str |
| bill_length_mm | float64 |
| bill_depth_mm | float64 |
| flipper_length_mm | int64 |
| body_mass_g | int64 |
| sex | str |
| year | int64 |
| species | str |
Declared validation design
Reserve 20% of the declared population for the final test using split seed 42. Keep class proportions with a stratified split.
Use the same five stratified training folds, shuffled with seed 42, for reference comparison, candidates and selection. Use out-of-fold training predictions for diagnosis. Select with macro F1 and compare with a most-frequent-class reference. Accuracy is supplementary. Open final-test evidence after selection and diagnosis.
Supporting concepts: A logistic workflow → · Validate k on a meaningful scale → · Nonlinear SVM behaviour → · One-R with numeric inputs → · Classification trees → · Match Naive Bayes to feature types → · QDA allows class-specific shapes →
Remember the idea
This checkpoint combines previously taught skills. Assemble the workflow; help remains available when needed.
Required Python variables and evidence
Use these names so Check can inspect your workflow. Each meaning is shown beside its name.
| Variable | Meaning |
|---|---|
| X | Feature dataframe for the declared population, preserving row indices. |
| X_test | Final-test feature rows from the declared split. |
| X_train | Training feature rows from the declared split. |
| chosen_name | Model ID nominated from training evidence. |
| cv_results | Candidate validation evidence; comparison workflows use a dataframe indexed by model ID. |
| df | Loaded and prepared input dataframe. |
| final_accuracy | Final-test accuracy, supplementary to macro F1. |
| final_f1 | Macro F1 on final-test class predictions. |
| final_model | Chosen pipeline fitted on training rows, after selection and diagnosis. |
| final_predictions | Unaltered predictions of final_model on X_test. |
| matrix | Confusion matrix of training-only diagnostic predictions. |
| reference_results | cross_validate result for the dummy reference; test_score contains five scores. |
| selected | Dictionary mapping each requested model ID to its chosen pipeline. |
| y | Target series, aligned with X. |
| y_test | Final-test targets, aligned with X_test. |
| y_train | Training targets, aligned with X_train. |
Hint 1 — Think
Reconstruct class-aware validation while investigating the role of context.
Hint 2 — Tools
Stratified folds, mixed logistic pipeline, C search, measurement ablation and confusion matrix.
Hint 3 — Approach
Compare reference and candidate, select C, evaluate the matched ablation and OOF class errors, then report final metrics.
Explained solution
from sklearn.base import clone
from sklearn.model_selection import train_test_split, StratifiedKFold, cross_validate, cross_val_predict, GridSearchCV
from sklearn.dummy import DummyClassifier
from sklearn.metrics import f1_score, accuracy_score, confusion_matrix
import matplotlib.pyplot as plt
X=df[['bill_length_mm', 'bill_depth_mm', 'flipper_length_mm', 'body_mass_g', 'sex', 'island', 'year']]
y=df['species']
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42,stratify=y)
folds=StratifiedKFold(n_splits=5,shuffle=True,random_state=42)
candidates={}
grids={}
# Prepare inputs and construct the candidate pipeline.
numeric_features = ['bill_length_mm', 'bill_depth_mm', 'flipper_length_mm', 'body_mass_g']
categorical_features = ['island', 'year']
encoded_binary_features = ['sex']
from sklearn.preprocessing import OneHotEncoder
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import StandardScaler
category_encoder = OneHotEncoder(handle_unknown='ignore', sparse_output=False, drop=None)
preprocessor = ColumnTransformer([('continuous', StandardScaler(), numeric_features), ('encoded', category_encoder, encoded_binary_features + categorical_features)], verbose_feature_names_out=False)
from sklearn.pipeline import Pipeline
from sklearn.linear_model import LogisticRegression
model = LogisticRegression(max_iter=2000, random_state=42)
pipeline = Pipeline([('prepare', preprocessor), ('model', model)])
model = pipeline
candidates['logistic']=model
grids['logistic']={'model__C': [0.1, 1.0, 10.0]}
# Validate initial candidates on matching training folds.
initial_results={name:cross_validate(candidate,X_train,y_train,cv=folds,scoring='f1_macro') for name,candidate in candidates.items()}
# Compare with a simple reference before tuning.
reference_results=cross_validate(DummyClassifier(strategy='most_frequent'),X_train,y_train,cv=folds,scoring='f1_macro')
rows=[]
selected={}
for name,candidate in candidates.items():
initial=initial_results[name]
grid=grids[name]
if grid:
search=GridSearchCV(candidate,grid,cv=folds,scoring='f1_macro',n_jobs=1)
search.fit(X_train,y_train)
selected[name]=search.best_estimator_
validation_score=float(search.best_score_)
settings=search.best_params_
else:
selected[name]=candidate
validation_score=float(np.mean(initial['test_score']))
settings='Keep defaults: no production search'
rows.append({'model':name,'initial_score':float(np.mean(initial['test_score'])),'selected_score':validation_score,'settings':str(settings)})
cv_results=pd.DataFrame(rows).set_index('model')
# This solution nominates by mean training-fold score. Other evidence-based choices can be defensible.
chosen_name=cv_results.selected_score.idxmax()
chosen=selected[chosen_name]
oof_predictions=cross_val_predict(chosen,X_train,y_train,cv=folds)
diagnostic_y=y_train
class_labels=sorted(y_train.unique())
matrix=confusion_matrix(diagnostic_y,oof_predictions,labels=class_labels)
fig,ax=plt.subplots(figsize=(6,4))
ax.imshow(matrix,cmap='Blues')
ax.set(xticks=range(len(class_labels)),yticks=range(len(class_labels)),xticklabels=class_labels,yticklabels=class_labels,xlabel='Predicted class',ylabel='Actual class',title='Training-only confusion matrix')
for (i,j),value in np.ndenumerate(matrix):
ax.text(j,i,str(value),ha='center',va='center',color='black')
fig.tight_layout()
fig.savefig('diagnostic.png',dpi=150,bbox_inches='tight')
final_model=clone(chosen).fit(X_train,y_train)
final_predictions=final_model.predict(X_test)
final_f1=f1_score(y_test,final_predictions,average='macro')
final_accuracy=accuracy_score(y_test,final_predictions)
print('Final macro F1:',final_f1,'; accuracy:',final_accuracy)
The ablation tests reliance on context using training evidence; stratification and macro F1 keep minority-class behaviour visible before final evaluation.
Helpful prior knowledge: Classification evidence retrieval · Neighbours, margins and rules retrieval · Trees and distributions retrieval These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.