Complete the Penguin neural classifier workflow with a reference, width selection, class diagnostics and final metrics.
This practice: Complete the Penguin neural classifier workflow with a reference, width selection, class diagnostics and final metrics. Next: Clustering and Discovery. Continue to Clustering and Discovery →
Given data · penguins
333 observations. One measured penguin. The dataframe df is supplied afresh for each Run.
Download source CSV · Source and original dictionary
| island | bill_length_mm | bill_depth_mm | flipper_length_mm | body_mass_g | sex | year | species |
|---|---|---|---|---|---|---|---|
| Torgersen | 39.1 | 18.7 | 181 | 3750 | male | 2007 | Adelie |
| Torgersen | 39.5 | 17.4 | 186 | 3800 | female | 2007 | Adelie |
| Torgersen | 40.3 | 18 | 195 | 3250 | female | 2007 | Adelie |
| Torgersen | 36.7 | 19.3 | 193 | 3450 | female | 2007 | Adelie |
| Torgersen | 39.3 | 20.6 | 190 | 3650 | male | 2007 | Adelie |
| Torgersen | 38.9 | 17.8 | 181 | 3625 | female | 2007 | Adelie |
| Torgersen | 39.2 | 19.6 | 195 | 4675 | male | 2007 | Adelie |
| Torgersen | 41.1 | 17.6 | 182 | 3200 | female | 2007 | Adelie |
Column meanings and units
Bill length/depth and flipper length: mm. Body mass: grams. Year and island: sampling context.
ML uses 333 complete cases. Removing incomplete records may change the represented population. Geographic context may not generalise to new islands.
| Column | Stored type |
|---|---|
| island | str |
| bill_length_mm | float64 |
| bill_depth_mm | float64 |
| flipper_length_mm | int64 |
| body_mass_g | int64 |
| sex | str |
| year | int64 |
| species | str |
Declared validation design
Reserve 20% of the declared population for the final test using split seed 42. Keep class proportions with a stratified split.
Use the same five stratified training folds, shuffled with seed 42, for reference comparison, candidates and selection. Use out-of-fold training predictions for diagnosis. Select with macro F1 and compare with a most-frequent-class reference. Accuracy is supplementary. Open final-test evidence after selection and diagnosis.
Supporting concepts: Neural classification workflow → · Convergence, early stopping and validation →
Remember the idea
This checkpoint combines previously taught skills. Assemble the workflow; help remains available when needed.
Required Python variables and evidence
Use these names so Check can inspect your workflow. Each meaning is shown beside its name.
| Variable | Meaning |
|---|---|
| X | Feature dataframe for the declared population, preserving row indices. |
| X_test | Final-test feature rows from the declared split. |
| X_train | Training feature rows from the declared split. |
| chosen_name | Model ID nominated from training evidence. |
| cv_results | Candidate validation evidence; comparison workflows use a dataframe indexed by model ID. |
| df | Loaded and prepared input dataframe. |
| final_accuracy | Final-test accuracy, supplementary to macro F1. |
| final_f1 | Macro F1 on final-test class predictions. |
| final_model | Chosen pipeline fitted on training rows, after selection and diagnosis. |
| final_predictions | Unaltered predictions of final_model on X_test. |
| loss_curves | Dictionary of available neural training-loss histories by model ID. |
| matrix | Confusion matrix of training-only diagnostic predictions. |
| reference_results | cross_validate result for the dummy reference; test_score contains five scores. |
| selected | Dictionary mapping each requested model ID to its chosen pipeline. |
| y | Target series, aligned with X. |
| y_test | Final-test targets, aligned with X_test. |
| y_train | Training targets, aligned with X_train. |
Hint 1 — Think
Reconstruct a classification workflow that inspects optimisation as well as predictive error.
Hint 2 — Tools
Scaled MLP pipeline, stratified width search, OOF confusion and loss histories.
Hint 3 — Approach
Compare reference and neural candidates, nominate width, inspect convergence and class errors, then fit the fixed recipe for final reporting.
Explained solution
from sklearn.base import clone
from sklearn.model_selection import train_test_split, StratifiedKFold, cross_validate, cross_val_predict, GridSearchCV
from sklearn.dummy import DummyClassifier
from sklearn.metrics import f1_score, accuracy_score, confusion_matrix
import matplotlib.pyplot as plt
X=df[['bill_length_mm', 'bill_depth_mm', 'flipper_length_mm', 'body_mass_g']]
y=df['species']
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42,stratify=y)
folds=StratifiedKFold(n_splits=5,shuffle=True,random_state=42)
candidates={}
grids={}
# Prepare inputs and construct the candidate pipeline.
from sklearn.preprocessing import StandardScaler
preprocessor = StandardScaler()
from sklearn.pipeline import Pipeline
from sklearn.neural_network import MLPClassifier
model = MLPClassifier(hidden_layer_sizes=(24,), max_iter=500, early_stopping=True, random_state=42)
pipeline = Pipeline([('prepare', preprocessor), ('model', model)])
model = pipeline
candidates['mlp_cls']=model
grids['mlp_cls']={'model__hidden_layer_sizes': [(16,), (24,)]}
# Validate initial candidates on matching training folds.
initial_results={name:cross_validate(candidate,X_train,y_train,cv=folds,scoring='f1_macro') for name,candidate in candidates.items()}
# Compare with a simple reference before tuning.
reference_results=cross_validate(DummyClassifier(strategy='most_frequent'),X_train,y_train,cv=folds,scoring='f1_macro')
rows=[]
selected={}
for name,candidate in candidates.items():
initial=initial_results[name]
grid=grids[name]
if grid:
search=GridSearchCV(candidate,grid,cv=folds,scoring='f1_macro',n_jobs=1)
search.fit(X_train,y_train)
selected[name]=search.best_estimator_
validation_score=float(search.best_score_)
settings=search.best_params_
else:
selected[name]=candidate
validation_score=float(np.mean(initial['test_score']))
settings='Keep defaults: no production search'
rows.append({'model':name,'initial_score':float(np.mean(initial['test_score'])),'selected_score':validation_score,'settings':str(settings)})
cv_results=pd.DataFrame(rows).set_index('model')
# This solution nominates by mean training-fold score. Other evidence-based choices can be defensible.
chosen_name=cv_results.selected_score.idxmax()
chosen=selected[chosen_name]
oof_predictions=cross_val_predict(chosen,X_train,y_train,cv=folds)
diagnostic_y=y_train
class_labels=sorted(y_train.unique())
matrix=confusion_matrix(diagnostic_y,oof_predictions,labels=class_labels)
fig,ax=plt.subplots(figsize=(6,4))
ax.imshow(matrix,cmap='Blues')
ax.set(xticks=range(len(class_labels)),yticks=range(len(class_labels)),xticklabels=class_labels,yticklabels=class_labels,xlabel='Predicted class',ylabel='Actual class',title='Training-only confusion matrix')
for (i,j),value in np.ndenumerate(matrix):
ax.text(j,i,str(value),ha='center',va='center',color='black')
fig.tight_layout()
fig.savefig('diagnostic.png',dpi=150,bbox_inches='tight')
final_model=clone(chosen).fit(X_train,y_train)
final_predictions=final_model.predict(X_test)
final_f1=f1_score(y_test,final_predictions,average='macro')
final_accuracy=accuracy_score(y_test,final_predictions)
print('Final macro F1:',final_f1,'; accuracy:',final_accuracy)
loss_curves={}
for name,candidate in selected.items():
fitted=candidate.named_steps['model']
if hasattr(fitted,'regressor_'):
fitted=fitted.regressor_
if hasattr(fitted,'loss_curve_'):
loss_curves[name]=list(fitted.loss_curve_)
Outer validation selects the complete workflow while loss traces qualify optimisation; final metrics come only after those development decisions.
Helpful prior knowledge: Neural classification workflow · Shared network retrieval These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.