Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Supervised Workflow lessonsQUESTIONS · MODELS · EVIDENCE
Select, diagnose and finish · ML-W13 · 12–18 MIN

Finish once, then report

Separate selection from final evidence.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeAdapt the Python
  3. TransferReason about the Python

Understand the idea

After selecting a workflow using training evidence, refit it on all training rows and predict final-test rows. Reuse those predictions for reporting; choosing new settings after inspecting final errors is exploratory.

Separate selection from final evidence.ObservationsTraining → fit / validateFinal testProtect final rows from preparation and selection.
Schematic · Separate selection from final evidence.Scroll the diagram horizontally if needed.

Python skill: Copies the chosen estimator configuration without carrying fitted state.

Meet the syntax

final_model = clone(chosen).fit(X_train, y_train)
clone(chosen)
Copies the chosen estimator configuration without carrying fitted state.
fit(X_train, y_train)
Refits that fixed choice on all training rows before the final evaluation.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

from sklearn.base import clone
from sklearn.metrics import root_mean_squared_error
final_model=clone(model).fit(X_train,y_train)
final_predictions=final_model.predict(X_test)
answer=root_mean_squared_error(y_test,final_predictions)

This practice: Read and run the Python. Next: Change · Finish once, then report.

Given data · LINE24

24 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

LINE24 · first 8 prepared rows
distanceduration
15.30501
1.4782611.2361
1.9565211.8488
2.4347814.809
2.9130410.7589
3.391319.4972
3.8695717.89
4.3478315.531

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
distancefloat64
durationfloat64
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.model_selection import train_test_split
X = df[['distance']]
y = df['duration']
X_train, X_test, y_train, y_test = train_test_split(X,y,test_size=.2,random_state=42)
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import KFold,cross_validate
folds=KFold(5,shuffle=True,random_state=42)
model=LinearRegression()

Your task · Follow

Refit the supplied chosen line and compute final RMSE.

Hint 1 — Think

The candidate choice is already fixed before the reserved rows are evaluated.

Hint 2 — Tools

clone, fit, predict and root_mean_squared_error.

Hint 3 — Approach

Clone the chosen recipe, fit all training rows, save final predictions and calculate their RMSE.

Explained solution
from sklearn.base import clone
from sklearn.metrics import root_mean_squared_error
final_model=clone(model).fit(X_train,y_train)
final_predictions=final_model.predict(X_test)
answer=root_mean_squared_error(y_test,final_predictions)

Cloning removes stale fitted state; the final evaluation estimates the fixed workflow rather than selecting another setting.

Helpful prior knowledge: Diagnose without opening the final test These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Refit the supplied chosen line and compute final RMSE.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.