Understand the idea
After selecting a workflow using training evidence, refit it on all training rows and predict final-test rows. Reuse those predictions for reporting; choosing new settings after inspecting final errors is exploratory.
Python skill: Copies the chosen estimator configuration without carrying fitted state.
Meet the syntax
final_model = clone(chosen).fit(X_train, y_train)clone(chosen)- Copies the chosen estimator configuration without carrying fitted state.
fit(X_train, y_train)- Refits that fixed choice on all training rows before the final evaluation.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
from sklearn.base import clone
from sklearn.metrics import root_mean_squared_error
final_model=clone(model).fit(X_train,y_train)
final_predictions=final_model.predict(X_test)
answer=root_mean_squared_error(y_test,final_predictions)
This practice: Read and run the Python. Next: Change · Finish once, then report.
Given data · LINE24
24 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.
| distance | duration |
|---|---|
| 1 | 5.30501 |
| 1.47826 | 11.2361 |
| 1.95652 | 11.8488 |
| 2.43478 | 14.809 |
| 2.91304 | 10.7589 |
| 3.3913 | 19.4972 |
| 3.86957 | 17.89 |
| 4.34783 | 15.531 |
Column meanings and units
Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.
Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.
| Column | Stored type |
|---|---|
| distance | float64 |
| duration | float64 |
Supplied setup · available if you need to inspect it
This code runs before your editor on every Run. These are the objects your exercise uses.
from sklearn.model_selection import train_test_split
X = df[['distance']]
y = df['duration']
X_train, X_test, y_train, y_test = train_test_split(X,y,test_size=.2,random_state=42)
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import KFold,cross_validate
folds=KFold(5,shuffle=True,random_state=42)
model=LinearRegression()
Your task · Follow
Refit the supplied chosen line and compute final RMSE.
Hint 1 — Think
The candidate choice is already fixed before the reserved rows are evaluated.
Hint 2 — Tools
clone, fit, predict and root_mean_squared_error.
Hint 3 — Approach
Clone the chosen recipe, fit all training rows, save final predictions and calculate their RMSE.
Explained solution
from sklearn.base import clone
from sklearn.metrics import root_mean_squared_error
final_model=clone(model).fit(X_train,y_train)
final_predictions=final_model.predict(X_test)
answer=root_mean_squared_error(y_test,final_predictions)
Cloning removes stale fitted state; the final evaluation estimates the fixed workflow rather than selecting another setting.
Helpful prior knowledge: Diagnose without opening the final test These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.