Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← ML Foundations lessonsQUESTIONS · MODELS · EVIDENCE
Honest evaluation · ML-F-K1 · 25–35 MIN

ML Foundations checkpoint

Retrieval 1

Exercises within this concept

  1. Retrieval 1Retrieve and applyCurrent exercise
Retrieval 1 · LINE24B

Define X/y, split, fit, predict and evaluate a line.

This practice: Define X/y, split, fit, predict and evaluate a line. Next: Supervised Workflow. Continue to Supervised Workflow →

Given data · LINE24B

24 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

LINE24B · first 8 prepared rows
distanceduration
19.73269
1.4782612.4693
1.9565210.113
2.4347810.5783
2.913048.76362
3.391319.0888
3.8695717.6587
4.3478319.6607

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
distancefloat64
durationfloat64

Declared validation design

Reserve 20% of the declared population for the final test using split seed 42. Use a random split.

Supporting concepts: Measuring prediction error → · Could we know this at prediction time? → · Your first complete classification workflow →

Remember the idea

This checkpoint combines previously taught skills. Assemble the workflow; help remains available when needed.

Define X/y, split, fit, predict and evaluate a line.Training XScale numbersEncode categoriesFit estimatorEach fold learns its own preparation.
Schematic · Define X/y, split, fit, predict and evaluate a line.Scroll the diagram horizontally if needed.

Required Python variables and evidence

Use these names so Check can inspect your workflow. Each meaning is shown beside its name.

Workflow evidence contract
VariableMeaning
XFeature dataframe for the declared population, preserving row indices.
X_testFinal-test feature rows from the declared split.
X_trainTraining feature rows from the declared split.
answerRequested result described in the task and deliverables.
dfLoaded and prepared input dataframe.
predictionsRequested result described in the task and deliverables.
yTarget series, aligned with X.
y_testFinal-test targets, aligned with X_test.
Hint 1 — Think

Reconstruct the sequence that keeps evaluation rows outside learning.

Hint 2 — Tools

Feature/target selection, joint splitting, LinearRegression and RMSE.

Hint 3 — Approach

Define the schema, reserve rows, fit the line, predict the held-away features and score aligned outcomes.

Explained solution
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from sklearn.metrics import root_mean_squared_error
X=df[['distance']]
y=df.duration
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42)
model=LinearRegression().fit(X_train,y_train)
predictions=model.predict(X_test)
answer=root_mean_squared_error(y_test,predictions)

The split creates the evaluation boundary before fitting; prediction and scoring then describe the same held-away observations in duration units.

Helpful prior knowledge: Foundations retrieval · Honest evaluation retrieval · Your first complete classification workflow These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Retrieval 1

Define X/y, split, fit, predict and evaluate a line.

answerpredictions

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.

What does the held-out delivery error support about predictions for new rows?

Use your Run output as evidence. This response is optional, not machine-graded or saved.