Define X/y, split, fit, predict and evaluate a line.
This practice: Define X/y, split, fit, predict and evaluate a line. Next: Supervised Workflow. Continue to Supervised Workflow →
Given data · LINE24B
24 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.
| distance | duration |
|---|---|
| 1 | 9.73269 |
| 1.47826 | 12.4693 |
| 1.95652 | 10.113 |
| 2.43478 | 10.5783 |
| 2.91304 | 8.76362 |
| 3.3913 | 19.0888 |
| 3.86957 | 17.6587 |
| 4.34783 | 19.6607 |
Column meanings and units
Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.
Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.
| Column | Stored type |
|---|---|
| distance | float64 |
| duration | float64 |
Declared validation design
Reserve 20% of the declared population for the final test using split seed 42. Use a random split.
Supporting concepts: Measuring prediction error → · Could we know this at prediction time? → · Your first complete classification workflow →
Remember the idea
This checkpoint combines previously taught skills. Assemble the workflow; help remains available when needed.
Required Python variables and evidence
Use these names so Check can inspect your workflow. Each meaning is shown beside its name.
| Variable | Meaning |
|---|---|
| X | Feature dataframe for the declared population, preserving row indices. |
| X_test | Final-test feature rows from the declared split. |
| X_train | Training feature rows from the declared split. |
| answer | Requested result described in the task and deliverables. |
| df | Loaded and prepared input dataframe. |
| predictions | Requested result described in the task and deliverables. |
| y | Target series, aligned with X. |
| y_test | Final-test targets, aligned with X_test. |
Hint 1 — Think
Reconstruct the sequence that keeps evaluation rows outside learning.
Hint 2 — Tools
Feature/target selection, joint splitting, LinearRegression and RMSE.
Hint 3 — Approach
Define the schema, reserve rows, fit the line, predict the held-away features and score aligned outcomes.
Explained solution
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from sklearn.metrics import root_mean_squared_error
X=df[['distance']]
y=df.duration
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42)
model=LinearRegression().fit(X_train,y_train)
predictions=model.predict(X_test)
answer=root_mean_squared_error(y_test,predictions)
The split creates the evaluation boundary before fitting; prediction and scoring then describe the same held-away observations in duration units.
Helpful prior knowledge: Foundations retrieval · Honest evaluation retrieval · Your first complete classification workflow These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.