Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← ML Foundations lessonsQUESTIONS · MODELS · EVIDENCE
Learning from examples · ML-F05 · 12–18 MIN

Measuring prediction error

Connect residuals and RMSE to target units.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeAdapt the Python
  3. TransferReason about the Python

Understand the idea

A residual is actual minus predicted. RMSE squares residuals, averages them, and returns to target units with a square root. Large errors have more influence.

Connect residuals and RMSE to target units.actual −predictedResiduals measure signed vertical differences.
Schematic · Connect residuals and RMSE to target units.Scroll the diagram horizontally if needed.

Python skill: Compares aligned actual values and predictions, returning one error summary in the target units.

Meet the syntax

root_mean_squared_error(actual, predicted)
residual = actual - predicted
root_mean_squared_error(actual, predicted)
Compares aligned actual values and predictions, returning one error summary in the target units.
actual
The observed outcomes for these evaluation rows.
predicted
The model outputs for those same rows, in the same order.
residual = actual - predicted
Keeps signed row-level errors; positive residuals mean the model predicted too little.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

answer=pd.DataFrame({'actual':y_test,'predicted':predictions,'residual':y_test-predictions})

This practice: Read and run the Python. Next: Change · Measuring prediction error.

Given data · LINE24

24 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

LINE24 · first 8 prepared rows
distanceduration
15.30501
1.4782611.2361
1.9565211.8488
2.4347814.809
2.9130410.7589
3.391319.4972
3.8695717.89
4.3478315.531

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
distancefloat64
durationfloat64
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.model_selection import train_test_split
X = df[['distance']]
y = df['duration']
X_train, X_test, y_train, y_test = train_test_split(X,y,test_size=.2,random_state=42)
from sklearn.linear_model import LinearRegression
model = LinearRegression().fit(X_train,y_train)
predictions = model.predict(X_test)

Your task · Follow

Create answer with actual, predicted and residual columns for evaluation rows.

Hint 1 — Think

The sign of an error matters when you want to see underprediction and overprediction.

Hint 2 — Tools

pd.DataFrame and aligned subtraction.

Hint 3 — Approach

Pair each evaluation outcome with its supplied prediction, then calculate actual minus predicted.

Explained solution
answer=pd.DataFrame({'actual':y_test,'predicted':predictions,'residual':y_test-predictions})

The table exposes row-level signed errors that a single aggregate error score would hide.

Helpful prior knowledge: Predicting new rows These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Create answer with actual, predicted and residual columns for evaluation rows.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.