Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Regression lessonsQUESTIONS · MODELS · EVIDENCE
Lines and evidence · ML-R02 · 12–18 MIN

RMSE and R² answer different questions

Distinguish target-unit errors from relative fit.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeReason about the Python
  3. TransferAdapt the Python

Understand the idea

RMSE measures errors in target units. R² compares squared error with variation around the evaluation-set mean. It can be negative. A trained mean dummy uses the training mean, so it is a distinct reference.

Distinguish target-unit errors from relative fit.RMSE: outcome unitsR²: relative squared errorHow large are errors?Compared with a mean?Use the same aligned outcomes and predictions.
Schematic · Distinguish target-unit errors from relative fit.Scroll the diagram horizontally if needed.

Python skill: Summarises error in original target units.

Meet the syntax

root_mean_squared_error(y_test, predictions)
r2_score(y_test, predictions)
root_mean_squared_error
Summarises error in original target units.
r2_score
Compares squared prediction error with variation around the evaluation-target mean; the result can be negative.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

from sklearn.metrics import root_mean_squared_error,r2_score
rmse=root_mean_squared_error(y_test,predictions)
r2=r2_score(y_test,predictions)

This practice: Read and run the Python. Next: Change · RMSE and R² answer different questions.

Given data · LINE24

24 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

LINE24 · first 8 prepared rows
distanceduration
15.30501
1.4782611.2361
1.9565211.8488
2.4347814.809
2.9130410.7589
3.391319.4972
3.8695717.89
4.3478315.531

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
distancefloat64
durationfloat64
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.model_selection import train_test_split
X = df[['distance']]
y = df['duration']
X_train, X_test, y_train, y_test = train_test_split(X,y,test_size=.2,random_state=42)
from sklearn.linear_model import LinearRegression
model = LinearRegression().fit(X_train,y_train)
predictions = model.predict(X_test)

Your task · Follow

Compute RMSE and R² for the supplied predictions; store them in rmse and r2.

Hint 1 — Think

One measure reports error in outcome units; the other compares squared error with a reference.

Hint 2 — Tools

root_mean_squared_error and r2_score.

Hint 3 — Approach

Apply both metrics to the same aligned outcomes and predictions, then store each named result.

Explained solution
from sklearn.metrics import root_mean_squared_error,r2_score
rmse=root_mean_squared_error(y_test,predictions)
r2=r2_score(y_test,predictions)

Using identical rows makes the two summaries complementary descriptions of the same prediction evidence.

Helpful prior knowledge: Read a fitted line These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Compute RMSE and R² for the supplied predictions; store them in rmse and r2.

rmser2

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.