Understand the idea
A pattern in residuals suggests the model has left systematic structure unexplained. Association, prediction, extrapolation and causation are different claims.
Python skill: Defines signed residuals: positive values mean underprediction.
Meet the syntax
residual = actual - predictedactual - predicted- Defines signed residuals: positive values mean underprediction.
residual- Keeps row-level errors so a plot can reveal patterns hidden by a single summary.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
from sklearn.linear_model import LinearRegression
import matplotlib.pyplot as plt
model=LinearRegression().fit(X_train,y_train)
answer=y_test-model.predict(X_test)
fig,ax=plt.subplots()
ax.scatter(X_test.x,answer)
ax.axhline(0,color='black')
ax.set(xlabel='x',ylabel='Residual',title='Line fitted to curved observations')
This practice: Read and run the Python. Next: Change · Read residual patterns and limits.
Given data · CURVE48
48 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.
| x | y |
|---|---|
| -3 | 15.9571 |
| -2.87234 | 13.0709 |
| -2.74468 | 14.9362 |
| -2.61702 | 14.45 |
| -2.48936 | 9.39011 |
| -2.3617 | 9.68978 |
| -2.23404 | 11.2101 |
| -2.10638 | 9.96814 |
Column meanings and units
Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.
Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.
| Column | Stored type |
|---|---|
| x | float64 |
| y | float64 |
Supplied setup · available if you need to inspect it
This code runs before your editor on every Run. These are the objects your exercise uses.
from sklearn.model_selection import train_test_split
X=df[['x']]
y=df.y
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42)
Your task · Follow
Fit a line to CURVE48 training rows and plot residuals against x.
Hint 1 — Think
A residual keeps the direction of the prediction error.
Hint 2 — Tools
LinearRegression, actual-minus-predicted subtraction and scatter.
Hint 3 — Approach
Fit the training line, compute held-away residuals and plot them against x with a zero reference.
Explained solution
from sklearn.linear_model import LinearRegression
import matplotlib.pyplot as plt
model=LinearRegression().fit(X_train,y_train)
answer=y_test-model.predict(X_test)
fig,ax=plt.subplots()
ax.scatter(X_test.x,answer)
ax.axhline(0,color='black')
ax.set(xlabel='x',ylabel='Residual',title='Line fitted to curved observations')
Signed errors reveal systematic curvature that an aggregate RMSE hides; these supplied diagnostic rows are exploratory evidence, not a final untouched-test claim.
Helpful prior knowledge: RMSE and R² answer different questions These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.