Understand the idea
A useful model can still omit important factors. Weak predictive results do not prove no relationship; fitted explanations do not establish interventions.
Python skill: Read two fitted summaries and distinguish a model coefficient from a causal claim.
Meet the syntax
model.coef_[0]
model.score(X, y)model.coef_[0]- The fitted change in prediction per one-unit increase in distance for this line.
model.score(X, y)- For LinearRegression this returns R² on the supplied rows; these are training rows here.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
slope = float(model.coef_[0])
training_r2 = float(model.score(X, y))
This practice: Read and run the Python. Next: Change · What fitted explanations cannot establish.
Given data · LINE24
24 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.
| distance | duration |
|---|---|
| 1 | 5.30501 |
| 1.47826 | 11.2361 |
| 1.95652 | 11.8488 |
| 2.43478 | 14.809 |
| 2.91304 | 10.7589 |
| 3.3913 | 19.4972 |
| 3.86957 | 17.89 |
| 4.34783 | 15.531 |
Column meanings and units
Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.
Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.
| Column | Stored type |
|---|---|
| distance | float64 |
| duration | float64 |
Supplied setup · available if you need to inspect it
This code runs before your editor on every Run. These are the objects your exercise uses.
from sklearn.linear_model import LinearRegression
X = df[['distance']]
y = df['duration']
model = LinearRegression().fit(X, y)Your task · Follow
Store the fitted distance coefficient in slope and R² on the training rows in training_r2. Explain why neither establishes a causal effect.
Hint 1 — Think
Read two fitted summaries and distinguish a model coefficient from a causal claim.
Hint 2 — Tools
Use model.coef_[0], model.score(X, y). Read the visible syntax meanings before editing.
Hint 3 — Approach
Store the fitted distance coefficient in slope and R² on the training rows in training_r2. Explain why neither establishes a causal effect. Keep the supplied row order and inspect the named output after running.
Explained solution
slope = float(model.coef_[0])
training_r2 = float(model.score(X, y))
A coefficient describes the fitted line for this dataset. Training R² describes how that line fits these seen rows. Neither controls confounders or shows what an intervention would do. Report the population and validation evidence separately from causal claims.
Helpful prior knowledge: Control tree complexity · Read residual patterns and limits These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.