Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Regression lessonsQUESTIONS · MODELS · EVIDENCE
Curves and trees · ML-R11 · 12–18 MIN

What fitted explanations cannot establish

Match claim strength to the evidence.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeExplain the Python result
  3. TransferExplain the Python result

Understand the idea

A useful model can still omit important factors. Weak predictive results do not prove no relationship; fitted explanations do not establish interventions.

Match claim strength to the evidence.Observed: predictive error under this evaluationQualify: population, model, inputs and uncertaintyPrediction evidence alone does not identify causal effects.
Schematic · Match claim strength to the evidence.Scroll the diagram horizontally if needed.

Python skill: Read two fitted summaries and distinguish a model coefficient from a causal claim.

Meet the syntax

model.coef_[0]
model.score(X, y)
model.coef_[0]
The fitted change in prediction per one-unit increase in distance for this line.
model.score(X, y)
For LinearRegression this returns R² on the supplied rows; these are training rows here.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

slope = float(model.coef_[0])
training_r2 = float(model.score(X, y))

This practice: Read and run the Python. Next: Change · What fitted explanations cannot establish.

Given data · LINE24

24 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

LINE24 · first 8 prepared rows
distanceduration
15.30501
1.4782611.2361
1.9565211.8488
2.4347814.809
2.9130410.7589
3.391319.4972
3.8695717.89
4.3478315.531

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
distancefloat64
durationfloat64
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.linear_model import LinearRegression
X = df[['distance']]
y = df['duration']
model = LinearRegression().fit(X, y)

Your task · Follow

Store the fitted distance coefficient in slope and R² on the training rows in training_r2. Explain why neither establishes a causal effect.

Hint 1 — Think

Read two fitted summaries and distinguish a model coefficient from a causal claim.

Hint 2 — Tools

Use model.coef_[0], model.score(X, y). Read the visible syntax meanings before editing.

Hint 3 — Approach

Store the fitted distance coefficient in slope and R² on the training rows in training_r2. Explain why neither establishes a causal effect. Keep the supplied row order and inspect the named output after running.

Explained solution
slope = float(model.coef_[0])
training_r2 = float(model.score(X, y))

A coefficient describes the fitted line for this dataset. Training R² describes how that line fits these seen rows. Neither controls confounders or shows what an intervention would do. Report the population and validation evidence separately from causal claims.

Helpful prior knowledge: Control tree complexity · Read residual patterns and limits These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Store the fitted distance coefficient in slope and R² on the training rows in training_r2. Explain why neither establishes a causal effect.

slopetraining_r2

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.