Understand the idea
A regression tree repeatedly splits rows using thresholds. Each terminal leaf predicts an average training target. Threshold order matters; feature scaling is unnecessary.
Python skill: Creates a regression tree whose terminal leaves predict fitted target averages.
Meet the syntax
DecisionTreeRegressor(random_state=42)
plot_tree(model)
model.apply(row)DecisionTreeRegressor- Creates a regression tree whose terminal leaves predict fitted target averages.
random_state=42- Makes the tree-building choices reproducible.
plot_tree(model)- Draws the fitted split structure, leaf values and sample counts.
model.apply(row)- Returns the terminal leaf index reached by each feature row.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
from sklearn.tree import DecisionTreeRegressor,plot_tree
import matplotlib.pyplot as plt
model=DecisionTreeRegressor(max_depth=3,random_state=42).fit(X_train,y_train)
fig,ax=plt.subplots(figsize=(9,4))
plot_tree(model,feature_names=['x'],ax=ax)
answer=model.predict(X_test)
This practice: Read and run the Python. Next: Change · Trees predict with leaf averages.
Regression tree · from idea to workflow
Question: Can threshold rules predict a quantity?
Mechanism: Repeated feature cuts form leaves that predict training averages.
Watch for: Deep leaves overfit small samples; predictions cannot smoothly extrapolate beyond learned leaves.
Interpret or debug · self-review: Compare training and validation errors and inspect a poorly populated leaf. Explain when a shallower tree is preferable.
Guided full workflow: Apply this model with a baseline, training validation and one final evaluation →
Independent full workflow: ML-X03 · ML-X04. First practise complete regression and classification in F11–F12; check readiness in W-K2–W-K3.
Given data · STEP60
60 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.
| x | y |
|---|---|
| 0 | 10.4571 |
| 0.169492 | 8.44002 |
| 0.338983 | 11.1257 |
| 0.508475 | 11.4108 |
| 0.677966 | 7.07345 |
| 0.847458 | 8.04673 |
| 1.01695 | 10.1918 |
| 1.18644 | 9.52564 |
Column meanings and units
Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.
Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.
| Column | Stored type |
|---|---|
| x | float64 |
| y | float64 |
Supplied setup · available if you need to inspect it
This code runs before your editor on every Run. These are the objects your exercise uses.
from sklearn.model_selection import train_test_split
X=df[['x']]
y=df.y
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42)
Your task · Follow
Fit a depth-three tree to STEP60 and draw its splits and leaf values.
Hint 1 — Think
Each leaf predicts a quantity learned from the training outcomes reaching it.
Hint 2 — Tools
DecisionTreeRegressor, max_depth and plot_tree.
Hint 3 — Approach
Fit the declared tree, render its labelled splits and predict the held-away rows.
Explained solution
from sklearn.tree import DecisionTreeRegressor,plot_tree
import matplotlib.pyplot as plt
model=DecisionTreeRegressor(max_depth=3,random_state=42).fit(X_train,y_train)
fig,ax=plt.subplots(figsize=(9,4))
plot_tree(model,feature_names=['x'],ax=ax)
answer=model.predict(X_test)
The diagram connects threshold decisions to leaf values; the depth cap limits how finely the training population is partitioned.
Helpful prior knowledge: Supervised Workflow checkpoint These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.