Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Regression lessonsQUESTIONS · MODELS · EVIDENCE
Curves and trees · ML-R09 · 25–40 MIN

Trees predict with leaf averages

Learn recursive splits and regression leaves within this branch.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeAdapt the Python
  3. TransferAdapt the Python
  4. ApplyBuild a guided full workflow

Understand the idea

A regression tree repeatedly splits rows using thresholds. Each terminal leaf predicts an average training target. Threshold order matters; feature scaling is unnecessary.

Learn recursive splits and regression leaves within this branch.x ≤ threshold?yesnoTargets 8,10,12 → mean 10Targets 18,20 → mean 19Each leaf predicts its training target average.
Schematic · Learn recursive splits and regression leaves within this branch.Scroll the diagram horizontally if needed.

Python skill: Creates a regression tree whose terminal leaves predict fitted target averages.

Meet the syntax

DecisionTreeRegressor(random_state=42)
plot_tree(model)
model.apply(row)
DecisionTreeRegressor
Creates a regression tree whose terminal leaves predict fitted target averages.
random_state=42
Makes the tree-building choices reproducible.
plot_tree(model)
Draws the fitted split structure, leaf values and sample counts.
model.apply(row)
Returns the terminal leaf index reached by each feature row.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

from sklearn.tree import DecisionTreeRegressor,plot_tree
import matplotlib.pyplot as plt
model=DecisionTreeRegressor(max_depth=3,random_state=42).fit(X_train,y_train)
fig,ax=plt.subplots(figsize=(9,4))
plot_tree(model,feature_names=['x'],ax=ax)
answer=model.predict(X_test)

This practice: Read and run the Python. Next: Change · Trees predict with leaf averages.

Regression tree · from idea to workflow

Question: Can threshold rules predict a quantity?

Mechanism: Repeated feature cuts form leaves that predict training averages.

Watch for: Deep leaves overfit small samples; predictions cannot smoothly extrapolate beyond learned leaves.

Interpret or debug · self-review: Compare training and validation errors and inspect a poorly populated leaf. Explain when a shallower tree is preferable.

Guided full workflow: Apply this model with a baseline, training validation and one final evaluation →

Independent full workflow: ML-X03 · ML-X04. First practise complete regression and classification in F11–F12; check readiness in W-K2–W-K3.

Given data · STEP60

60 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

STEP60 · first 8 prepared rows
xy
010.4571
0.1694928.44002
0.33898311.1257
0.50847511.4108
0.6779667.07345
0.8474588.04673
1.0169510.1918
1.186449.52564

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
xfloat64
yfloat64
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.model_selection import train_test_split
X=df[['x']]
y=df.y
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42)

Your task · Follow

Fit a depth-three tree to STEP60 and draw its splits and leaf values.

Hint 1 — Think

Each leaf predicts a quantity learned from the training outcomes reaching it.

Hint 2 — Tools

DecisionTreeRegressor, max_depth and plot_tree.

Hint 3 — Approach

Fit the declared tree, render its labelled splits and predict the held-away rows.

Explained solution
from sklearn.tree import DecisionTreeRegressor,plot_tree
import matplotlib.pyplot as plt
model=DecisionTreeRegressor(max_depth=3,random_state=42).fit(X_train,y_train)
fig,ax=plt.subplots(figsize=(9,4))
plot_tree(model,feature_names=['x'],ax=ax)
answer=model.predict(X_test)

The diagram connects threshold decisions to leaf values; the depth cap limits how finely the training population is partitioned.

Helpful prior knowledge: Supervised Workflow checkpoint These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Fit a depth-three tree to STEP60 and draw its splits and leaf values.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.