Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Regression lessonsQUESTIONS · MODELS · EVIDENCE
Curves and trees · ML-R10 · 12–18 MIN

Control tree complexity

Use depth and minimum leaf size with validation evidence.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeAdapt the Python
  3. TransferExplain the Python result

Understand the idea

A deep tree can memorise small groups. Depth and minimum leaf size constrain flexibility. Impurity importance describes this fitted model, not causal importance.

Use depth and minimum leaf size with validation evidence.x ≤ threshold?yesnoTargets 8,10,12 → mean 10Targets 18,20 → mean 19Each leaf predicts its training target average.
Schematic · Use depth and minimum leaf size with validation evidence.Scroll the diagram horizontally if needed.

Python skill: Names the tree parameter that limits the number of splits along a path.

Meet the syntax

GridSearchCV(DecisionTreeRegressor(random_state=42), {'max_depth':[3,5,None]}, ...)
{'min_samples_leaf': [1, 5, 10]}
'max_depth'
Names the tree parameter that limits the number of splits along a path.
[3,5,None]
Compares two explicit depth limits with no depth cap; validation must judge flexibility.
{'min_samples_leaf': [1, 5, 10]}
Compares minimum observations allowed in a fitted leaf, a separate constraint from maximum depth.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

search=GridSearchCV(DecisionTreeRegressor(random_state=42),{'max_depth':[3,5,None]},cv=folds,scoring='neg_root_mean_squared_error').fit(X_train,y_train)
answer=-search.cv_results_['mean_test_score']

This practice: Read and run the Python. Next: Change · Control tree complexity.

Given data · STEP60

60 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

STEP60 · first 8 prepared rows
xy
010.4571
0.1694928.44002
0.33898311.1257
0.50847511.4108
0.6779667.07345
0.8474588.04673
1.0169510.1918
1.186449.52564

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
xfloat64
yfloat64
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.model_selection import train_test_split
X=df[['x']]
y=df.y
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42)

from sklearn.tree import DecisionTreeRegressor
from sklearn.model_selection import GridSearchCV,KFold
folds=KFold(5,shuffle=True,random_state=42)

Your task · Follow

Search the production depth choices with five training folds.

Hint 1 — Think

Depth should be chosen from held-out training evidence rather than test error.

Hint 2 — Tools

GridSearchCV, max_depth and negative-RMSE scoring.

Hint 3 — Approach

Compare the stated depth grid with the supplied folds, then convert mean scores to positive errors.

Explained solution
search=GridSearchCV(DecisionTreeRegressor(random_state=42),{'max_depth':[3,5,None]},cv=folds,scoring='neg_root_mean_squared_error').fit(X_train,y_train)
answer=-search.cv_results_['mean_test_score']

Matching folds makes complexity comparisons fair; sign conversion reports error in the familiar smaller-is-better direction.

Helpful prior knowledge: Trees predict with leaf averages · How a search makes a choice These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Search the production depth choices with five training folds.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.