Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Classification lessonsQUESTIONS · MODELS · EVIDENCE
Margins and rules · ML-C14 · 18–25 MIN

Classification trees

Learn recursive splits, impurity, class leaves and validation within Classification.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeReason about the Python
  3. PractiseAdapt the Python
  4. TransferExplain the Python result

Understand the idea

A classification tree asks a sequence of threshold questions. A useful split separates class counts, reducing impurity. A leaf predicts from the classes reaching it. Scaling is unnecessary; depth and minimum leaf size constrain overfitting.

Learn recursive splits, impurity, class leaves and validation within Classification.Width ≤ threshold?yesno8 A, 1 B → predict A1 A, 9 B → predict BPurer child class counts reduce impurity.
Schematic · Learn recursive splits, impurity, class leaves and validation within Classification.Scroll the diagram horizontally if needed.

Python skill: Creates a tree that splits feature space and predicts from each leaf’s class distribution.

Meet the syntax

DecisionTreeClassifier(random_state=42)
plot_tree(model)
model.apply(row)
DecisionTreeClassifier
Creates a tree that splits feature space and predicts from each leaf’s class distribution.
random_state=42
Makes fitted split choices reproducible.
plot_tree(model)
Displays thresholds, class counts and leaf predictions of a fitted tree.
model.apply(row)
Finds the leaf reached by a row; predict returns the class selected at that leaf.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

from sklearn.tree import DecisionTreeClassifier
model=DecisionTreeClassifier(max_depth=3,random_state=42).fit(X_train,y_train)
leaf=int(model.apply(X_test.iloc[:1])[0])
label=model.predict(X_test.iloc[:1])[0]

This practice: Read and run the Python. Next: Change · Classification trees.

Classification tree · from idea to workflow

Question: Can a short sequence of conditions predict a class?

Mechanism: Each split concentrates class labels; leaves predict class evidence.

Watch for: Tiny leaves memorise observations; class imbalance can hide rare-class failures.

Interpret or debug · self-review: Inspect minority-class validation errors and depth. Repair overfitting using training-only evidence.

Independent full workflow: ML-X06. First practise complete regression and classification in F11–F12; check readiness in W-K2–W-K3.

Given data · CLASS180

180 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

CLASS180 · first 8 prepared rows
lengthwidthlabel
236.033-0.805568A
581.2970.728558A
-1511.27-1.00866A
99.0248-0.24496A
-13.0141-0.660765A
681.1790.602475A
51.14720.873157A
362.131-0.665605A

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
lengthfloat64
widthfloat64
labelstr
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.model_selection import train_test_split
X=df[['length','width']]
y=df.label
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42,stratify=y)

Your task · Follow

Fit a tree with maximum depth 3. For the first test row, store its terminal node ID in leaf and its predicted class in label.

Hint 1 — Think

A tree follows feature thresholds to a leaf with a learned class prediction.

Hint 2 — Tools

DecisionTreeClassifier, apply and predict.

Hint 3 — Approach

Fit the declared depth-limited tree, then obtain the first test row’s leaf and label.

Explained solution
from sklearn.tree import DecisionTreeClassifier
model=DecisionTreeClassifier(max_depth=3,random_state=42).fit(X_train,y_train)
leaf=int(model.apply(X_test.iloc[:1])[0])
label=model.predict(X_test.iloc[:1])[0]

The two outputs connect tree traversal with classification without requiring any regression-tree prerequisite.

Helpful prior knowledge: Macro F1 and imbalance · How a search makes a choice These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Fit a tree with maximum depth 3. For the first test row, store its terminal node ID in leaf and its predicted class in label.

leaflabel

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.