Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← ML Foundations lessonsQUESTIONS · MODELS · EVIDENCE
Honest evaluation · ML-F08 · 12–18 MIN

Predicting classes

Distinguish class labels from quantities.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeReason about the Python
  3. TransferExplain the Python result

Understand the idea

A classifier returns labels. The supplied tree recipe is only an example of the familiar fit/predict interface here; tree mechanics and depth choices come later.

Distinguish class labels from quantities.MeasurementsFitted classifierClass labelNew row 1 → ANew row 2 → CA label names a category, even when stored as a number.
Schematic · Distinguish class labels from quantities.Scroll the diagram horizontally if needed.

Python skill: Learns the supplied classification recipe from feature rows and class labels.

Meet the syntax

classifier.fit(X, y)
classifier.predict(new_rows)
classifier.fit(X, y)
Learns the supplied classification recipe from feature rows and class labels.
classifier.predict(new_rows)
Returns one predicted class label for each new feature row.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

from sklearn.tree import DecisionTreeClassifier
classifier=DecisionTreeClassifier(max_depth=2,random_state=42)
classifier.fit(df[['length','width']],df.label)
answer=classifier.predict(df[['length','width']])

This practice: Read and run the Python. Next: Change · Predicting classes.

Given data · CLASS18

18 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

CLASS18 · all rows
lengthwidthlabel
236.033-0.805568A
581.2970.728558A
-1511.27-1.00866A
99.0248-0.24496A
-13.0141-0.660765A
681.1790.602475A
2051.151.87316B
2362.130.334395B
2285.630.257253B
2680.440.961328B
1856.810.472554B
2946.980.880302B
-331.7812.72724C
412.3253.28307C
319.7013.33371C
1658.912.68519C
-396.7822.36965C
477.1363.8745C

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
lengthfloat64
widthfloat64
labelstr

Your task · Follow

Use the supplied classifier recipe to predict these observations’ class labels.

Hint 1 — Think

The output is a class label, even though the recipe uses numeric measurements.

Hint 2 — Tools

The supplied DecisionTreeClassifier recipe, fit and predict.

Hint 3 — Approach

Use the given classifier settings with the two measurement columns and label target, then obtain one label per row.

Explained solution
from sklearn.tree import DecisionTreeClassifier
classifier=DecisionTreeClassifier(max_depth=2,random_state=42)
classifier.fit(df[['length','width']],df.label)
answer=classifier.predict(df[['length','width']])

This exercise isolates the classifier interface; the supplied tree recipe is not a request to choose or tune a tree.

Helpful prior knowledge: Making a reproducible split These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Use the supplied classifier recipe to predict these observations’ class labels.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.