Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← ML Foundations lessonsQUESTIONS · MODELS · EVIDENCE
Honest evaluation · ML-F09 · 12–18 MIN

Preserving class representation

Apply and explain stratification.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeAdapt the Python
  3. TransferAdapt the Python

Understand the idea

Stratification preserves approximate class proportions across a random split. It does not create additional minority examples or solve all sampling problems.

Apply and explain stratification.Illustrative population: 8 A + 4 BTraining: 6 A + 3 BTest: 2 A + 1 BStratification preserves class proportions approximately.
Schematic · Apply and explain stratification.Scroll the diagram horizontally if needed.

Python skill: Uses the target labels to approximately preserve class proportions in each split.

Meet the syntax

train_test_split(X, y, stratify=y, random_state=42)
stratify=y
Uses the target labels to approximately preserve class proportions in each split.
random_state=42
Keeps the comparison of split strategies reproducible.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

from sklearn.model_selection import train_test_split
X=df[['length','width']]
y=df.label
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42,stratify=y)
answer=y_test.value_counts().sort_index()

This practice: Read and run the Python. Next: Change · Preserving class representation.

Given data · CLASS180

180 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

CLASS180 · first 8 prepared rows
lengthwidthlabel
236.033-0.805568A
581.2970.728558A
-1511.27-1.00866A
99.0248-0.24496A
-13.0141-0.660765A
681.1790.602475A
51.14720.873157A
362.131-0.665605A

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
lengthfloat64
widthfloat64
labelstr

Your task · Follow

Split CLASS180 with stratification and inspect test class counts.

Hint 1 — Think

Small class groups can disappear from an evaluation partition by chance.

Hint 2 — Tools

train_test_split with stratify and value_counts.

Hint 3 — Approach

Split features and labels together using the declared seed and stratification, then count test labels.

Explained solution
from sklearn.model_selection import train_test_split
X=df[['length','width']]
y=df.label
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42,stratify=y)
answer=y_test.value_counts().sort_index()

Stratification approximately preserves class representation; counting the labels makes the partition visible.

Helpful prior knowledge: Predicting classes These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Split CLASS180 with stratification and inspect test class counts.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.