Understand the idea
Stratification preserves approximate class proportions across a random split. It does not create additional minority examples or solve all sampling problems.
Python skill: Uses the target labels to approximately preserve class proportions in each split.
Meet the syntax
train_test_split(X, y, stratify=y, random_state=42)stratify=y- Uses the target labels to approximately preserve class proportions in each split.
random_state=42- Keeps the comparison of split strategies reproducible.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
from sklearn.model_selection import train_test_split
X=df[['length','width']]
y=df.label
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42,stratify=y)
answer=y_test.value_counts().sort_index()
This practice: Read and run the Python. Next: Change · Preserving class representation.
Given data · CLASS180
180 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.
| length | width | label |
|---|---|---|
| 236.033 | -0.805568 | A |
| 581.297 | 0.728558 | A |
| -1511.27 | -1.00866 | A |
| 99.0248 | -0.24496 | A |
| -13.0141 | -0.660765 | A |
| 681.179 | 0.602475 | A |
| 51.1472 | 0.873157 | A |
| 362.131 | -0.665605 | A |
Column meanings and units
Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.
Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.
| Column | Stored type |
|---|---|
| length | float64 |
| width | float64 |
| label | str |
Your task · Follow
Split CLASS180 with stratification and inspect test class counts.
Hint 1 — Think
Small class groups can disappear from an evaluation partition by chance.
Hint 2 — Tools
train_test_split with stratify and value_counts.
Hint 3 — Approach
Split features and labels together using the declared seed and stratification, then count test labels.
Explained solution
from sklearn.model_selection import train_test_split
X=df[['length','width']]
y=df.label
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42,stratify=y)
answer=y_test.value_counts().sort_index()
Stratification approximately preserves class representation; counting the labels makes the partition visible.
Helpful prior knowledge: Predicting classes These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.