Understand the idea
Production fits the scale on the exploratory population, then transforms a reproducible sample of at most 500 rows for the hierarchy. The resulting labels belong to those sampled rows.
Python skill: Selects rows while preserving their original index labels.
Meet the syntax
sample = X.sample(min(500,len(X)), random_state=42)X.sample- Selects rows while preserving their original index labels.
min(500,len(X))- Caps the sample size at the smaller of 500 and the available row count.
random_state=42- Makes the selected sample reproducible.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
answer=X.sample(100,random_state=42)
This practice: Read and run the Python. Next: Change · Sampled hierarchies describe sampled rows.
Given data · penguins
333 observations. One measured penguin. The dataframe df is supplied afresh for each Run.
Download source CSV · Source and original dictionary
Reference labels are omitted from this preview and must remain outside fitting.
| island | bill_length_mm | bill_depth_mm | flipper_length_mm | body_mass_g | sex | year |
|---|---|---|---|---|---|---|
| Torgersen | 39.1 | 18.7 | 181 | 3750 | male | 2007 |
| Torgersen | 39.5 | 17.4 | 186 | 3800 | female | 2007 |
| Torgersen | 40.3 | 18 | 195 | 3250 | female | 2007 |
| Torgersen | 36.7 | 19.3 | 193 | 3450 | female | 2007 |
| Torgersen | 39.3 | 20.6 | 190 | 3650 | male | 2007 |
| Torgersen | 38.9 | 17.8 | 181 | 3625 | female | 2007 |
| Torgersen | 39.2 | 19.6 | 195 | 4675 | male | 2007 |
| Torgersen | 41.1 | 17.6 | 182 | 3200 | female | 2007 |
Column meanings and units
Bill length/depth and flipper length: mm. Body mass: grams. Year and island: sampling context.
ML uses 333 complete cases. Removing incomplete records may change the represented population. Geographic context may not generalise to new islands.
| Column | Stored type |
|---|---|
| island | str |
| bill_length_mm | float64 |
| bill_depth_mm | float64 |
| flipper_length_mm | int64 |
| body_mass_g | int64 |
| sex | str |
| year | int64 |
Supplied setup · available if you need to inspect it
This code runs before your editor on every Run. These are the objects your exercise uses.
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
X=df[['bill_length_mm','bill_depth_mm','flipper_length_mm','body_mass_g']]
scaled=StandardScaler().fit_transform(X)
labels=KMeans(n_clusters=3,n_init=20,random_state=42).fit_predict(scaled)
Your task · Follow
Sample 100 Penguin measurement rows while retaining their original index.
Hint 1 — Think
Original indices preserve the link between sampled observations and their measurements.
Hint 2 — Tools
DataFrame.sample with random_state.
Hint 3 — Approach
Take the requested reproducible sample without resetting its index.
Explained solution
answer=X.sample(100,random_state=42)
Retaining row identity lets later assignments and original-unit profiles remain aligned to the sampled population.
Helpful prior knowledge: Turn a hierarchy into groups · Describe clusters in meaningful units These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.