Understand the idea
In supervised learning we had y. Here the goal is to describe structure in X. Cluster IDs are arbitrary identifiers, not predicted real-world classes. Reference labels must not guide fitting or selection.
Python skill: Contains only the measurements chosen to define similarity.
Meet the syntax
X = df[measurement_columns]measurement_columns- Contains only the measurements chosen to define similarity.
df[measurement_columns]- Creates X without a supervised target or hidden reference label.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
answer=df[['bill_length_mm','bill_depth_mm','flipper_length_mm','body_mass_g']]
This practice: Read and run the Python. Next: Change · Discovery has no prediction target.
Given data · penguins
333 observations. One measured penguin. The dataframe df is supplied afresh for each Run.
Download source CSV · Source and original dictionary
Reference labels are omitted from this preview and must remain outside fitting.
| island | bill_length_mm | bill_depth_mm | flipper_length_mm | body_mass_g | sex | year |
|---|---|---|---|---|---|---|
| Torgersen | 39.1 | 18.7 | 181 | 3750 | male | 2007 |
| Torgersen | 39.5 | 17.4 | 186 | 3800 | female | 2007 |
| Torgersen | 40.3 | 18 | 195 | 3250 | female | 2007 |
| Torgersen | 36.7 | 19.3 | 193 | 3450 | female | 2007 |
| Torgersen | 39.3 | 20.6 | 190 | 3650 | male | 2007 |
| Torgersen | 38.9 | 17.8 | 181 | 3625 | female | 2007 |
| Torgersen | 39.2 | 19.6 | 195 | 4675 | male | 2007 |
| Torgersen | 41.1 | 17.6 | 182 | 3200 | female | 2007 |
Column meanings and units
Bill length/depth and flipper length: mm. Body mass: grams. Year and island: sampling context.
ML uses 333 complete cases. Removing incomplete records may change the represented population. Geographic context may not generalise to new islands.
| Column | Stored type |
|---|---|
| island | str |
| bill_length_mm | float64 |
| bill_depth_mm | float64 |
| flipper_length_mm | int64 |
| body_mass_g | int64 |
| sex | str |
| year | int64 |
Your task · Follow
Store only bill length, bill depth, flipper length and body mass in answer, keeping every Penguin row and the stated column order.
Hint 1 — Think
Reference labels and context fields would change the meaning of measurement-based discovery.
Hint 2 — Tools
Explicit measurement-column selection.
Hint 3 — Approach
Keep only the four declared numeric measurements and inspect the resulting schema.
Explained solution
answer=df[['bill_length_mm','bill_depth_mm','flipper_length_mm','body_mass_g']]
Excluding species and context keeps grouping based on the intended measured similarities rather than known labels.
Helpful prior knowledge: ML Foundations checkpoint These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.