Understand the idea
Cluster labels may be permuted without changing the grouping. Report original-unit feature summaries and population sizes so the result is interpretable.
Python skill: Groups original-unit rows by their aligned fitted assignments.
Meet the syntax
X.groupby(labels).mean()
pd.Series(labels).value_counts()X.groupby(labels)- Groups original-unit rows by their aligned fitted assignments.
.mean()- Computes each group’s mean per measurement; it does not change or refit the groups.
pd.Series(labels).value_counts()- Counts how many fitted observations belong to each named group.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
answer=X.groupby(labels).mean()
sizes=pd.Series(labels).value_counts().sort_index()
This practice: Read and run the Python. Next: Change · Describe clusters in meaningful units.
Given data · penguins
333 observations. One measured penguin. The dataframe df is supplied afresh for each Run.
Download source CSV · Source and original dictionary
Reference labels are omitted from this preview and must remain outside fitting.
| island | bill_length_mm | bill_depth_mm | flipper_length_mm | body_mass_g | sex | year |
|---|---|---|---|---|---|---|
| Torgersen | 39.1 | 18.7 | 181 | 3750 | male | 2007 |
| Torgersen | 39.5 | 17.4 | 186 | 3800 | female | 2007 |
| Torgersen | 40.3 | 18 | 195 | 3250 | female | 2007 |
| Torgersen | 36.7 | 19.3 | 193 | 3450 | female | 2007 |
| Torgersen | 39.3 | 20.6 | 190 | 3650 | male | 2007 |
| Torgersen | 38.9 | 17.8 | 181 | 3625 | female | 2007 |
| Torgersen | 39.2 | 19.6 | 195 | 4675 | male | 2007 |
| Torgersen | 41.1 | 17.6 | 182 | 3200 | female | 2007 |
Column meanings and units
Bill length/depth and flipper length: mm. Body mass: grams. Year and island: sampling context.
ML uses 333 complete cases. Removing incomplete records may change the represented population. Geographic context may not generalise to new islands.
| Column | Stored type |
|---|---|
| island | str |
| bill_length_mm | float64 |
| bill_depth_mm | float64 |
| flipper_length_mm | int64 |
| body_mass_g | int64 |
| sex | str |
| year | int64 |
Supplied setup · available if you need to inspect it
This code runs before your editor on every Run. These are the objects your exercise uses.
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
X=df[['bill_length_mm','bill_depth_mm','flipper_length_mm','body_mass_g']]
scaled=StandardScaler().fit_transform(X)
labels=KMeans(n_clusters=3,n_init=20,random_state=42).fit_predict(scaled)
Your task · Follow
Report cluster sizes and original-unit means.
Hint 1 — Think
Assignments must align with the original rows whose profiles you report.
Hint 2 — Tools
groupby.mean and value_counts.
Hint 3 — Approach
Use the supplied labels to aggregate original measurements and count group membership.
Explained solution
answer=X.groupby(labels).mean()
sizes=pd.Series(labels).value_counts().sort_index()
Original-unit means explain group characteristics while sizes reveal whether a profile describes many rows or a tiny subset.
Helpful prior knowledge: Choosing k requires evidence and judgement These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.