Understand the idea
Reference labels may be inspected after fitting for interpretation, but must not secretly determine features, k or the hierarchy cut. Discovery describes a chosen population under chosen measurements and geometry.
Python skill: Use assign() and groupby() to interpret clusters in original units after fitting.
Meet the syntax
X.assign(cluster=labels)
groupby('cluster')
.mean()X.assign(cluster=labels)- Attaches the assignments to the same observations; order must stay aligned.
groupby('cluster')- Groups the original measurements by their fitted assignment.
.mean()- Summarises each group using the original measurement units.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
labelled = X.assign(cluster=labels)
answer = labelled.groupby('cluster')[['length', 'width']].mean()
This practice: Read and run the Python. Next: Change · Interpret discovery without inventing truth.
Given data · CLASS180
180 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.
Reference labels are omitted from this preview and must remain outside fitting.
| length | width | label |
|---|---|---|
| 236.033 | -0.805568 | A |
| 581.297 | 0.728558 | A |
| -1511.27 | -1.00866 | A |
| 99.0248 | -0.24496 | A |
| -13.0141 | -0.660765 | A |
| 681.179 | 0.602475 | A |
| 51.1472 | 0.873157 | A |
| 362.131 | -0.665605 | A |
Column meanings and units
Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.
Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.
| Column | Stored type |
|---|---|
| length | float64 |
| width | float64 |
| label | str |
Supplied setup · available if you need to inspect it
This code runs before your editor on every Run. These are the objects your exercise uses.
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
X = df[['length', 'width']]
scaled = StandardScaler().fit_transform(X)
labels = KMeans(n_clusters=3, n_init=20, random_state=42).fit_predict(scaled)Your task · Follow
Join the supplied fitted group assignments back to the original measurements and store the group means in answer.
Hint 1 — Think
Use assign() and groupby() to interpret clusters in original units after fitting.
Hint 2 — Tools
Use X.assign(cluster=labels), groupby('cluster'), .mean(). Read the visible syntax meanings before editing.
Hint 3 — Approach
Join the supplied fitted group assignments back to the original measurements and store the group means in answer. Keep the supplied row order and inspect the named output after running.
Explained solution
labelled = X.assign(cluster=labels)
answer = labelled.groupby('cluster')[['length', 'width']].mean()
The fit used measurements alone. These means describe its groups, whose numeric IDs are arbitrary. External labels may be compared afterwards, but cannot turn discovered groups into proven natural classes.
Helpful prior knowledge: When K-Means geometry misleads · Sampled hierarchies describe sampled rows These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.