Understand the idea
K-Means alternates assignments and centroid updates to reduce within-cluster squared distances. Production uses multiple initialisations and a reproducible seed. k=3 is a starting choice, not a discovered truth.
Python skill: Requests three fitted centroids.
Meet the syntax
KMeans(n_clusters=3, n_init=20, random_state=42)
model.cluster_centers_
model.labels_n_clusters=3- Requests three fitted centroids.
n_init=20- Tries 20 initialisations and retains the lowest-inertia fit.
random_state=42- Makes those initialisation choices repeatable.
model.cluster_centers_- Stores fitted centroids in the scaled coordinate system.
model.labels_- Stores one cluster assignment per fitted observation, in input order.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
model=KMeans(n_clusters=3,n_init=20,random_state=42).fit(scaled)
answer=model.cluster_centers_
This practice: Read and run the Python. Next: Change · K-Means assigns points to centroids.
K-Means · from idea to workflow
Question: What compact groups describe these observations?
Mechanism: Alternate nearest-centre assignments and mean-centre updates in scaled space.
Watch for: Round distance-based groups may not reflect useful categories; k is a resolution choice.
Interpret or debug · self-review: Compare silhouette, sizes and original-unit profiles. Explain why reference labels must stay outside fitting and why no final classification accuracy is claimed.
Independent full workflow: ML-X17. Use the Discovery route to frame a question without a prediction target.
Given data · CLUSTER36
36 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.
Reference labels are omitted from this preview and must remain outside fitting.
| length_mm | width_cm |
|---|---|
| 1106.65 | 6.36006 |
| 1262.66 | 13.292 |
| 317.138 | 5.44237 |
| 1044.74 | 8.89315 |
| 994.12 | 7.01435 |
| 1307.79 | 12.7223 |
| 1023.11 | 13.9453 |
| 1163.63 | 6.99248 |
Column meanings and units
Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.
Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.
| Column | Stored type |
|---|---|
| length_mm | float64 |
| width_cm | float64 |
Supplied setup · available if you need to inspect it
This code runs before your editor on every Run. These are the objects your exercise uses.
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
X=df.copy()
scaler=StandardScaler()
scaled=scaler.fit_transform(X)
Your task · Follow
Fit three K-Means clusters to scaled CLUSTER36. Store the fitted centroid coordinates in answer.
Hint 1 — Think
Centroids live in the same coordinate system used for fitting.
Hint 2 — Tools
KMeans, n_clusters, n_init and cluster_centers_.
Hint 3 — Approach
Fit the declared reproducible three-centroid model to scaled values and inspect its centres.
Explained solution
model=KMeans(n_clusters=3,n_init=20,random_state=42).fit(scaled)
answer=model.cluster_centers_
Multiple starts reduce dependence on a single initialisation; the resulting centres are in standardised units.
Helpful prior knowledge: Distance depends on scale and context These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.