Understand the idea
Different starting centres can find different local solutions. Multiple starts help, but do not remove the preference for compact centroid-based groups or resolve the meaning of k.
Python skill: Uses a single initialisation so seed sensitivity is visible in this exercise.
Meet the syntax
KMeans(n_clusters=4, n_init=1, random_state=seed)
model.inertia_n_init=1- Uses a single initialisation so seed sensitivity is visible in this exercise.
random_state=seed- Changes the starting random choices for each repeated fit.
model.inertia_- Records each fit’s compactness so instability can be compared.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
answer=[]
for seed in [1,2,3]:
model=KMeans(n_clusters=4,n_init=1,random_state=seed).fit(scaled)
answer.append(model.inertia_)
This practice: Read and run the Python. Next: Change · When K-Means geometry misleads.
Given data · CLUSTER36
36 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.
Reference labels are omitted from this preview and must remain outside fitting.
| length_mm | width_cm |
|---|---|
| 1106.65 | 6.36006 |
| 1262.66 | 13.292 |
| 317.138 | 5.44237 |
| 1044.74 | 8.89315 |
| 994.12 | 7.01435 |
| 1307.79 | 12.7223 |
| 1023.11 | 13.9453 |
| 1163.63 | 6.99248 |
Column meanings and units
Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.
Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.
| Column | Stored type |
|---|---|
| length_mm | float64 |
| width_cm | float64 |
Supplied setup · available if you need to inspect it
This code runs before your editor on every Run. These are the objects your exercise uses.
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
X=df.copy()
scaler=StandardScaler()
scaled=scaler.fit_transform(X)
Your task · Follow
Fit four-cluster K-Means with one initialisation for each seed 1, 2 and 3. Store each inertia in answer in seed order.
Hint 1 — Think
Different initial centres can lead to different local solutions.
Hint 2 — Tools
KMeans with n_init=1, random_state and inertia_.
Hint 3 — Approach
Keep k and data fixed while fitting each declared seed, then compare the objectives.
Explained solution
answer=[]
for seed in [1,2,3]:
model=KMeans(n_clusters=4,n_init=1,random_state=seed).fit(scaled)
answer.append(model.inertia_)
One start exposes initialisation sensitivity that a multi-start run is designed to reduce.
Helpful prior knowledge: Describe clusters in meaningful units These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.