Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Clustering and Discovery lessonsQUESTIONS · MODELS · EVIDENCE
K-Means · ML-U06 · 12–18 MIN

When K-Means geometry misleads

Recognise initialisation, scale and shape limitations.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeExplain the Python result
  3. TransferExplain the Python result

Understand the idea

Different starting centres can find different local solutions. Multiple starts help, but do not remove the preference for compact centroid-based groups or resolve the meaning of k.

Recognise initialisation, scale and shape limitations.×××Distance, neighbourhoods and centres depend on scale.
Schematic · Recognise initialisation, scale and shape limitations.Scroll the diagram horizontally if needed.

Python skill: Uses a single initialisation so seed sensitivity is visible in this exercise.

Meet the syntax

KMeans(n_clusters=4, n_init=1, random_state=seed)
model.inertia_
n_init=1
Uses a single initialisation so seed sensitivity is visible in this exercise.
random_state=seed
Changes the starting random choices for each repeated fit.
model.inertia_
Records each fit’s compactness so instability can be compared.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

answer=[]
for seed in [1,2,3]:
    model=KMeans(n_clusters=4,n_init=1,random_state=seed).fit(scaled)
    answer.append(model.inertia_)

This practice: Read and run the Python. Next: Change · When K-Means geometry misleads.

Given data · CLUSTER36

36 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

Reference labels are omitted from this preview and must remain outside fitting.

CLUSTER36 · first 8 prepared rows
length_mmwidth_cm
1106.656.36006
1262.6613.292
317.1385.44237
1044.748.89315
994.127.01435
1307.7912.7223
1023.1113.9453
1163.636.99248

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
length_mmfloat64
width_cmfloat64
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
X=df.copy()
scaler=StandardScaler()
scaled=scaler.fit_transform(X)

Your task · Follow

Fit four-cluster K-Means with one initialisation for each seed 1, 2 and 3. Store each inertia in answer in seed order.

Hint 1 — Think

Different initial centres can lead to different local solutions.

Hint 2 — Tools

KMeans with n_init=1, random_state and inertia_.

Hint 3 — Approach

Keep k and data fixed while fitting each declared seed, then compare the objectives.

Explained solution
answer=[]
for seed in [1,2,3]:
    model=KMeans(n_clusters=4,n_init=1,random_state=seed).fit(scaled)
    answer.append(model.inertia_)

One start exposes initialisation sensitivity that a multi-start run is designed to reduce.

Helpful prior knowledge: Describe clusters in meaningful units These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Fit four-cluster K-Means with one initialisation for each seed 1, 2 and 3. Store each inertia in answer in seed order.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.