Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Clustering and Discovery lessonsQUESTIONS · MODELS · EVIDENCE
K-Means · ML-U05 · 12–18 MIN

Describe clusters in meaningful units

Use sizes and original-unit profiles with named features.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeAdapt the Python
  3. TransferExplain the Python result

Understand the idea

Cluster labels may be permuted without changing the grouping. Report original-unit feature summaries and population sizes so the result is interpretable.

Use sizes and original-unit profiles with named features.GroupSizeMean length (mm)0n₀μ₀1n₁μ₁2n₂μ₂Aggregate original measurements with aligned assignments.
Schematic · Use sizes and original-unit profiles with named features.Scroll the diagram horizontally if needed.

Python skill: Groups original-unit rows by their aligned fitted assignments.

Meet the syntax

X.groupby(labels).mean()
pd.Series(labels).value_counts()
X.groupby(labels)
Groups original-unit rows by their aligned fitted assignments.
.mean()
Computes each group’s mean per measurement; it does not change or refit the groups.
pd.Series(labels).value_counts()
Counts how many fitted observations belong to each named group.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

answer=X.groupby(labels).mean()
sizes=pd.Series(labels).value_counts().sort_index()

This practice: Read and run the Python. Next: Change · Describe clusters in meaningful units.

Given data · penguins

333 observations. One measured penguin. The dataframe df is supplied afresh for each Run.

Download source CSV · Source and original dictionary

Reference labels are omitted from this preview and must remain outside fitting.

penguins · first 8 prepared rows
islandbill_length_mmbill_depth_mmflipper_length_mmbody_mass_gsexyear
Torgersen39.118.71813750male2007
Torgersen39.517.41863800female2007
Torgersen40.3181953250female2007
Torgersen36.719.31933450female2007
Torgersen39.320.61903650male2007
Torgersen38.917.81813625female2007
Torgersen39.219.61954675male2007
Torgersen41.117.61823200female2007

Column meanings and units

Bill length/depth and flipper length: mm. Body mass: grams. Year and island: sampling context.

ML uses 333 complete cases. Removing incomplete records may change the represented population. Geographic context may not generalise to new islands.

Input schema
ColumnStored type
islandstr
bill_length_mmfloat64
bill_depth_mmfloat64
flipper_length_mmint64
body_mass_gint64
sexstr
yearint64
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
X=df[['bill_length_mm','bill_depth_mm','flipper_length_mm','body_mass_g']]
scaled=StandardScaler().fit_transform(X)
labels=KMeans(n_clusters=3,n_init=20,random_state=42).fit_predict(scaled)

Your task · Follow

Report cluster sizes and original-unit means.

Hint 1 — Think

Assignments must align with the original rows whose profiles you report.

Hint 2 — Tools

groupby.mean and value_counts.

Hint 3 — Approach

Use the supplied labels to aggregate original measurements and count group membership.

Explained solution
answer=X.groupby(labels).mean()
sizes=pd.Series(labels).value_counts().sort_index()

Original-unit means explain group characteristics while sizes reveal whether a profile describes many rows or a tiny subset.

Helpful prior knowledge: Choosing k requires evidence and judgement These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Report cluster sizes and original-unit means.

answersizes

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.