Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Clustering and Discovery lessonsQUESTIONS · MODELS · EVIDENCE
Hierarchical discovery · ML-U09 · 12–18 MIN

Sampled hierarchies describe sampled rows

Keep population, sample and labels aligned.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeAdapt the Python
  3. TransferExplain the Python result

Understand the idea

Production fits the scale on the exploratory population, then transforms a reproducible sample of at most 500 rows for the hierarchy. The resulting labels belong to those sampled rows.

Keep population, sample and labels aligned.Population: N rowsSample: n rowsn aligned labelsPreserve sampled row identities through fitting and profiles.The hierarchy describes the sampled observations.
Schematic · Keep population, sample and labels aligned.Scroll the diagram horizontally if needed.

Python skill: Selects rows while preserving their original index labels.

Meet the syntax

sample = X.sample(min(500,len(X)), random_state=42)
X.sample
Selects rows while preserving their original index labels.
min(500,len(X))
Caps the sample size at the smaller of 500 and the available row count.
random_state=42
Makes the selected sample reproducible.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

answer=X.sample(100,random_state=42)

This practice: Read and run the Python. Next: Change · Sampled hierarchies describe sampled rows.

Given data · penguins

333 observations. One measured penguin. The dataframe df is supplied afresh for each Run.

Download source CSV · Source and original dictionary

Reference labels are omitted from this preview and must remain outside fitting.

penguins · first 8 prepared rows
islandbill_length_mmbill_depth_mmflipper_length_mmbody_mass_gsexyear
Torgersen39.118.71813750male2007
Torgersen39.517.41863800female2007
Torgersen40.3181953250female2007
Torgersen36.719.31933450female2007
Torgersen39.320.61903650male2007
Torgersen38.917.81813625female2007
Torgersen39.219.61954675male2007
Torgersen41.117.61823200female2007

Column meanings and units

Bill length/depth and flipper length: mm. Body mass: grams. Year and island: sampling context.

ML uses 333 complete cases. Removing incomplete records may change the represented population. Geographic context may not generalise to new islands.

Input schema
ColumnStored type
islandstr
bill_length_mmfloat64
bill_depth_mmfloat64
flipper_length_mmint64
body_mass_gint64
sexstr
yearint64
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
X=df[['bill_length_mm','bill_depth_mm','flipper_length_mm','body_mass_g']]
scaled=StandardScaler().fit_transform(X)
labels=KMeans(n_clusters=3,n_init=20,random_state=42).fit_predict(scaled)

Your task · Follow

Sample 100 Penguin measurement rows while retaining their original index.

Hint 1 — Think

Original indices preserve the link between sampled observations and their measurements.

Hint 2 — Tools

DataFrame.sample with random_state.

Hint 3 — Approach

Take the requested reproducible sample without resetting its index.

Explained solution
answer=X.sample(100,random_state=42)

Retaining row identity lets later assignments and original-unit profiles remain aligned to the sampled population.

Helpful prior knowledge: Turn a hierarchy into groups · Describe clusters in meaningful units These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Sample 100 Penguin measurement rows while retaining their original index.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.