Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Clustering and Discovery lessonsQUESTIONS · MODELS · EVIDENCE
A different question · ML-U01 · 12–18 MIN

Discovery has no prediction target

State the exploratory population and question.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeExplain the Python result
  3. TransferExplain the Python result

Understand the idea

In supervised learning we had y. Here the goal is to describe structure in X. Cluster IDs are arbitrary identifiers, not predicted real-world classes. Reference labels must not guide fitting or selection.

State the exploratory population and question.QuestionPredict quantityPredict labelDiscover groupsReduce dimensions
Schematic · State the exploratory population and question.Scroll the diagram horizontally if needed.

Python skill: Contains only the measurements chosen to define similarity.

Meet the syntax

X = df[measurement_columns]
measurement_columns
Contains only the measurements chosen to define similarity.
df[measurement_columns]
Creates X without a supervised target or hidden reference label.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

answer=df[['bill_length_mm','bill_depth_mm','flipper_length_mm','body_mass_g']]

This practice: Read and run the Python. Next: Change · Discovery has no prediction target.

Given data · penguins

333 observations. One measured penguin. The dataframe df is supplied afresh for each Run.

Download source CSV · Source and original dictionary

Reference labels are omitted from this preview and must remain outside fitting.

penguins · first 8 prepared rows
islandbill_length_mmbill_depth_mmflipper_length_mmbody_mass_gsexyear
Torgersen39.118.71813750male2007
Torgersen39.517.41863800female2007
Torgersen40.3181953250female2007
Torgersen36.719.31933450female2007
Torgersen39.320.61903650male2007
Torgersen38.917.81813625female2007
Torgersen39.219.61954675male2007
Torgersen41.117.61823200female2007

Column meanings and units

Bill length/depth and flipper length: mm. Body mass: grams. Year and island: sampling context.

ML uses 333 complete cases. Removing incomplete records may change the represented population. Geographic context may not generalise to new islands.

Input schema
ColumnStored type
islandstr
bill_length_mmfloat64
bill_depth_mmfloat64
flipper_length_mmint64
body_mass_gint64
sexstr
yearint64

Your task · Follow

Store only bill length, bill depth, flipper length and body mass in answer, keeping every Penguin row and the stated column order.

Hint 1 — Think

Reference labels and context fields would change the meaning of measurement-based discovery.

Hint 2 — Tools

Explicit measurement-column selection.

Hint 3 — Approach

Keep only the four declared numeric measurements and inspect the resulting schema.

Explained solution
answer=df[['bill_length_mm','bill_depth_mm','flipper_length_mm','body_mass_g']]

Excluding species and context keeps grouping based on the intended measured similarities rather than known labels.

Helpful prior knowledge: ML Foundations checkpoint These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Store only bill length, bill depth, flipper length and body mass in answer, keeping every Penguin row and the stated column order.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.