Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Clustering and Discovery lessonsQUESTIONS · MODELS · EVIDENCE
A different question · ML-U02 · 12–18 MIN

Distance depends on scale and context

Fit a meaningful distance representation for the exploratory population.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeAdapt the Python
  3. TransferReason about the Python

Understand the idea

In discovery, fitting a scale on the declared population is part of describing that population. If PCA or clustering features later enter a supervised workflow, fit those transformations inside its training folds.

Fit a meaningful distance representation for the exploratory population.Raw: unequal unitsStandardisedFit statistics at the boundary appropriate to the question.
Schematic · Fit a meaningful distance representation for the exploratory population.Scroll the diagram horizontally if needed.

Python skill: Learns a scale from the declared discovery population and transforms those same rows.

Meet the syntax

scaled = StandardScaler().fit_transform(X)
StandardScaler().fit_transform(X)
Learns a scale from the declared discovery population and transforms those same rows.
scaled
Contains comparable-unit coordinates for distance calculations; retain X for original-unit descriptions.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

from sklearn.preprocessing import StandardScaler
scaler=StandardScaler()
answer=scaler.fit_transform(df)

This practice: Read and run the Python. Next: Change · Distance depends on scale and context.

Given data · CLUSTER36

36 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

Reference labels are omitted from this preview and must remain outside fitting.

CLUSTER36 · first 8 prepared rows
length_mmwidth_cm
1106.656.36006
1262.6613.292
317.1385.44237
1044.748.89315
994.127.01435
1307.7912.7223
1023.1113.9453
1163.636.99248

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
length_mmfloat64
width_cmfloat64

Your task · Follow

Standardise CLUSTER36’s unequal-unit measurements.

Hint 1 — Think

A large numerical unit can dominate distances without being substantively more important.

Hint 2 — Tools

StandardScaler.fit_transform.

Hint 3 — Approach

Learn a standardised representation of the declared discovery population.

Explained solution
from sklearn.preprocessing import StandardScaler
scaler=StandardScaler()
answer=scaler.fit_transform(df)

For this descriptive task the supplied population defines the scale; later supervised use would require a different fitting boundary.

Helpful prior knowledge: Discovery has no prediction target · Learn a scale from training rows These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Standardise CLUSTER36’s unequal-unit measurements.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.