Understand the idea
In discovery, fitting a scale on the declared population is part of describing that population. If PCA or clustering features later enter a supervised workflow, fit those transformations inside its training folds.
Python skill: Learns a scale from the declared discovery population and transforms those same rows.
Meet the syntax
scaled = StandardScaler().fit_transform(X)StandardScaler().fit_transform(X)- Learns a scale from the declared discovery population and transforms those same rows.
scaled- Contains comparable-unit coordinates for distance calculations; retain X for original-unit descriptions.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
from sklearn.preprocessing import StandardScaler
scaler=StandardScaler()
answer=scaler.fit_transform(df)
This practice: Read and run the Python. Next: Change · Distance depends on scale and context.
Given data · CLUSTER36
36 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.
Reference labels are omitted from this preview and must remain outside fitting.
| length_mm | width_cm |
|---|---|
| 1106.65 | 6.36006 |
| 1262.66 | 13.292 |
| 317.138 | 5.44237 |
| 1044.74 | 8.89315 |
| 994.12 | 7.01435 |
| 1307.79 | 12.7223 |
| 1023.11 | 13.9453 |
| 1163.63 | 6.99248 |
Column meanings and units
Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.
Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.
| Column | Stored type |
|---|---|
| length_mm | float64 |
| width_cm | float64 |
Your task · Follow
Standardise CLUSTER36’s unequal-unit measurements.
Hint 1 — Think
A large numerical unit can dominate distances without being substantively more important.
Hint 2 — Tools
StandardScaler.fit_transform.
Hint 3 — Approach
Learn a standardised representation of the declared discovery population.
Explained solution
from sklearn.preprocessing import StandardScaler
scaler=StandardScaler()
answer=scaler.fit_transform(df)
For this descriptive task the supplied population defines the scale; later supervised use would require a different fitting boundary.
Helpful prior knowledge: Discovery has no prediction target · Learn a scale from training rows These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.