Understand the idea
Agglomerative clustering begins with individual observations and repeatedly merges groups. Ward chooses merges based on the increase in within-cluster variation. Dendrogram height represents its merge distance, not probability.
Python skill: Builds hierarchical merges that minimise increases in within-group squared distance.
Meet the syntax
linkage(scaled, method='ward')
dendrogram(linkage_matrix)linkage(scaled, method='ward')- Builds hierarchical merges that minimise increases in within-group squared distance.
dendrogram(linkage_matrix)- Draws the stored merge structure; leaf ordering is not an independent measurement axis.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
answer=linkage_matrix
This practice: Read and run the Python. Next: Change · Hierarchical merging.
Hierarchical clustering · from idea to workflow
Question: How do observations merge into groups at different resolutions?
Mechanism: Ward merges groups to limit increases in within-group squared variation.
Watch for: Scale, unusual points and sampling change merges; a cut does not discover a uniquely true class.
Interpret or debug · self-review: Read merge heights and compare two cuts. Explain how the sample limits the scope of group profiles.
Guided full workflow: Build a sampled hierarchy, compare cuts and describe groups in original units →
Independent full workflow: ML-X18. Use the Discovery route to frame a question without a prediction target.
Given data · CLUSTER36
36 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.
Reference labels are omitted from this preview and must remain outside fitting.
| length_mm | width_cm |
|---|---|
| 1106.65 | 6.36006 |
| 1262.66 | 13.292 |
| 317.138 | 5.44237 |
| 1044.74 | 8.89315 |
| 994.12 | 7.01435 |
| 1307.79 | 12.7223 |
| 1023.11 | 13.9453 |
| 1163.63 | 6.99248 |
Column meanings and units
Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.
Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.
| Column | Stored type |
|---|---|
| length_mm | float64 |
| width_cm | float64 |
Supplied setup · available if you need to inspect it
This code runs before your editor on every Run. These are the objects your exercise uses.
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
X=df.copy()
scaler=StandardScaler()
scaled=scaler.fit_transform(X)
from scipy.cluster.hierarchy import linkage,dendrogram,cut_tree
linkage_matrix=linkage(scaled,method='ward')
Your task · Follow
Store the supplied Ward linkage matrix for scaled CLUSTER36 in answer.
Hint 1 — Think
A linkage matrix records merges rather than one final assignment per row.
Hint 2 — Tools
The supplied Ward linkage_matrix.
Hint 3 — Approach
Inspect and return the prepared merge record for the scaled population.
Explained solution
answer=linkage_matrix
Each merge row represents the hierarchical construction; a later cut is needed to obtain a chosen grouping.
Helpful prior knowledge: Distance depends on scale and context These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.