Scale the transfer population, compare k evidence, profile groups and qualify their interpretation.
This practice: Scale the transfer population, compare k evidence, profile groups and qualify their interpretation. Next: PCA. Continue to PCA →
Given data · CLUSTER36B
36 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.
Reference labels are omitted from this preview and must remain outside fitting.
| length_mm | width_cm |
|---|---|
| 1085.48 | 12.3736 |
| 795.065 | 6.81964 |
| 302.857 | 13.4007 |
| 1005.83 | 10.7201 |
| 725.742 | 14.2927 |
| 1330.12 | 9.57362 |
| 805.448 | 8.75158 |
| 720.481 | 10.093 |
Column meanings and units
Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.
Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.
| Column | Stored type |
|---|---|
| length_mm | float64 |
| width_cm | float64 |
Supporting concepts: Describe clusters in meaningful units → · Interpret discovery without inventing truth →
Remember the idea
This checkpoint combines previously taught skills. Assemble the workflow; help remains available when needed.
Required Python variables and evidence
Use these names so Check can inspect your workflow. Each meaning is shown beside its name.
| Variable | Meaning |
|---|---|
| X | Feature dataframe for the declared population, preserving row indices. |
| df | Loaded and prepared input dataframe. |
| evidence | Dataframe with k, inertia and silhouette columns for k=2–8. |
| labels | One cluster label per fitted row; equivalent consistent label names are accepted. |
| model | Fitted estimator or pipeline for the requested model family. |
| profiles | Original-unit means grouped by aligned cluster labels. |
| scaled | Standardised features in the same row/column order as X. |
| sizes | Group sizes indexed by cluster label. |
Hint 1 — Think
Reconstruct discovery without importing a supervised target or a unique correct grouping.
Hint 2 — Tools
Scaling, KMeans k comparison, silhouette, sizes and original-unit profiles.
Hint 3 — Approach
Define the population, compare resolutions, choose a defensible grouping and report its profiles with a qualified visual.
Explained solution
from sklearn.preprocessing import StandardScaler
from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score
import matplotlib.pyplot as plt
X=df.copy()
scaler=StandardScaler()
scaled=scaler.fit_transform(X)
rows=[]
for k in range(2,min(8,len(X)-1)+1):
candidate=KMeans(n_clusters=k,n_init=20,random_state=42).fit(scaled)
score=silhouette_score(scaled,candidate.labels_,sample_size=min(2000,len(X)),random_state=42)
rows.append({'k':k,'inertia':candidate.inertia_,'silhouette':score})
evidence=pd.DataFrame(rows)
# Three is an illustrative choice to discuss against the evidence, not a unique optimum.
model=KMeans(n_clusters=3,n_init=20,random_state=42).fit(scaled)
labels=model.labels_
sizes=pd.Series(labels).value_counts().sort_index()
profiles=X.groupby(labels).mean()
fig,ax=plt.subplots(figsize=(6,4))
for label in np.unique(labels):
group=X.loc[labels==label]
ax.scatter(group.length_mm,group.width_cm,label='Group '+str(label),marker=['o','s','^'][int(label)%3])
ax.set(xlabel='Length (mm)',ylabel='Width (cm)',title='Exploratory measurement groups')
ax.legend()
fig.savefig('clusters.png',dpi=150,bbox_inches='tight')
print(evidence)
The workflow joins geometric evidence with substantive interpretation; neither the selected k nor the group IDs claim a hidden class truth.
Helpful prior knowledge: K-Means retrieval · Hierarchy retrieval These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.