Fit Penguin PCA, retain at least 90% variance, label weights and report the separate 2D view.
This practice: Fit Penguin PCA, retain at least 90% variance, label weights and report the separate 2D view. Next: Choose and Explain Models. Continue to Choose and Explain Models →
Given data · penguins
333 observations. One measured penguin. The dataframe df is supplied afresh for each Run.
Download source CSV · Source and original dictionary
Reference labels are omitted from this preview and must remain outside fitting.
| island | bill_length_mm | bill_depth_mm | flipper_length_mm | body_mass_g | sex | year |
|---|---|---|---|---|---|---|
| Torgersen | 39.1 | 18.7 | 181 | 3750 | male | 2007 |
| Torgersen | 39.5 | 17.4 | 186 | 3800 | female | 2007 |
| Torgersen | 40.3 | 18 | 195 | 3250 | female | 2007 |
| Torgersen | 36.7 | 19.3 | 193 | 3450 | female | 2007 |
| Torgersen | 39.3 | 20.6 | 190 | 3650 | male | 2007 |
| Torgersen | 38.9 | 17.8 | 181 | 3625 | female | 2007 |
| Torgersen | 39.2 | 19.6 | 195 | 4675 | male | 2007 |
| Torgersen | 41.1 | 17.6 | 182 | 3200 | female | 2007 |
Column meanings and units
Bill length/depth and flipper length: mm. Body mass: grams. Year and island: sampling context.
ML uses 333 complete cases. Removing incomplete records may change the represented population. Geographic context may not generalise to new islands.
| Column | Stored type |
|---|---|
| island | str |
| bill_length_mm | float64 |
| bill_depth_mm | float64 |
| flipper_length_mm | int64 |
| body_mass_g | int64 |
| sex | str |
| year | int64 |
Supporting concepts: Two dimensions are a view →
Remember the idea
This checkpoint combines previously taught skills. Assemble the workflow; help remains available when needed.
Required Python variables and evidence
Use these names so Check can inspect your workflow. Each meaning is shown beside its name.
| Variable | Meaning |
|---|---|
| X | Feature dataframe for the declared population, preserving row indices. |
| cumulative | Cumulative sum of ratios. |
| pca | PCA fitted on scaled measurements. |
| ratios | Explained variance ratios in descending component order. |
| reduced | Scores for the retained component prefix. |
| retained | Smallest number of components reaching at least 90% variance. |
| scaled | Standardised features in the same row/column order as X. |
| scores2 | Row-by-two component scores, with signs consistent with weights. |
| weights | Feature-by-two dataframe of PC1/PC2 axis weights, labelled by feature and component. |
Hint 1 — Think
Reconstruct the distinction between retained representation and a convenient picture.
Hint 2 — Tools
Measurement selection, scaling, PCA, cumulative variance and labelled weights/scores.
Hint 3 — Approach
Fit the measurement representation, find the smallest 90% prefix and report it separately from the first-two-axis view.
Explained solution
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
import matplotlib.pyplot as plt
X = df[['bill_length_mm','bill_depth_mm','flipper_length_mm','body_mass_g']]
scaler = StandardScaler()
scaled = scaler.fit_transform(X)
pca = PCA().fit(scaled)
scores = pca.transform(scaled)
ratios = pca.explained_variance_ratio_
cumulative = np.cumsum(ratios)
retained = int(np.searchsorted(cumulative,0.9)+1)
reduced = scores[:,:retained]
weights = pd.DataFrame(pca.components_[:2].T.copy(),index=X.columns,columns=['PC1','PC2'])
scores2 = scores[:,:2].copy()
fig, ax = plt.subplots(figsize=(6,4))
ax.scatter(scores2[:,0],scores2[:,1],s=12)
ax.set(title='Two-component view of measurements',xlabel='PC1 score',ylabel='PC2 score')
print('Retained dimensions:',retained,'; variance visible in 2D:',ratios[:2].sum())
The retained matrix fulfils the variance rule while the labelled two-dimensional display communicates only the variation its axes contain.
Helpful prior knowledge: PCA retrieval These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.