Retrieve without the worked example: Store the smallest number of PCA components reaching 80% variance in retained_80 and the number reaching 95% in retained_95.
Use the new retrieval population shown here.
Retrieve earlier concepts before combining them.
This practice: Retrieve earlier concepts before combining them. Next: Retrieval 2 · PCA retrieval.
Given data · PCA48_REVIEW
48 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.
Reference labels are omitted from this preview and must remain outside fitting.
| a | b | c | d | e |
|---|---|---|---|---|
| 102.621 | 50.9668 | 22.5861 | 44.2048 | 0.872538 |
| 89.9035 | 43.797 | 20.1521 | 39.6665 | -0.76894 |
| 107.356 | 53.8675 | 21.059 | 41.7489 | 0.85829 |
| 109.521 | 54.1374 | 22.5529 | 44.9654 | 1.4285 |
| 80.6208 | 39.8638 | 14.1887 | 29.433 | -3.32762 |
| 86.2193 | 44.2712 | 18.703 | 37.7121 | -1.63585 |
| 101.474 | 50.6632 | 18.125 | 36.0109 | -0.348408 |
| 96.4609 | 49.0096 | 17.4031 | 33.8596 | -0.966253 |
Column meanings and units
Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.
Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.
| Column | Stored type |
|---|---|
| a | float64 |
| b | float64 |
| c | float64 |
| d | float64 |
| e | float64 |
Supplied setup · available if you need to inspect it
This code runs before your editor on every Run. These are the objects your exercise uses.
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
X=df.copy()
scaler=StandardScaler()
scaled=scaler.fit_transform(X)
pca=PCA().fit(scaled)
scores=pca.transform(scaled)
Supporting concepts: Two dimensions are a view →
Remember the idea
Use the inputs and evidence to recover the method. Hints and explained solutions remain collapsed; exact phrasing is not graded.
Hint 1 — Think
Recall how a threshold becomes the smallest sufficient component count.
Hint 2 — Tools
Cumulative variance and first-threshold lookup.
Hint 3 — Approach
Apply the prefix rule to each requested threshold using the same fitted ratios.
Explained solution
cumulative=np.cumsum(pca.explained_variance_ratio_)
retained_80=int(np.searchsorted(cumulative,.8)+1)
retained_95=int(np.searchsorted(cumulative,.95)+1)
The counts compare retention requirements without refitting or confusing a component position with the number retained.
Helpful prior knowledge: Two dimensions are a view These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.