Understand the idea
Production uses the smallest prefix reaching at least 90% cumulative variance. The threshold is a stated compression criterion, not a universal optimum.
Python skill: Finds the first zero-based position reaching the variance threshold.
Meet the syntax
retained = np.searchsorted(cumulative, 0.9) + 1
np.cumsum(ratios)np.searchsorted(cumulative, 0.9)- Finds the first zero-based position reaching the variance threshold.
+ 1- Converts that position into a component count for the retained prefix.
np.cumsum(ratios)- Accumulates explained variance in component order before applying the retention threshold.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
cumulative=np.cumsum(pca.explained_variance_ratio_)
retained=int(np.searchsorted(cumulative,.9)+1)
answer=scores[:,:retained]
This practice: Read and run the Python. Next: Change · Select a retained representation.
Given data · PCA48
48 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.
Reference labels are omitted from this preview and must remain outside fitting.
| a | b | c | d | e |
|---|---|---|---|---|
| 103.047 | 50.8622 | 22.7157 | 44.6184 | 0.906993 |
| 89.6002 | 44.3015 | 20.2703 | 40.1253 | -0.680116 |
| 107.505 | 53.9521 | 21.1565 | 41.7009 | 0.818161 |
| 109.406 | 54.2501 | 22.5252 | 44.9095 | 1.39291 |
| 80.4896 | 40.0557 | 14.1714 | 29.4087 | -3.27953 |
| 86.9782 | 44.1387 | 18.7213 | 37.5997 | -1.70077 |
| 101.278 | 50.4611 | 18.1185 | 36.0784 | -0.343557 |
| 96.8376 | 48.7875 | 17.4445 | 33.8533 | -0.987809 |
Column meanings and units
Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.
Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.
| Column | Stored type |
|---|---|
| a | float64 |
| b | float64 |
| c | float64 |
| d | float64 |
| e | float64 |
Supplied setup · available if you need to inspect it
This code runs before your editor on every Run. These are the objects your exercise uses.
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
X=df.copy()
scaler=StandardScaler()
scaled=scaler.fit_transform(X)
pca=PCA().fit(scaled)
scores=pca.transform(scaled)
Your task · Follow
Keep the fewest leading PCA components that explain at least 90% variance. Store that count in retained and the matching row scores in answer.
Hint 1 — Think
The rule asks for the first cumulative total that reaches the threshold.
Hint 2 — Tools
np.cumsum, np.searchsorted and score-column slicing.
Hint 3 — Approach
Accumulate variance ratios, locate the first qualifying component count and retain that prefix of scores.
Explained solution
cumulative=np.cumsum(pca.explained_variance_ratio_)
retained=int(np.searchsorted(cumulative,.9)+1)
answer=scores[:,:retained]
Adding one converts a zero-based position into a dimension count; the prefix retains the required variation with the fewest leading axes.
Helpful prior knowledge: Explained variance is not predictive accuracy These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.