Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← PCA lessonsQUESTIONS · MODELS · EVIDENCE
Reading a representation · ML-P-R1 · 15–20 MIN

PCA retrieval

Retrieval 1

Exercises within this concept

  1. Retrieval 1Retrieve and applyCurrent exercise
  2. Retrieval 2Retrieve and apply
  3. Retrieval 3Retrieve and apply
Retrieval 1 · PCA48_REVIEW

Retrieve without the worked example: Store the smallest number of PCA components reaching 80% variance in retained_80 and the number reaching 95% in retained_95.

Use the new retrieval population shown here.

Retrieve earlier concepts before combining them.

This practice: Retrieve earlier concepts before combining them. Next: Retrieval 2 · PCA retrieval.

Given data · PCA48_REVIEW

48 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

Reference labels are omitted from this preview and must remain outside fitting.

PCA48_REVIEW · first 8 prepared rows
abcde
102.62150.966822.586144.20480.872538
89.903543.79720.152139.6665-0.76894
107.35653.867521.05941.74890.85829
109.52154.137422.552944.96541.4285
80.620839.863814.188729.433-3.32762
86.219344.271218.70337.7121-1.63585
101.47450.663218.12536.0109-0.348408
96.460949.009617.403133.8596-0.966253

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
afloat64
bfloat64
cfloat64
dfloat64
efloat64
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
X=df.copy()
scaler=StandardScaler()
scaled=scaler.fit_transform(X)
pca=PCA().fit(scaled)
scores=pca.transform(scaled)

Supporting concepts: Two dimensions are a view →

Remember the idea

Use the inputs and evidence to recover the method. Hints and explained solutions remain collapsed; exact phrasing is not graded.

Retrieve earlier concepts before combining them.Candidate ACandidate BFitValidateCostCompare matching evidence; smaller error can cost more.
Schematic · Retrieve earlier concepts before combining them.Scroll the diagram horizontally if needed.
Hint 1 — Think

Recall how a threshold becomes the smallest sufficient component count.

Hint 2 — Tools

Cumulative variance and first-threshold lookup.

Hint 3 — Approach

Apply the prefix rule to each requested threshold using the same fitted ratios.

Explained solution
cumulative=np.cumsum(pca.explained_variance_ratio_)
retained_80=int(np.searchsorted(cumulative,.8)+1)
retained_95=int(np.searchsorted(cumulative,.95)+1)

The counts compare retention requirements without refitting or confusing a component position with the number retained.

Helpful prior knowledge: Two dimensions are a view These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Retrieval 1

Retrieve without the worked example: Store the smallest number of PCA components reaching 80% variance in retained_80 and the number reaching 95% in retained_95. Use the new retrieval population shown here.

retained_80retained_95

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.

What changes when you retain 95% rather than 80% of the measured variation?

Use your Run output as evidence. This response is optional, not machine-graded or saved.