Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← PCA lessonsQUESTIONS · MODELS · EVIDENCE
How much to retain · ML-P03 · 12–18 MIN

Explained variance is not predictive accuracy

Interpret per-component variation.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeAdapt the Python
  3. TransferReason about the Python

Understand the idea

Explained variance ratios measure variation along fitted axes relative to total variation in the prepared inputs. They say nothing directly about predicting a target.

Interpret per-component variation.PC1PC2Variance by componentNew axes combine measurements.
Schematic · Interpret per-component variation.Scroll the diagram horizontally if needed.

Python skill: Reads each component’s share of input variance in fitted order; these values are not prediction scores.

Meet the syntax

pca.explained_variance_ratio_
pca.explained_variance_ratio_
Reads each component’s share of input variance in fitted order; these values are not prediction scores.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

answer=pca.explained_variance_ratio_

This practice: Read and run the Python. Next: Change · Explained variance is not predictive accuracy.

Given data · PCA48

48 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

Reference labels are omitted from this preview and must remain outside fitting.

PCA48 · first 8 prepared rows
abcde
103.04750.862222.715744.61840.906993
89.600244.301520.270340.1253-0.680116
107.50553.952121.156541.70090.818161
109.40654.250122.525244.90951.39291
80.489640.055714.171429.4087-3.27953
86.978244.138718.721337.5997-1.70077
101.27850.461118.118536.0784-0.343557
96.837648.787517.444533.8533-0.987809

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
afloat64
bfloat64
cfloat64
dfloat64
efloat64
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
X=df.copy()
scaler=StandardScaler()
scaled=scaler.fit_transform(X)
pca=PCA().fit(scaled)
scores=pca.transform(scaled)

Your task · Follow

Inspect the explained variance ratios.

Hint 1 — Think

Variance ratios refer to the representation on which PCA was fitted.

Hint 2 — Tools

explained_variance_ratio_.

Hint 3 — Approach

Read the fitted per-component ratio array in component order.

Explained solution
answer=pca.explained_variance_ratio_

The ratios describe how input variation is distributed across axes; they are not prediction scores.

Helpful prior knowledge: Fit a reusable PCA representation These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Inspect the explained variance ratios.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.