Understand the idea
Standardisation changes which variation PCA emphasises. Fit the representation on the intended population; transform new rows using those same fitted statistics and axes.
Python skill: Learns component axes from the supplied scaled population.
Meet the syntax
pca = PCA().fit(scaled)
scores = pca.transform(scaled)PCA().fit(scaled)- Learns component axes from the supplied scaled population.
pca.transform(scaled)- Projects rows onto the already fitted axes without learning new axes.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
answer=scores
This practice: Read and run the Python. Next: Change · Fit a reusable PCA representation.
Principal component analysis · from idea to workflow
Question: Can fewer axes summarise variation in the measurements?
Mechanism: Orthogonal combinations successively retain as much variance as possible.
Watch for: High variance need not be useful for prediction; unscaled units can dominate.
Interpret or debug · self-review: Read weights and retained variance together. Explain why a component is a combination rather than an original feature, and why PCA is neither a target predictor nor a cluster label.
Guided full workflow: Scale measurements, fit PCA, retain dimensions and interpret component weights →
Independent full workflow: ML-X19. Use the Discovery route to frame a question without a prediction target.
Given data · PCA48
48 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.
Reference labels are omitted from this preview and must remain outside fitting.
| a | b | c | d | e |
|---|---|---|---|---|
| 103.047 | 50.8622 | 22.7157 | 44.6184 | 0.906993 |
| 89.6002 | 44.3015 | 20.2703 | 40.1253 | -0.680116 |
| 107.505 | 53.9521 | 21.1565 | 41.7009 | 0.818161 |
| 109.406 | 54.2501 | 22.5252 | 44.9095 | 1.39291 |
| 80.4896 | 40.0557 | 14.1714 | 29.4087 | -3.27953 |
| 86.9782 | 44.1387 | 18.7213 | 37.5997 | -1.70077 |
| 101.278 | 50.4611 | 18.1185 | 36.0784 | -0.343557 |
| 96.8376 | 48.7875 | 17.4445 | 33.8533 | -0.987809 |
Column meanings and units
Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.
Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.
| Column | Stored type |
|---|---|
| a | float64 |
| b | float64 |
| c | float64 |
| d | float64 |
| e | float64 |
Supplied setup · available if you need to inspect it
This code runs before your editor on every Run. These are the objects your exercise uses.
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
X=df.copy()
scaler=StandardScaler()
scaled=scaler.fit_transform(X)
pca=PCA().fit(scaled)
scores=pca.transform(scaled)
Your task · Follow
Store the supplied PCA component scores for every PCA48 row in answer.
Hint 1 — Think
Scores are each observation’s coordinates along the fitted component axes.
Hint 2 — Tools
The supplied PCA scores array.
Hint 3 — Approach
Return the prepared component-score matrix and inspect its row/column shape.
Explained solution
answer=scores
Rows remain observations while columns become component coordinates rather than original measurements.
Helpful prior knowledge: New coordinates, not selected original columns These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.