Understand the idea
PCA rotates numeric measurements into component axes and can retain fewer axes. Components combine original features; they are not selected original columns or cluster assignments.
Python skill: Call fit_transform() to learn a representation and return coordinates in the new feature space.
Meet the syntax
StandardScaler().fit_transform(X)
PCA(n_components=2)
pca.fit_transform(scaled)StandardScaler().fit_transform(X)- Learns a common scale from the declared discovery population and transforms it.
PCA(n_components=2)- Requests two new axes, each combining the original measurements.
pca.fit_transform(scaled)- Learns these axes and returns two coordinates per input row; it does not select two original columns.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
X = df.select_dtypes(include='number')
scaled = StandardScaler().fit_transform(X)
pca = PCA(n_components=2)
answer = pca.fit_transform(scaled)
This practice: Read and run the Python. Next: Change · New coordinates, not selected original columns.
Given data · PCA48
48 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.
Reference labels are omitted from this preview and must remain outside fitting.
| a | b | c | d | e |
|---|---|---|---|---|
| 103.047 | 50.8622 | 22.7157 | 44.6184 | 0.906993 |
| 89.6002 | 44.3015 | 20.2703 | 40.1253 | -0.680116 |
| 107.505 | 53.9521 | 21.1565 | 41.7009 | 0.818161 |
| 109.406 | 54.2501 | 22.5252 | 44.9095 | 1.39291 |
| 80.4896 | 40.0557 | 14.1714 | 29.4087 | -3.27953 |
| 86.9782 | 44.1387 | 18.7213 | 37.5997 | -1.70077 |
| 101.278 | 50.4611 | 18.1185 | 36.0784 | -0.343557 |
| 96.8376 | 48.7875 | 17.4445 | 33.8533 | -0.987809 |
Column meanings and units
Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.
Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.
| Column | Stored type |
|---|---|
| a | float64 |
| b | float64 |
| c | float64 |
| d | float64 |
| e | float64 |
Your task · Follow
Standardise the supplied measurements, create two PCA coordinates for every row, and store the coordinates in answer.
Hint 1 — Think
Call fit_transform() to learn a representation and return coordinates in the new feature space.
Hint 2 — Tools
Use StandardScaler().fit_transform(X), PCA(n_components=2), pca.fit_transform(scaled). Read the visible syntax meanings before editing.
Hint 3 — Approach
Standardise the supplied measurements, create two PCA coordinates for every row, and store the coordinates in answer. Keep the supplied row order and inspect the named output after running.
Explained solution
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
X = df.select_dtypes(include='number')
scaled = StandardScaler().fit_transform(X)
pca = PCA(n_components=2)
answer = pca.fit_transform(scaled)
The output has one row per observation and two new coordinate columns. Each coordinate combines original measurements. This first two-axis picture introduces the API; later lessons use explained variance to decide how many components to retain.
Helpful prior knowledge: Distance depends on scale and context These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.