Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← PCA lessonsQUESTIONS · MODELS · EVIDENCE
Reading a representation · ML-P05 · 12–18 MIN

Component weights and scores

Distinguish feature-axis weights from row coordinates.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeReason about the Python
  3. TransferAdapt the Python

Understand the idea

Production calls components_.T “loadings”: these are component-axis weights. Other statistical conventions use scaled loadings. Scores locate observations on the axes. An axis sign can flip without changing its information.

Distinguish feature-axis weights from row coordinates.Weights: features × axesScores: rows × axesHow is each axis defined?Where is each observation?Label feature names, row identities and component order.
Schematic · Distinguish feature-axis weights from row coordinates.Scroll the diagram horizontally if needed.

Python skill: Transposes component-by-feature weights into feature-by-component columns.

Meet the syntax

weights = pd.DataFrame(pca.components_.T, index=X.columns)
pca.components_.T
Transposes component-by-feature weights into feature-by-component columns.
index=X.columns
Labels each weight row with its original measurement name.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

answer=pd.DataFrame(pca.components_[:2].T,index=X.columns,columns=['PC1','PC2'])

This practice: Read and run the Python. Next: Change · Component weights and scores.

Given data · PCA48

48 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

Reference labels are omitted from this preview and must remain outside fitting.

PCA48 · first 8 prepared rows
abcde
103.04750.862222.715744.61840.906993
89.600244.301520.270340.1253-0.680116
107.50553.952121.156541.70090.818161
109.40654.250122.525244.90951.39291
80.489640.055714.171429.4087-3.27953
86.978244.138718.721337.5997-1.70077
101.27850.461118.118536.0784-0.343557
96.837648.787517.444533.8533-0.987809

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
afloat64
bfloat64
cfloat64
dfloat64
efloat64
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
X=df.copy()
scaler=StandardScaler()
scaled=scaler.fit_transform(X)
pca=PCA().fit(scaled)
scores=pca.transform(scaled)

Your task · Follow

Build a labelled first-two-component weight table.

Hint 1 — Think

components_ stores axes as rows, while the requested table uses features as rows.

Hint 2 — Tools

Transpose, original feature names and component labels.

Hint 3 — Approach

Take the first two axes, transpose them and attach feature/PC names.

Explained solution
answer=pd.DataFrame(pca.components_[:2].T,index=X.columns,columns=['PC1','PC2'])

The labelled weights distinguish how each original measurement contributes to each new axis.

Helpful prior knowledge: Fit a reusable PCA representation These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Build a labelled first-two-component weight table.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.