Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Classification lessonsQUESTIONS · MODELS · EVIDENCE
Class distributions · ML-C17 · 12–18 MIN

LDA shares a covariance shape

Understand shared class geometry and linear boundaries.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeAdapt the Python
  3. TransferAdapt the Python

Understand the idea

LDA gives classes different means but a shared covariance shape. Production uses the lsqr solver and optionally shrinkage. This module uses LDA as a classifier, not as a PCA replacement.

Understand shared class geometry and linear boundaries.Class AClass BLDA: different class means, one shared covariance shape.
Schematic · Understand shared class geometry and linear boundaries.Scroll the diagram horizontally if needed.

Python skill: Uses the least-squares LDA solver, which supports shrinkage.

Meet the syntax

LinearDiscriminantAnalysis(solver='lsqr', shrinkage='auto')
model.means_
solver='lsqr'
Uses the least-squares LDA solver, which supports shrinkage.
shrinkage='auto'
Estimates covariance shrinkage from the fitting rows.
LinearDiscriminantAnalysis
Fits class means with a shared covariance structure.
model.means_
Reads the fitted class-by-feature mean matrix; label its axes with classes_ and feature names.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

answer=pd.DataFrame(model.means_,index=model.classes_,columns=X_train.columns)

This practice: Read and run the Python. Next: Change · LDA shares a covariance shape.

Linear discriminant analysis · from idea to workflow

Question: Can class means and shared variation separate labels?

Mechanism: Use class means with a shared within-class covariance to form linear boundaries.

Watch for: Different covariance shapes or redundant inputs can undermine the model.

Interpret or debug · self-review: Compare class spreads before defending a shared covariance. Use validation to test the modelling simplification.

Independent full workflow: ML-X14. First practise complete regression and classification in F11–F12; check readiness in W-K2–W-K3.

Given data · COV90_SHARED

90 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

COV90_SHARED · first 8 prepared rows
x1x2label
-0.743901-0.783885A
-2.12019-0.179627A
1.340741.03856A
-0.952592-0.268834A
-0.558928-0.411972A
-2.15048-0.372685A
-1.62080.506436A
-0.975229-0.834521A

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
x1float64
x2float64
labelstr
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.model_selection import train_test_split
X=df[['x1','x2']]
y=df.label
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42,stratify=y)
from sklearn.discriminant_analysis import LinearDiscriminantAnalysis
model=LinearDiscriminantAnalysis(solver='lsqr').fit(X_train,y_train)

Your task · Follow

Build answer as a table of fitted LDA means, with model.classes_ as rows and X_train column names as columns.

Hint 1 — Think

Class means need the same feature names and class order used by the fitted model.

Hint 2 — Tools

means_, classes_ and a labelled dataframe.

Hint 3 — Approach

Read the fitted means and attach class rows plus training-feature columns.

Explained solution
answer=pd.DataFrame(model.means_,index=model.classes_,columns=X_train.columns)

The table exposes each class centre while leaving the shared covariance assumption conceptually separate.

Helpful prior knowledge: Gaussian Naive Bayes These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Build answer as a table of fitted LDA means, with model.classes_ as rows and X_train column names as columns.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.