Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Classification lessonsQUESTIONS · MODELS · EVIDENCE
Class distributions · ML-C19 · 12–18 MIN

Covariance support and regularisation limits

Qualify numerical stability and unit effects.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeExplain the Python result
  3. TransferExplain the Python result

Understand the idea

Regularisation can stabilise a covariance estimate but cannot manufacture missing class information. Adding a regularising identity term also means unit choices can matter.

Qualify numerical stability and unit effects.Class AClass BCompare means, covariance shapes and assumptions.
Schematic · Qualify numerical stability and unit effects.Scroll the diagram horizontally if needed.

Python skill: Filter by class and use cov() to inspect the covariance information available to QDA.

Meet the syntax

df['label'].eq('A')
class_a.cov()
df['label'].eq('A')
Creates a Boolean mask selecting examples from one class.
class_a.cov()
Returns feature-by-feature sample covariance. Its values depend on the measurement units.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

counts = df['label'].value_counts()
class_a = df.loc[df['label'].eq('A'), ['length', 'width']]
answer = class_a.cov()

This practice: Read and run the Python. Next: Change · Covariance support and regularisation limits.

Given data · CLASS180

180 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

CLASS180 · first 8 prepared rows
lengthwidthlabel
236.033-0.805568A
581.2970.728558A
-1511.27-1.00866A
99.0248-0.24496A
-13.0141-0.660765A
681.1790.602475A
51.14720.873157A
362.131-0.665605A

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
lengthfloat64
widthfloat64
labelstr

Your task · Follow

Count observations in each class and calculate the covariance of the two input measurements within class A. Store the covariance in answer.

Hint 1 — Think

Filter by class and use cov() to inspect the covariance information available to QDA.

Hint 2 — Tools

Use df['label'].eq('A'), class_a.cov(). Read the visible syntax meanings before editing.

Hint 3 — Approach

Count observations in each class and calculate the covariance of the two input measurements within class A. Store the covariance in answer. Keep the supplied row order and inspect the named output after running.

Explained solution
counts = df['label'].value_counts()
class_a = df.loc[df['label'].eq('A'), ['length', 'width']]
answer = class_a.cov()

QDA estimates a separate covariance structure for each class. A small class or nearly redundant measurements can make that estimate unstable. This table shows what is being estimated, not proof of stability. Use regularisation and validation; adding a parameter does not create new observations.

Helpful prior knowledge: QDA allows class-specific shapes These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Count observations in each class and calculate the covariance of the two input measurements within class A. Store the covariance in answer.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.