Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Classification lessonsQUESTIONS · MODELS · EVIDENCE
Probabilities and neighbours · ML-C04 · 12–18 MIN

Labels and probabilities

Align probability columns with class labels.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeAdapt the Python
  3. TransferAdapt the Python

Understand the idea

predict_proba columns follow classes_, not a guessed order. Thresholds turn binary probabilities into labels; multiclass prediction usually chooses the largest class score.

Align probability columns with class labels.Probability for class BRow 1Row 2Row 3A threshold turns a probability into a class decision.
Schematic · Align probability columns with class labels.Scroll the diagram horizontally if needed.

Python skill: Gives the fitted class order used by probability columns.

Meet the syntax

model.classes_
model.predict_proba(X_test)
model.classes_
Gives the fitted class order used by probability columns.
model.predict_proba(X_test)
Returns a row-by-class array of probabilities, not one label per row.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

answer=pd.DataFrame(model.predict_proba(X_test),index=X_test.index,columns=model.classes_)

This practice: Read and run the Python. Next: Change · Labels and probabilities.

Given data · CLASS180

180 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

CLASS180 · first 8 prepared rows
lengthwidthlabel
236.033-0.805568A
581.2970.728558A
-1511.27-1.00866A
99.0248-0.24496A
-13.0141-0.660765A
681.1790.602475A
51.14720.873157A
362.131-0.665605A

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
lengthfloat64
widthfloat64
labelstr
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.model_selection import train_test_split
X=df[['length','width']]
y=df.label
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42,stratify=y)
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
model=Pipeline([('scale',StandardScaler()),('model',LogisticRegression(max_iter=2000,random_state=42))]).fit(X_train,y_train)

Your task · Follow

Create a dataframe of test probabilities with the actual class labels as columns.

Hint 1 — Think

Probability columns follow the estimator’s fitted class order.

Hint 2 — Tools

predict_proba, classes_ and dataframe index.

Hint 3 — Approach

Build the probability table with test-row indices and fitted class names.

Explained solution
answer=pd.DataFrame(model.predict_proba(X_test),index=X_test.index,columns=model.classes_)

Explicit row and class labels prevent a probability column from being attributed to the wrong observation or class.

Helpful prior knowledge: Macro F1 and imbalance These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Create a dataframe of test probabilities with the actual class labels as columns.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.