Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Classification lessonsQUESTIONS · MODELS · EVIDENCE
Probabilities and neighbours · ML-C05 · 12–18 MIN

Logistic regression

Understand a regularised linear class boundary.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeExplain the Python result
  3. TransferAdapt the Python

Understand the idea

Logistic regression is a classifier despite its name. It models class log-odds with a linear score and uses regularisation. Scaling helps optimisation and makes regularisation more comparable across inputs.

Understand a regularised linear class boundary.Class AClass BLogistic classification uses a linear decision boundary.
Schematic · Understand a regularised linear class boundary.Scroll the diagram horizontally if needed.

Python skill: Creates a classification estimator despite the word regression in its name.

Meet the syntax

LogisticRegression(max_iter=2000, random_state=42)
LogisticRegression
Creates a classification estimator despite the word regression in its name.
max_iter=2000
Sets an optimisation iteration limit; it does not guarantee convergence.
random_state=42
Makes supported stochastic fitting choices repeatable.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
model=Pipeline([('scale',StandardScaler()),('model',LogisticRegression(max_iter=2000,random_state=42))]).fit(X_train,y_train)
answer=model.predict(X_test)

This practice: Read and run the Python. Next: Change · Logistic regression.

Logistic classification · from idea to workflow

Question: Which class is plausible from available inputs?

Mechanism: A linear score becomes class probabilities and then a decision.

Watch for: A linear boundary can miss curved separation; probability and decision cost are different.

Interpret or debug · self-review: Explain which error changes when the decision threshold moves, and why threshold selection belongs inside training validation.

Independent full workflow: ML-X05 · ML-X16. First practise complete regression and classification in F11–F12; check readiness in W-K2–W-K3.

Given data · CLASS180

180 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

CLASS180 · first 8 prepared rows
lengthwidthlabel
236.033-0.805568A
581.2970.728558A
-1511.27-1.00866A
99.0248-0.24496A
-13.0141-0.660765A
681.1790.602475A
51.14720.873157A
362.131-0.665605A

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
lengthfloat64
widthfloat64
labelstr
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.model_selection import train_test_split
X=df[['length','width']]
y=df.label
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42,stratify=y)

Your task · Follow

Fit a scaled logistic classifier and predict test labels.

Hint 1 — Think

The classifier’s optimisation uses the prepared feature scale learned from training rows.

Hint 2 — Tools

Pipeline, StandardScaler and LogisticRegression.

Hint 3 — Approach

Build the named scaling/classification steps, fit the training pair and predict through the complete pipeline.

Explained solution
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
model=Pipeline([('scale',StandardScaler()),('model',LogisticRegression(max_iter=2000,random_state=42))]).fit(X_train,y_train)
answer=model.predict(X_test)

The pipeline reuses training statistics for test rows while the iteration budget supports convergence.

Helpful prior knowledge: Labels and probabilities These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Fit a scaled logistic classifier and predict test labels.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.