Understand the idea
Logistic regression is a classifier despite its name. It models class log-odds with a linear score and uses regularisation. Scaling helps optimisation and makes regularisation more comparable across inputs.
Python skill: Creates a classification estimator despite the word regression in its name.
Meet the syntax
LogisticRegression(max_iter=2000, random_state=42)LogisticRegression- Creates a classification estimator despite the word regression in its name.
max_iter=2000- Sets an optimisation iteration limit; it does not guarantee convergence.
random_state=42- Makes supported stochastic fitting choices repeatable.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
model=Pipeline([('scale',StandardScaler()),('model',LogisticRegression(max_iter=2000,random_state=42))]).fit(X_train,y_train)
answer=model.predict(X_test)
This practice: Read and run the Python. Next: Change · Logistic regression.
Logistic classification · from idea to workflow
Question: Which class is plausible from available inputs?
Mechanism: A linear score becomes class probabilities and then a decision.
Watch for: A linear boundary can miss curved separation; probability and decision cost are different.
Interpret or debug · self-review: Explain which error changes when the decision threshold moves, and why threshold selection belongs inside training validation.
Independent full workflow: ML-X05 · ML-X16. First practise complete regression and classification in F11–F12; check readiness in W-K2–W-K3.
Given data · CLASS180
180 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.
| length | width | label |
|---|---|---|
| 236.033 | -0.805568 | A |
| 581.297 | 0.728558 | A |
| -1511.27 | -1.00866 | A |
| 99.0248 | -0.24496 | A |
| -13.0141 | -0.660765 | A |
| 681.179 | 0.602475 | A |
| 51.1472 | 0.873157 | A |
| 362.131 | -0.665605 | A |
Column meanings and units
Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.
Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.
| Column | Stored type |
|---|---|
| length | float64 |
| width | float64 |
| label | str |
Supplied setup · available if you need to inspect it
This code runs before your editor on every Run. These are the objects your exercise uses.
from sklearn.model_selection import train_test_split
X=df[['length','width']]
y=df.label
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42,stratify=y)
Your task · Follow
Fit a scaled logistic classifier and predict test labels.
Hint 1 — Think
The classifier’s optimisation uses the prepared feature scale learned from training rows.
Hint 2 — Tools
Pipeline, StandardScaler and LogisticRegression.
Hint 3 — Approach
Build the named scaling/classification steps, fit the training pair and predict through the complete pipeline.
Explained solution
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
model=Pipeline([('scale',StandardScaler()),('model',LogisticRegression(max_iter=2000,random_state=42))]).fit(X_train,y_train)
answer=model.predict(X_test)
The pipeline reuses training statistics for test rows while the iteration budget supports convergence.
Helpful prior knowledge: Labels and probabilities These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.