Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Classification lessonsQUESTIONS · MODELS · EVIDENCE
Probabilities and neighbours · ML-C09 · 12–18 MIN

Validate k on a meaningful scale

Choose k with fold-local scaling.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeReason about the Python
  3. TransferExplain the Python result

Understand the idea

Validation rows must not become training neighbours. Scaling belongs inside the candidate pipeline so each fold learns its own metric scale.

Choose k with fold-local scaling.?○ Class A □ Class B? New rowNearby training labels vote in the prepared feature space.
Schematic · Choose k with fold-local scaling.Scroll the diagram horizontally if needed.

Python skill: Fits the distance scale within each pipeline fit, including each validation fold.

Meet the syntax

Pipeline([('scale',StandardScaler()),('model',KNeighborsClassifier())])
('scale',StandardScaler())
Fits the distance scale within each pipeline fit, including each validation fold.
('model',KNeighborsClassifier())
Places neighbour voting after scaling; model__n_neighbors can address its k setting.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.neighbors import KNeighborsClassifier
from sklearn.model_selection import GridSearchCV,StratifiedKFold
model=Pipeline([('scale',StandardScaler()),('model',KNeighborsClassifier())])
search=GridSearchCV(model,{'model__n_neighbors':[3,5,9]},cv=StratifiedKFold(5,shuffle=True,random_state=42),scoring='f1_macro').fit(X_train,y_train)
answer=search.cv_results_['mean_test_score']

This practice: Read and run the Python. Next: Change · Validate k on a meaningful scale.

Given data · CLASS180

180 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

CLASS180 · first 8 prepared rows
lengthwidthlabel
236.033-0.805568A
581.2970.728558A
-1511.27-1.00866A
99.0248-0.24496A
-13.0141-0.660765A
681.1790.602475A
51.14720.873157A
362.131-0.665605A

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
lengthfloat64
widthfloat64
labelstr
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.model_selection import train_test_split
X=df[['length','width']]
y=df.label
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42,stratify=y)

Your task · Follow

Search production k values 3,5,9 on CLASS180.

Hint 1 — Think

Scaling must be learned separately inside each validation fit.

Hint 2 — Tools

Pipeline, GridSearchCV and model__n_neighbors.

Hint 3 — Approach

Compare the declared k values in a scaling/KNN pipeline using seeded stratified macro-F1 folds.

Explained solution
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.neighbors import KNeighborsClassifier
from sklearn.model_selection import GridSearchCV,StratifiedKFold
model=Pipeline([('scale',StandardScaler()),('model',KNeighborsClassifier())])
search=GridSearchCV(model,{'model__n_neighbors':[3,5,9]},cv=StratifiedKFold(5,shuffle=True,random_state=42),scoring='f1_macro').fit(X_train,y_train)
answer=search.cv_results_['mean_test_score']

Fold-local scaling keeps validation observations from changing distances used to fit their own predictor.

Helpful prior knowledge: Neighbour voting These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Search production k values 3,5,9 on CLASS180.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.