Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Classification lessonsQUESTIONS · MODELS · EVIDENCE
Probabilities and neighbours · ML-C08 · 12–18 MIN

Neighbour voting

Explain a classifier’s local distance-based decision.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeAdapt the Python
  3. TransferExplain the Python result

Understand the idea

KNN stores training examples and predicts from nearby examples. k controls the neighbourhood size. Distance depends on input scales and can become less discriminating in many dimensions.

Explain a classifier’s local distance-based decision.?○ Class A □ Class B? New rowNearby training labels vote in the prepared feature space.
Schematic · Explain a classifier’s local distance-based decision.Scroll the diagram horizontally if needed.

Python skill: Predicts a class by votes from fitted training neighbours.

Meet the syntax

KNeighborsClassifier(n_neighbors=3)
model.kneighbors(row, return_distance=False)
KNeighborsClassifier
Predicts a class by votes from fitted training neighbours.
n_neighbors=3
Uses the three nearest training observations for each prediction.
model.kneighbors(row, return_distance=False)
Returns positions of the nearest fitted training observations in the prepared feature space.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

from sklearn.neighbors import KNeighborsClassifier
from sklearn.preprocessing import StandardScaler
scaler=StandardScaler().fit(X_train)
model=KNeighborsClassifier(n_neighbors=3).fit(scaler.transform(X_train),y_train)
answer=model.kneighbors(scaler.transform(X_test.iloc[:1]),return_distance=False)

This practice: Read and run the Python. Next: Change · Neighbour voting.

Nearest neighbours · from idea to workflow

Question: Do nearby measured examples suggest a label?

Mechanism: Scaled distances identify neighbours whose class labels vote.

Watch for: Units and irrelevant features change neighbourhoods; small k can be noisy.

Interpret or debug · self-review: If a measurement changes units, explain why the unscaled prediction can change. Compare k with a fold-local scaler.

Independent full workflow: ML-X07. First practise complete regression and classification in F11–F12; check readiness in W-K2–W-K3.

Given data · CLASS18

18 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

CLASS18 · all rows
lengthwidthlabel
236.033-0.805568A
581.2970.728558A
-1511.27-1.00866A
99.0248-0.24496A
-13.0141-0.660765A
681.1790.602475A
2051.151.87316B
2362.130.334395B
2285.630.257253B
2680.440.961328B
1856.810.472554B
2946.980.880302B
-331.7812.72724C
412.3253.28307C
319.7013.33371C
1658.912.68519C
-396.7822.36965C
477.1363.8745C

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
lengthfloat64
widthfloat64
labelstr
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.model_selection import train_test_split
X=df[['length','width']]
y=df.label
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42,stratify=y)

Your task · Follow

Fit three-neighbour KNN on scaled training features. Store the positions of the three nearest training rows for the first test row in answer.

Hint 1 — Think

Neighbour indices refer to the fitted training representation.

Hint 2 — Tools

StandardScaler, KNeighborsClassifier and kneighbors.

Hint 3 — Approach

Learn the scale on training rows, fit three-neighbour KNN and query the transformed first test row.

Explained solution
from sklearn.neighbors import KNeighborsClassifier
from sklearn.preprocessing import StandardScaler
scaler=StandardScaler().fit(X_train)
model=KNeighborsClassifier(n_neighbors=3).fit(scaler.transform(X_train),y_train)
answer=model.kneighbors(scaler.transform(X_test.iloc[:1]),return_distance=False)

Using the same coordinate system makes the returned neighbour positions meaningful for the fitted training table.

Helpful prior knowledge: Macro F1 and imbalance · Learn a scale from training rows These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Fit three-neighbour KNN on scaled training features. Store the positions of the three nearest training rows for the first test row in answer.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.