Understand the idea
KNN stores training examples and predicts from nearby examples. k controls the neighbourhood size. Distance depends on input scales and can become less discriminating in many dimensions.
Python skill: Predicts a class by votes from fitted training neighbours.
Meet the syntax
KNeighborsClassifier(n_neighbors=3)
model.kneighbors(row, return_distance=False)KNeighborsClassifier- Predicts a class by votes from fitted training neighbours.
n_neighbors=3- Uses the three nearest training observations for each prediction.
model.kneighbors(row, return_distance=False)- Returns positions of the nearest fitted training observations in the prepared feature space.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
from sklearn.neighbors import KNeighborsClassifier
from sklearn.preprocessing import StandardScaler
scaler=StandardScaler().fit(X_train)
model=KNeighborsClassifier(n_neighbors=3).fit(scaler.transform(X_train),y_train)
answer=model.kneighbors(scaler.transform(X_test.iloc[:1]),return_distance=False)
This practice: Read and run the Python. Next: Change · Neighbour voting.
Nearest neighbours · from idea to workflow
Question: Do nearby measured examples suggest a label?
Mechanism: Scaled distances identify neighbours whose class labels vote.
Watch for: Units and irrelevant features change neighbourhoods; small k can be noisy.
Interpret or debug · self-review: If a measurement changes units, explain why the unscaled prediction can change. Compare k with a fold-local scaler.
Independent full workflow: ML-X07. First practise complete regression and classification in F11–F12; check readiness in W-K2–W-K3.
Given data · CLASS18
18 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.
| length | width | label |
|---|---|---|
| 236.033 | -0.805568 | A |
| 581.297 | 0.728558 | A |
| -1511.27 | -1.00866 | A |
| 99.0248 | -0.24496 | A |
| -13.0141 | -0.660765 | A |
| 681.179 | 0.602475 | A |
| 2051.15 | 1.87316 | B |
| 2362.13 | 0.334395 | B |
| 2285.63 | 0.257253 | B |
| 2680.44 | 0.961328 | B |
| 1856.81 | 0.472554 | B |
| 2946.98 | 0.880302 | B |
| -331.781 | 2.72724 | C |
| 412.325 | 3.28307 | C |
| 319.701 | 3.33371 | C |
| 1658.91 | 2.68519 | C |
| -396.782 | 2.36965 | C |
| 477.136 | 3.8745 | C |
Column meanings and units
Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.
Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.
| Column | Stored type |
|---|---|
| length | float64 |
| width | float64 |
| label | str |
Supplied setup · available if you need to inspect it
This code runs before your editor on every Run. These are the objects your exercise uses.
from sklearn.model_selection import train_test_split
X=df[['length','width']]
y=df.label
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42,stratify=y)
Your task · Follow
Fit three-neighbour KNN on scaled training features. Store the positions of the three nearest training rows for the first test row in answer.
Hint 1 — Think
Neighbour indices refer to the fitted training representation.
Hint 2 — Tools
StandardScaler, KNeighborsClassifier and kneighbors.
Hint 3 — Approach
Learn the scale on training rows, fit three-neighbour KNN and query the transformed first test row.
Explained solution
from sklearn.neighbors import KNeighborsClassifier
from sklearn.preprocessing import StandardScaler
scaler=StandardScaler().fit(X_train)
model=KNeighborsClassifier(n_neighbors=3).fit(scaler.transform(X_train),y_train)
answer=model.kneighbors(scaler.transform(X_test.iloc[:1]),return_distance=False)
Using the same coordinate system makes the returned neighbour positions meaningful for the fitted training table.
Helpful prior knowledge: Macro F1 and imbalance · Learn a scale from training rows These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.