Understand the idea
SVM fits boundaries influenced by support vectors near margins. Production SVC uses an RBF kernel and does not enable probability estimation. Decision values are not probabilities.
Python skill: Creates a support-vector classifier; its default kernel is RBF.
Meet the syntax
SVC(random_state=42)
model.decision_function(X_test)
model.named_steps["model"].n_support_SVC- Creates a support-vector classifier; its default kernel is RBF.
model.decision_function(X_test)- Returns signed margin-based decision values; these are not probabilities.
model.named_steps["model"].n_support_- Reads the fitted SVC support-vector count per class from the final pipeline step.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
answer=model.predict(X_test)
This practice: Read and run the Python. Next: Change · Margins and support vectors.
Support vector classification · from idea to workflow
Question: Can a margin separate classes in the prepared feature space?
Mechanism: Support vectors define a margin; a kernel can make the boundary curved.
Watch for: Unscaled inputs distort geometry; a flexible kernel can overfit.
Interpret or debug · self-review: Compare a linear and curved boundary using matching folds. Explain why a perfect training score is insufficient.
Guided full workflow: Apply this model with a baseline, training validation and one final evaluation →
Independent full workflow: ML-X08. First practise complete regression and classification in F11–F12; check readiness in W-K2–W-K3.
Given data · CLASS180
180 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.
| length | width | label |
|---|---|---|
| 236.033 | -0.805568 | A |
| 581.297 | 0.728558 | A |
| -1511.27 | -1.00866 | A |
| 99.0248 | -0.24496 | A |
| -13.0141 | -0.660765 | A |
| 681.179 | 0.602475 | A |
| 51.1472 | 0.873157 | A |
| 362.131 | -0.665605 | A |
Column meanings and units
Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.
Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.
| Column | Stored type |
|---|---|
| length | float64 |
| width | float64 |
| label | str |
Supplied setup · available if you need to inspect it
This code runs before your editor on every Run. These are the objects your exercise uses.
from sklearn.model_selection import train_test_split
X=df[['length','width']]
y=df.label
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42,stratify=y)
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.svm import SVC
model=Pipeline([('scale',StandardScaler()),('model',SVC(random_state=42))]).fit(X_train,y_train)
Your task · Follow
Use the supplied fitted scaled SVM to predict held-away labels.
Hint 1 — Think
The supplied fitted pipeline already contains the scale and margin model.
Hint 2 — Tools
Pipeline.predict.
Hint 3 — Approach
Pass the held-away features through the existing fitted pipeline.
Explained solution
answer=model.predict(X_test)
Prediction applies the learned preparation and SVM decision rule without refitting on evaluation rows.
Helpful prior knowledge: Macro F1 and imbalance · Learn a scale from training rows These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.