Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Classification lessonsQUESTIONS · MODELS · EVIDENCE
Margins and rules · ML-C11 · 12–18 MIN

Nonlinear SVM behaviour

Relate C, kernel scale and cost to validation.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeReason about the Python
  3. TransferExplain the Python result

Understand the idea

RBF kernels support nonlinear boundaries. C controls the penalty/regularisation trade-off; gamma controls the locality of the kernel. Production searches C and keeps gamma="scale".

Relate C, kernel scale and cost to validation.An RBF boundary can curve; validate its flexibility.
Schematic · Relate C, kernel scale and cost to validation.Scroll the diagram horizontally if needed.

Python skill: Controls the penalty for training margin violations in the pipeline SVM.

Meet the syntax

GridSearchCV(model, {'model__C':[.5,2,10]}, scoring='f1_macro', cv=folds)
'model__C'
Controls the penalty for training margin violations in the pipeline SVM.
scoring='f1_macro'
Judges settings by class-balanced validation evidence, not by margin magnitude.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

from sklearn.model_selection import GridSearchCV,StratifiedKFold
search=GridSearchCV(model,{'model__C':[.5,2,10]},cv=StratifiedKFold(5,shuffle=True,random_state=42),scoring='f1_macro').fit(X_train,y_train)
answer=search.cv_results_['mean_test_score']

This practice: Read and run the Python. Next: Change · Nonlinear SVM behaviour.

Given data · CLASS180

180 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

CLASS180 · first 8 prepared rows
lengthwidthlabel
236.033-0.805568A
581.2970.728558A
-1511.27-1.00866A
99.0248-0.24496A
-13.0141-0.660765A
681.1790.602475A
51.14720.873157A
362.131-0.665605A

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
lengthfloat64
widthfloat64
labelstr
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.model_selection import train_test_split
X=df[['length','width']]
y=df.label
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42,stratify=y)
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.svm import SVC
model=Pipeline([('scale',StandardScaler()),('model',SVC(random_state=42))]).fit(X_train,y_train)

Your task · Follow

Search C=.5,2,10 using five stratified folds.

Hint 1 — Think

C changes the margin/error trade-off and must be chosen before final testing.

Hint 2 — Tools

GridSearchCV, StratifiedKFold and model__C.

Hint 3 — Approach

Evaluate the three declared C values with training-only macro-F1 folds.

Explained solution
from sklearn.model_selection import GridSearchCV,StratifiedKFold
search=GridSearchCV(model,{'model__C':[.5,2,10]},cv=StratifiedKFold(5,shuffle=True,random_state=42),scoring='f1_macro').fit(X_train,y_train)
answer=search.cv_results_['mean_test_score']

The shared folds compare regularisation choices while the pipeline keeps scaling within each fit.

Helpful prior knowledge: Margins and support vectors These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Search C=.5,2,10 using five stratified folds.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.