Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Supervised Workflow lessonsQUESTIONS · MODELS · EVIDENCE
Select, diagnose and finish · ML-W11 · 12–18 MIN

How a search makes a choice

Understand search boundaries before model-specific tuning.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeReason about the Python
  3. TransferExplain the Python result

Understand the idea

GridSearchCV compares declared candidate settings through training-fold validation, selects by the scorer and refits the chosen candidate on training rows. LinearRegression has no search in the shared checkpoint; keeping defaults is a valid decision.

Understand search boundaries before model-specific tuning.Fold 1VTTTTFold 2TVTTTFold 3TTVTTFold 4TTTVTFold 5TTTTVFinalT: fit on training rows V: validate Final: held away
Schematic · Understand search boundaries before model-specific tuning.Scroll the diagram horizontally if needed.

Python skill: Use a parameter dictionary and GridSearchCV to compare settings inside training folds.

Meet the syntax

GridSearchCV
'model__fit_intercept'
[True, False]
search.cv_results_
GridSearchCV
Fits each candidate setting on the training side of each fold, compares validation scores and refits the selected candidate.
'model__fit_intercept'
The double underscore addresses the setting inside the pipeline step named model.
[True, False]
The two candidate values for this teaching experiment. Production linear workflows keep their defaults.
search.cv_results_
The fitted search’s result dictionary; mean_test_score is validation evidence, not the sealed final test.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

from sklearn.model_selection import GridSearchCV, KFold
from sklearn.linear_model import LinearRegression
from sklearn.pipeline import Pipeline
model = Pipeline([('model', LinearRegression())])
folds = KFold(n_splits=5, shuffle=True, random_state=42)
search = GridSearchCV(model, {'model__fit_intercept': [True, False]}, cv=folds, scoring='neg_root_mean_squared_error')
search.fit(X_train, y_train)
answer = pd.DataFrame(search.cv_results_)[['params', 'mean_test_score']]

This practice: Read and run the Python. Next: Change · How a search makes a choice.

Given data · LINE24

24 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

LINE24 · first 8 prepared rows
distanceduration
15.30501
1.4782611.2361
1.9565211.8488
2.4347814.809
2.9130410.7589
3.391319.4972
3.8695717.89
4.3478315.531

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
distancefloat64
durationfloat64
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(df[['distance']], df['duration'], test_size=0.2, random_state=42)

Your task · Follow

Compare fitting a line with and without an intercept using five training folds. Store the candidate parameters and mean validation scores in answer.

Hint 1 — Think

Use a parameter dictionary and GridSearchCV to compare settings inside training folds.

Hint 2 — Tools

Use GridSearchCV, 'model__fit_intercept', [True, False], search.cv_results_. Read the visible syntax meanings before editing.

Hint 3 — Approach

Compare fitting a line with and without an intercept using five training folds. Store the candidate parameters and mean validation scores in answer. Keep the supplied row order and inspect the named output after running.

Explained solution
from sklearn.model_selection import GridSearchCV, KFold
from sklearn.linear_model import LinearRegression
from sklearn.pipeline import Pipeline
model = Pipeline([('model', LinearRegression())])
folds = KFold(n_splits=5, shuffle=True, random_state=42)
search = GridSearchCV(model, {'model__fit_intercept': [True, False]}, cv=folds, scoring='neg_root_mean_squared_error')
search.fit(X_train, y_train)
answer = pd.DataFrame(search.cv_results_)[['params', 'mean_test_score']]

Each intercept choice is evaluated on identical training folds. Negative RMSE is larger when the error is smaller. The search never receives X_test or y_test. This small experiment teaches search mechanics; the playground’s linear workflow keeps its defaults.

Helpful prior knowledge: Settings, learned values and fit quality These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Compare fitting a line with and without an intercept using five training folds. Store the candidate parameters and mean validation scores in answer.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.