Understand the idea
GridSearchCV compares declared candidate settings through training-fold validation, selects by the scorer and refits the chosen candidate on training rows. LinearRegression has no search in the shared checkpoint; keeping defaults is a valid decision.
Python skill: Use a parameter dictionary and GridSearchCV to compare settings inside training folds.
Meet the syntax
GridSearchCV
'model__fit_intercept'
[True, False]
search.cv_results_GridSearchCV- Fits each candidate setting on the training side of each fold, compares validation scores and refits the selected candidate.
'model__fit_intercept'- The double underscore addresses the setting inside the pipeline step named model.
[True, False]- The two candidate values for this teaching experiment. Production linear workflows keep their defaults.
search.cv_results_- The fitted search’s result dictionary; mean_test_score is validation evidence, not the sealed final test.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
from sklearn.model_selection import GridSearchCV, KFold
from sklearn.linear_model import LinearRegression
from sklearn.pipeline import Pipeline
model = Pipeline([('model', LinearRegression())])
folds = KFold(n_splits=5, shuffle=True, random_state=42)
search = GridSearchCV(model, {'model__fit_intercept': [True, False]}, cv=folds, scoring='neg_root_mean_squared_error')
search.fit(X_train, y_train)
answer = pd.DataFrame(search.cv_results_)[['params', 'mean_test_score']]
This practice: Read and run the Python. Next: Change · How a search makes a choice.
Given data · LINE24
24 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.
| distance | duration |
|---|---|
| 1 | 5.30501 |
| 1.47826 | 11.2361 |
| 1.95652 | 11.8488 |
| 2.43478 | 14.809 |
| 2.91304 | 10.7589 |
| 3.3913 | 19.4972 |
| 3.86957 | 17.89 |
| 4.34783 | 15.531 |
Column meanings and units
Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.
Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.
| Column | Stored type |
|---|---|
| distance | float64 |
| duration | float64 |
Supplied setup · available if you need to inspect it
This code runs before your editor on every Run. These are the objects your exercise uses.
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(df[['distance']], df['duration'], test_size=0.2, random_state=42)Your task · Follow
Compare fitting a line with and without an intercept using five training folds. Store the candidate parameters and mean validation scores in answer.
Hint 1 — Think
Use a parameter dictionary and GridSearchCV to compare settings inside training folds.
Hint 2 — Tools
Use GridSearchCV, 'model__fit_intercept', [True, False], search.cv_results_. Read the visible syntax meanings before editing.
Hint 3 — Approach
Compare fitting a line with and without an intercept using five training folds. Store the candidate parameters and mean validation scores in answer. Keep the supplied row order and inspect the named output after running.
Explained solution
from sklearn.model_selection import GridSearchCV, KFold
from sklearn.linear_model import LinearRegression
from sklearn.pipeline import Pipeline
model = Pipeline([('model', LinearRegression())])
folds = KFold(n_splits=5, shuffle=True, random_state=42)
search = GridSearchCV(model, {'model__fit_intercept': [True, False]}, cv=folds, scoring='neg_root_mean_squared_error')
search.fit(X_train, y_train)
answer = pd.DataFrame(search.cv_results_)[['params', 'mean_test_score']]
Each intercept choice is evaluated on identical training folds. Negative RMSE is larger when the error is smaller. The search never receives X_test or y_test. This small experiment teaches search mechanics; the playground’s linear workflow keeps its defaults.
Helpful prior knowledge: Settings, learned values and fit quality These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.