Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Supervised Workflow lessonsQUESTIONS · MODELS · EVIDENCE
Validate and compare · ML-W09 · 18–25 MIN

Read cross-validation evidence

Interpret scores and compare on matching folds.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeAdapt the Python
  3. PractiseAdapt the Python
  4. TransferReason about the Python

Understand the idea

sklearn scorers follow “larger is better”, so RMSE scoring is negated. The test_score key returned by cross_validate refers to fold validation, not the final test. A dummy reference uses the same folds.

Interpret scores and compare on matching folds.Candidate ACandidate BFitValidateCostCompare matching evidence; smaller error can cost more.
Schematic · Interpret scores and compare on matching folds.Scroll the diagram horizontally if needed.

Python skill: Fits independent copies on the training part of each fold and returns timing/score arrays.

Meet the syntax

cross_validate(model, X_train, y_train, cv=folds, scoring='neg_root_mean_squared_error')
cross_validate
Fits independent copies on the training part of each fold and returns timing/score arrays.
cv=folds
Reuses the declared fold design across comparisons.
scoring='neg_root_mean_squared_error'
Uses negative RMSE so larger scores are better; negate test_score to report positive RMSE.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

answer=cross_validate(model,X_train,y_train,cv=folds,scoring='neg_root_mean_squared_error')

This practice: Read and run the Python. Next: Change · Read cross-validation evidence.

Given data · LINE24

24 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

LINE24 · first 8 prepared rows
distanceduration
15.30501
1.4782611.2361
1.9565211.8488
2.4347814.809
2.9130410.7589
3.391319.4972
3.8695717.89
4.3478315.531

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
distancefloat64
durationfloat64
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.model_selection import train_test_split
X = df[['distance']]
y = df['duration']
X_train, X_test, y_train, y_test = train_test_split(X,y,test_size=.2,random_state=42)
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import KFold,cross_validate
folds=KFold(5,shuffle=True,random_state=42)
model=LinearRegression()

Your task · Follow

Cross-validate the supplied linear model and store the result dictionary.

Hint 1 — Think

Cross-validation returns separate evidence for each fit rather than one final-test score.

Hint 2 — Tools

cross_validate and the negative-RMSE scorer.

Hint 3 — Approach

Evaluate the supplied model on the declared training folds and retain the result dictionary.

Explained solution
answer=cross_validate(model,X_train,y_train,cv=folds,scoring='neg_root_mean_squared_error')

Independent fold fits provide comparable held-out training evidence while leaving the final test unused.

Helpful prior knowledge: What a fold does · A useful reference These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Cross-validate the supplied linear model and store the result dictionary.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.