Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Supervised Workflow lessonsQUESTIONS · MODELS · EVIDENCE
Validate and compare · ML-W08 · 12–18 MIN

What a fold does

Separate fitting, validation and final-test row roles.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeReason about the Python
  3. TransferAdapt the Python

Understand the idea

Cross-validation repeatedly holds away part of the training population. The final test is outside this process. Each fold fits a fresh estimator.

Separate fitting, validation and final-test row roles.Fold 1VTTTTFold 2TVTTTFold 3TTVTTFold 4TTTVTFold 5TTTTVFinalT: fit on training rows V: validate Final: held away
Schematic · Separate fitting, validation and final-test row roles.Scroll the diagram horizontally if needed.

Python skill: Produces positional training/validation index pairs; it does not fit a model.

Meet the syntax

KFold(n_splits=5, shuffle=True, random_state=42)
folds.split(X_train)
KFold
Produces positional training/validation index pairs; it does not fit a model.
n_splits=5
Divides the supplied training population into five validation folds.
shuffle=True
Shuffles ordinary unordered observations before partitioning.
random_state=42
Makes the shuffled folds repeatable.
folds.split(X_train)
Yields positional training and validation indices for the supplied development table.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

from sklearn.model_selection import KFold
answer=list(KFold(3,shuffle=True,random_state=42).split(X_train))

This practice: Read and run the Python. Next: Change · What a fold does.

Given data · LINE24

24 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

LINE24 · first 8 prepared rows
distanceduration
15.30501
1.4782611.2361
1.9565211.8488
2.4347814.809
2.9130410.7589
3.391319.4972
3.8695717.89
4.3478315.531

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
distancefloat64
durationfloat64
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.model_selection import train_test_split
X = df[['distance']]
y = df['duration']
X_train, X_test, y_train, y_test = train_test_split(X,y,test_size=.2,random_state=42)

Your task · Follow

Use three shuffled KFold splits with seed 42 on X_train. Store the three (training positions, validation positions) pairs in answer.

Hint 1 — Think

Fold indices describe roles within the training table; they are not new observations.

Hint 2 — Tools

KFold.split and a list of index-array pairs.

Hint 3 — Approach

Construct the requested shuffled three-fold splitter and inspect every pair returned for X_train.

Explained solution
from sklearn.model_selection import KFold
answer=list(KFold(3,shuffle=True,random_state=42).split(X_train))

Each pair separates a fitting subset from a validation subset without involving the reserved test population.

Helpful prior knowledge: Learn missing-value replacements safely These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Use three shuffled KFold splits with seed 42 on X_train. Store the three (training positions, validation positions) pairs in answer.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.