Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← ML Foundations lessonsQUESTIONS · MODELS · EVIDENCE
Honest evaluation · ML-F07 · 12–18 MIN

Making a reproducible split

Protect evaluation rows before fitting.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeAdapt the Python
  3. TransferAdapt the Python

Understand the idea

A random state makes this split reproducible. It does not make a poor study design valid. Split before learning preprocessing statistics or model parameters.

Protect evaluation rows before fitting.ObservationsTraining → fit / validateFinal testProtect final rows from preparation and selection.
Schematic · Protect evaluation rows before fitting.Scroll the diagram horizontally if needed.

Python skill: Splits features and targets together so row alignment is preserved.

Meet the syntax

train_test_split(X, y, test_size=0.2, random_state=42)
train_test_split(X, y
Splits features and targets together so row alignment is preserved.
test_size=0.2
Reserves 20% of rows for the test partition.
random_state=42
Makes this random partition reproducible; it does not guarantee a representative sample.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

from sklearn.model_selection import train_test_split
X = df[['distance']]
y = df['duration']
X_train, X_test, y_train, y_test = train_test_split(X,y,test_size=.2,random_state=42)

This practice: Read and run the Python. Next: Change · Making a reproducible split.

Given data · LINE24

24 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

LINE24 · first 8 prepared rows
distanceduration
15.30501
1.4782611.2361
1.9565211.8488
2.4347814.809
2.9130410.7589
3.391319.4972
3.8695717.89
4.3478315.531

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
distancefloat64
durationfloat64

Your task · Follow

Split distance/duration into 80% training and 20% test with seed 42.

Hint 1 — Think

Splitting X and y separately risks breaking which outcome belongs to which row.

Hint 2 — Tools

train_test_split, test_size and random_state.

Hint 3 — Approach

Select the legitimate input and target, split them in one call and retain all four returned objects.

Explained solution
from sklearn.model_selection import train_test_split
X = df[['distance']]
y = df['duration']
X_train, X_test, y_train, y_test = train_test_split(X,y,test_size=.2,random_state=42)

A joint split preserves alignment and a fixed seed makes the declared evaluation boundary reproducible.

Helpful prior knowledge: Seen is not unseen These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Split distance/duration into 80% training and 20% test with seed 42.

X_trainX_test

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.