Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Supervised Workflow lessonsQUESTIONS · MODELS · EVIDENCE
Select, diagnose and finish · ML-W14 · 18–25 MIN

Respect time

Validate in the direction the model will be used.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeAdapt the Python
  3. PractiseAdapt the Python
  4. TransferExplain the Python result

Understand the idea

Random shuffling can leak future patterns into past training. TimeSeriesSplit uses forward validation. Its validation blocks do not cover every training row, so use the last block for the demonstrated diagnostic.

Validate in the direction the model will be used.Earlier trainingLaterEarlier trainingLaterEarlier trainingLaterTime
Schematic · Validate in the direction the model will be used.Scroll the diagram horizontally if needed.

Python skill: Produces expanding earlier training blocks followed by later validation blocks.

Meet the syntax

TimeSeriesSplit(n_splits=5)
TimeSeriesSplit
Produces expanding earlier training blocks followed by later validation blocks.
n_splits=5
Requests five forward validation blocks; no shuffling is used.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

cut=int(len(df)*.8)
train=df.iloc[:cut]
test=df.iloc[cut:]

This practice: Read and run the Python. Next: Change · Respect time.

Given data · TIME240

240 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

TIME240 · first 8 prepared rows
timehourtemperaturedemand
2026-01-01T00:00:00.000020182.438
2026-01-01T01:00:00.000121.5529178.092
2026-01-01T02:00:00.000223198.404
2026-01-01T03:00:00.000324.2426205.095
2026-01-01T04:00:00.000425.1962185.976
2026-01-01T05:00:00.000525.7956193.765
2026-01-01T06:00:00.000626206.223
2026-01-01T07:00:00.000725.7956202.052

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
timedatetime64[us]
hourint64
temperaturefloat64
demandfloat64

Your task · Follow

Store the first 80% of TIME240 in train and the final 20% in test, keeping time order.

Hint 1 — Think

Later observations must not teach the earlier development period.

Hint 2 — Tools

Positional slicing and time-column extrema.

Hint 3 — Approach

Find the stated chronological cut, create the earlier and later blocks and compare their boundary times.

Explained solution
cut=int(len(df)*.8)
train=df.iloc[:cut]
test=df.iloc[cut:]

Contiguous time slicing preserves the order relevant to a forward-use claim.

Helpful prior knowledge: Finish once, then report These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Store the first 80% of TIME240 in train and the final 20% in test, keeping time order.

traintest

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.