Understand the idea
Random shuffling can leak future patterns into past training. TimeSeriesSplit uses forward validation. Its validation blocks do not cover every training row, so use the last block for the demonstrated diagnostic.
Python skill: Produces expanding earlier training blocks followed by later validation blocks.
Meet the syntax
TimeSeriesSplit(n_splits=5)TimeSeriesSplit- Produces expanding earlier training blocks followed by later validation blocks.
n_splits=5- Requests five forward validation blocks; no shuffling is used.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
cut=int(len(df)*.8)
train=df.iloc[:cut]
test=df.iloc[cut:]
This practice: Read and run the Python. Next: Change · Respect time.
Given data · TIME240
240 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.
| time | hour | temperature | demand |
|---|---|---|---|
| 2026-01-01T00:00:00.000 | 0 | 20 | 182.438 |
| 2026-01-01T01:00:00.000 | 1 | 21.5529 | 178.092 |
| 2026-01-01T02:00:00.000 | 2 | 23 | 198.404 |
| 2026-01-01T03:00:00.000 | 3 | 24.2426 | 205.095 |
| 2026-01-01T04:00:00.000 | 4 | 25.1962 | 185.976 |
| 2026-01-01T05:00:00.000 | 5 | 25.7956 | 193.765 |
| 2026-01-01T06:00:00.000 | 6 | 26 | 206.223 |
| 2026-01-01T07:00:00.000 | 7 | 25.7956 | 202.052 |
Column meanings and units
Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.
Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.
| Column | Stored type |
|---|---|
| time | datetime64[us] |
| hour | int64 |
| temperature | float64 |
| demand | float64 |
Your task · Follow
Store the first 80% of TIME240 in train and the final 20% in test, keeping time order.
Hint 1 — Think
Later observations must not teach the earlier development period.
Hint 2 — Tools
Positional slicing and time-column extrema.
Hint 3 — Approach
Find the stated chronological cut, create the earlier and later blocks and compare their boundary times.
Explained solution
cut=int(len(df)*.8)
train=df.iloc[:cut]
test=df.iloc[cut:]
Contiguous time slicing preserves the order relevant to a forward-use claim.
Helpful prior knowledge: Finish once, then report These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.