Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Supervised Workflow lessonsQUESTIONS · MODELS · EVIDENCE
Prepare · ML-W-R1 · 15–20 MIN

Preparation retrieval

Retrieval 1

Exercises within this concept

  1. Retrieval 1Retrieve and applyCurrent exercise
  2. Retrieval 2Retrieve and apply
  3. Retrieval 3Retrieve and apply
Retrieval 1 · LINE24_REVIEW

Retrieve without the worked example: Repair full-table scaling: learn the scale from X_train and transform X_test.

Use the new retrieval population shown here.

Retrieve earlier concepts before combining them.

This practice: Retrieve earlier concepts before combining them. Next: Retrieval 2 · Preparation retrieval.

Given data · LINE24_REVIEW

24 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

LINE24_REVIEW · first 8 prepared rows
distanceduration
0.8151325.47563
1.6098310.5409
1.8919912.1471
2.4848114.9638
2.9699311.5503
3.0621219.2776
3.9543418.0319
4.1844516.0477

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
distancefloat64
durationfloat64
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.model_selection import train_test_split
X = df[['distance']]
y = df['duration']
X_train, X_test, y_train, y_test = train_test_split(X,y,test_size=.2,random_state=42)

Supporting concepts: Learn missing-value replacements safely →

Remember the idea

Use the inputs and evidence to recover the method. Hints and explained solutions remain collapsed; exact phrasing is not graded.

Retrieve earlier concepts before combining them.Candidate ACandidate BFitValidateCostCompare matching evidence; smaller error can cost more.
Schematic · Retrieve earlier concepts before combining them.Scroll the diagram horizontally if needed.
Hint 1 — Think

Recall which rows are allowed to define the scale.

Hint 2 — Tools

StandardScaler.fit versus transform.

Hint 3 — Approach

Learn preparation from the development table and reuse it on the reserved table.

Explained solution
from sklearn.preprocessing import StandardScaler
scaler=StandardScaler().fit(X_train)
answer=scaler.transform(X_test)

This restores the information boundary by preventing evaluation rows from contributing means or spreads.

Helpful prior knowledge: Learn missing-value replacements safely These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Retrieval 1

Retrieve without the worked example: Repair full-table scaling: learn the scale from X_train and transform X_test. Use the new retrieval population shown here.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.