Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Supervised Workflow lessonsQUESTIONS · MODELS · EVIDENCE
Prepare · ML-W03 · 12–18 MIN

Learn a scale from training rows

Distinguish fitting a scale from transforming with it.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeAdapt the Python
  3. TransferAdapt the Python

Understand the idea

StandardScaler learns means and scales during fit; transform reuses them. In supervised prediction learn them on training rows. In discovery fit them on the declared exploratory population, as U02 explains. You can enter here directly from Foundations.

Distinguish fitting a scale from transforming with it.Raw: unequal unitsStandardisedFit statistics at the boundary appropriate to the question.
Schematic · Distinguish fitting a scale from transforming with it.Scroll the diagram horizontally if needed.

Python skill: Learns a mean and scale from each training feature column.

Meet the syntax

scaler.fit(X_train)
scaler.transform(X_new)
StandardScaler()
scaler.fit(X_train)
Learns a mean and scale from each training feature column.
scaler.transform(X_new)
Applies the already learned scale; it does not refit on the new rows.
StandardScaler()
Creates an unfitted scaler; it learns statistics only when fit is called.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

from sklearn.preprocessing import StandardScaler
scaler=StandardScaler().fit(X_train)
answer=scaler.transform(X_train)

This practice: Read and run the Python. Next: Change · Learn a scale from training rows.

Given data · LINE24

24 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

LINE24 · first 8 prepared rows
distanceduration
15.30501
1.4782611.2361
1.9565211.8488
2.4347814.809
2.9130410.7589
3.391319.4972
3.8695717.89
4.3478315.531

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
distancefloat64
durationfloat64
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.model_selection import train_test_split
X = df[['distance']]
y = df['duration']
X_train, X_test, y_train, y_test = train_test_split(X,y,test_size=.2,random_state=42)

Your task · Follow

Fit a scaler on the supplied training distance column; store its transformed values.

Hint 1 — Think

A scale is learned information, not merely a cosmetic unit conversion.

Hint 2 — Tools

StandardScaler.fit and transform.

Hint 3 — Approach

Fit on the supplied training column, then apply that fitted scaler to those rows.

Explained solution
from sklearn.preprocessing import StandardScaler
scaler=StandardScaler().fit(X_train)
answer=scaler.transform(X_train)

The mean and spread come exclusively from the declared fitting population.

Helpful prior knowledge: ML Foundations checkpoint These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Fit a scaler on the supplied training distance column; store its transformed values.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.