Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Supervised Workflow lessonsQUESTIONS · MODELS · EVIDENCE
Prepare · ML-W01 · 12–18 MIN

Explore training data

Keep input exploration inside the training population.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeAdapt the Python
  3. TransferExplain the Python result

Understand the idea

Training summaries help define preparation and expose unusual values. Keep final-test rows out of decisions made from exploratory patterns.

Keep input exploration inside the training population.ObservationsTraining → fit / validateFinal testProtect final rows from preparation and selection.
Schematic · Keep input exploration inside the training population.Scroll the diagram horizontally if needed.

Python skill: Summarises numeric training columns without inspecting the reserved test population.

Meet the syntax

X_train.describe()
X_train.plot(...)
fig, ax = plt.subplots()
ax.scatter(x, y)
ax.set(xlabel="Feature", ylabel="Outcome")
X_train.describe()
Summarises numeric training columns without inspecting the reserved test population.
X_train.plot(...)
Starts a plot from the training table; choose the variables and labels for the question.
fig, ax = plt.subplots()
Creates a figure and axes for the training-data view.
ax.scatter(x, y)
Plots paired values; supply training inputs and their aligned outcomes.
ax.set(xlabel="Feature", ylabel="Outcome")
Labels the plotted quantities so the figure can be interpreted.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

answer=X_train.describe()
import matplotlib.pyplot as plt
fig,ax=plt.subplots()
ax.scatter(X_train.distance,y_train)
ax.set(xlabel='Distance',ylabel='Duration',title='Training delivery observations')

This practice: Read and run the Python. Next: Change · Explore training data.

Given data · MIX60

60 observations. One synthetic delivery observation generated for practice. The dataframe df is supplied afresh for each Run.

MIX60 · first 8 prepared rows
distanceweightserviceweekendduration
15.70526.75035standard061.7704
9.338694.81674express135.1694
17.31345.73931economy059.0766
14.257.69699standard156.5785
2.789376.42024express019.5577
19.53685.62508economy177.0289
15.46175.68023standard056.8109
15.93523.17871express152.08

Column meanings and units

Distance, weight and duration use the fixture’s numeric units; no kilometres, kilograms, minutes or other physical units are specified. RMSE is reported in the same synthetic duration units as the target.

These deterministic teaching observations do not describe real deliveries. Service effects and the alternating weekend flag are built into the generated response; they do not establish real-world causal effects.

distancefloat64
Numeric delivery-distance inputUnit / values: Synthetic distance units; physical unit unspecified
weightfloat64
Numeric parcel-weight inputUnit / values: Synthetic weight units; physical unit unspecified
servicestr
Delivery-service categoryUnit / values: standard / express / economy
weekendint64
Binary weekend input, alternating in the fixtureUnit / values: 0 / 1 indicator
durationfloat64
Numeric delivery-duration targetUnit / values: Synthetic duration units; physical unit unspecified
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.model_selection import train_test_split
X = df[['distance','weight','service','weekend']]
y = df['duration']
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42)

Your task · Follow

Summarise the numeric training inputs and draw distance against training duration.

Hint 1 — Think

Exploration can influence later choices, so use only the development population.

Hint 2 — Tools

describe, scatter, axis labels and X_train/y_train.

Hint 3 — Approach

Summarise numeric training inputs, then plot training distance against its aligned training target.

Explained solution
answer=X_train.describe()
import matplotlib.pyplot as plt
fig,ax=plt.subplots()
ax.scatter(X_train.distance,y_train)
ax.set(xlabel='Distance',ylabel='Duration',title='Training delivery observations')

Both the summary and figure describe rows available for development; the final test stays outside exploratory decisions.

Helpful prior knowledge: ML Foundations checkpoint These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Summarise the numeric training inputs and draw distance against training duration.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.