Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← ML Foundations lessonsQUESTIONS · MODELS · EVIDENCE
Questions and tables · ML-F01 · 12–18 MIN

What question are we answering?

Distinguish prediction, grouping and reduction.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeExplain the Python result
  3. TransferExplain the Python result

Understand the idea

Supervised learning uses examples with a known target. Regression predicts quantities; classification predicts labels. Clustering describes groups without a target. PCA creates a lower-dimensional representation.

Distinguish prediction, grouping and reduction.QuestionPredict quantityPredict labelDiscover groupsReduce dimensions
Schematic · Distinguish prediction, grouping and reduction.Scroll the diagram horizontally if needed.

Python skill: Use assignment, a string and a dataframe column to turn a prediction question into Python.

Meet the syntax

target_name
'duration'
df[target_name]
.head()
target_name
A variable is a name for a value; = assigns the string on its right.
'duration'
Quotes make a string: the exact column name holding the quantity to predict.
df[target_name]
Square brackets select the column named by this variable.
.head()
A dot accesses a method; parentheses call it. head shows the first five rows.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

target_name = 'duration'
answer = df[target_name].head()

This practice: Read and run the Python. Next: Change · What question are we answering?

Given data · LINE24

24 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

LINE24 · first 8 prepared rows
distanceduration
15.30501
1.4782611.2361
1.9565211.8488
2.4347814.809
2.9130410.7589
3.391319.4972
3.8695717.89
4.3478315.531

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
distancefloat64
durationfloat64

Your task · Follow

Store the target column name duration in target_name. Select its first five values into answer.

Hint 1 — Think

Use assignment, a string and a dataframe column to turn a prediction question into Python.

Hint 2 — Tools

Use target_name, 'duration', df[target_name], .head(). Read the visible syntax meanings before editing.

Hint 3 — Approach

Store the target column name duration in target_name. Select its first five values into answer. Keep the supplied row order and inspect the named output after running.

Explained solution
target_name = 'duration'
answer = df[target_name].head()

The question is to predict delivery duration. target_name stores the column label, not the values. Selecting df[target_name] returns the observed outcomes. Duration is a quantity, so this is regression. A species label would instead make it classification. Clustering and PCA do not use a prediction target.

Helpful prior knowledge: Data Foundations: inspecting, preparing and plotting tables. These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Store the target column name duration in target_name. Select its first five values into answer.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.