Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Choose and Explain Models lessonsQUESTIONS · MODELS · EVIDENCE
Frame and shortlist · ML-M01 · 18–25 MIN

Choose the task before the family

Build a complete model-choice mental model.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeExplain the Python result
  3. PractiseReason about the Python
  4. TransferExplain the Python result

Understand the idea

Begin with the question: predict a quantity, predict a label, discover groups or reduce a representation. Then consider sample size, feature types, scale, flexibility, interpretability, probabilities, assumptions and cost. Core is essential within a relevant pathway, not a requirement to finish every model.

Build a complete model-choice mental model.QuestionPredict quantityPredict labelDiscover groupsReduce dimensions
Schematic · Build a complete model-choice mental model.Scroll the diagram horizontally if needed.

Python skill: Use a dictionary and comprehension to organise estimators that answer the same prediction question.

Meet the syntax

candidates.items()
type(model).__name__
candidates.items()
Iterates over each candidate name and estimator together.
type(model).__name__
Reads the estimator class name for the comparison inventory.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

from sklearn.linear_model import LinearRegression
from sklearn.tree import DecisionTreeRegressor
candidates = {'line': LinearRegression(), 'tree': DecisionTreeRegressor(random_state=42)}
answer = {name: type(model).__name__ for name, model in candidates.items()}

This practice: Read and run the Python. Next: Change · Choose the task before the family.

Given data · LINE24

24 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

LINE24 · first 8 prepared rows
distanceduration
15.30501
1.4782611.2361
1.9565211.8488
2.4347814.809
2.9130410.7589
3.391319.4972
3.8695717.89
4.3478315.531

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
distancefloat64
durationfloat64

Your task · Follow

Create candidate estimators for a numeric target: a line under key line and a regression tree under key tree. Store their class names in a dictionary named answer with those same keys.

Hint 1 — Think

Use a dictionary and comprehension to organise estimators that answer the same prediction question.

Hint 2 — Tools

Use candidates.items(), type(model).__name__. Read the visible syntax meanings before editing.

Hint 3 — Approach

Create candidate estimators for a numeric target: a line under key line and a regression tree under key tree. Store their class names in a dictionary named answer with those same keys. Keep the supplied row order and inspect the named output after running.

Explained solution
from sklearn.linear_model import LinearRegression
from sklearn.tree import DecisionTreeRegressor
candidates = {'line': LinearRegression(), 'tree': DecisionTreeRegressor(random_state=42)}
answer = {name: type(model).__name__ for name, model in candidates.items()}

Both candidates predict a numeric outcome. The dictionary keeps their names attached to their estimator objects. Neither has been fitted or judged. Frame the task first, then compare plausible candidates using common folds and a suitable metric.

Helpful prior knowledge: Regression checkpoint · Classification checkpoint · Neural regression checkpoint · Neural classification checkpoint · Discovery checkpoint · PCA checkpoint These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Create candidate estimators for a numeric target: a line under key line and a regression tree under key tree. Store their class names in a dictionary named answer with those same keys.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.