Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Supervised Workflow lessonsQUESTIONS · MODELS · EVIDENCE
Prepare · ML-W02 · 12–18 MIN

A useful reference

Compare with a simple predictor that ignores X.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeAdapt the Python
  3. TransferReason about the Python

Understand the idea

A mean regressor and a most-frequent classifier establish reference behaviour. A feature-based model needs evidence that its additional structure is useful.

Compare with a simple predictor that ignores X.Training outcomesLearn a constantRegression: training meanClassification: most frequent training classEvaluate the reference on the same rows as the candidate.
Schematic · Compare with a simple predictor that ignores X.Scroll the diagram horizontally if needed.

Python skill: Creates a reference that learns only the training-target mean.

Meet the syntax

DummyRegressor(strategy='mean')
DummyClassifier(strategy='most_frequent')
DummyRegressor(strategy='mean')
Creates a reference that learns only the training-target mean.
DummyClassifier(strategy='most_frequent')
Creates a reference that always predicts the most common training class.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

from sklearn.dummy import DummyRegressor
reference=DummyRegressor(strategy='mean').fit(X_train,y_train)
answer=reference.predict(X_test)

This practice: Read and run the Python. Next: Change · A useful reference.

Given data · LINE24

24 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

LINE24 · first 8 prepared rows
distanceduration
15.30501
1.4782611.2361
1.9565211.8488
2.4347814.809
2.9130410.7589
3.391319.4972
3.8695717.89
4.3478315.531

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
distancefloat64
durationfloat64
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.model_selection import train_test_split
X = df[['distance']]
y = df['duration']
X_train, X_test, y_train, y_test = train_test_split(X,y,test_size=.2,random_state=42)

Your task · Follow

Fit a mean reference on training rows and predict the evaluation rows.

Hint 1 — Think

A useful reference must learn its constant from training outcomes only.

Hint 2 — Tools

DummyRegressor with the mean strategy, fit and predict.

Hint 3 — Approach

Fit the reference on the training pair and obtain predictions for the supplied evaluation features.

Explained solution
from sklearn.dummy import DummyRegressor
reference=DummyRegressor(strategy='mean').fit(X_train,y_train)
answer=reference.predict(X_test)

A learned training mean provides a simple comparison without borrowing the evaluation outcomes.

Helpful prior knowledge: Explore training data These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Fit a mean reference on training rows and predict the evaluation rows.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.