Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← ML Foundations lessonsQUESTIONS · MODELS · EVIDENCE
Honest evaluation · ML-F10 · 12–18 MIN

Could we know this at prediction time?

Recognise leakage and context-dependent shortcuts.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeReason about the Python
  3. TransferExplain the Python result

Understand the idea

A feature must exist at the intended prediction time and must not encode the answer. A genuine predictor can still be a shortcut that will fail in another setting.

Recognise leakage and context-dependent shortcuts.Available inputsPrediction timeLater outcomeCould this input be known when the prediction is made?
Schematic · Recognise leakage and context-dependent shortcuts.Scroll the diagram horizontally if needed.

Python skill: A list of columns that would be known when making the prediction.

Meet the syntax

X = df[available_predictors]
available_predictors
A list of columns that would be known when making the prediction.
df[available_predictors]
Selects only those legitimate inputs; a target-derived column must stay out.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

answer=df[['sugarpercent','pricepercent']]

This practice: Read and run the Python. Next: Change · Could we know this at prediction time?

Given data · candy_class

85 observations. One candy product in the survey. The dataframe df is supplied afresh for each Run.

Download source CSV · Source and original dictionary

candy_class · first 8 prepared rows
competitornamechocolatefruitycaramelpeanutyalmondynougatcrispedricewaferhardbarpluribussugarpercentpricepercentwinpercentpopular
100 Grand1010010100.7320.8666.971750% or above
3 Musketeers1000100100.6040.51167.602950% or above
One dime0000000000.0110.11632.2611below 50%
One quarter0000000000.0110.51146.1165below 50%
Air Heads0100000000.9060.51152.341550% or above
Almond Joy1001000100.4650.76750.347550% or above
Baby Ruth1011100100.6040.76756.914550% or above
Boston Baked Beans0001000010.3130.51123.4178below 50%

Column meanings and units

sugarpercent: sugar percentile. pricepercent: price percentile. winpercent: percentage of survey matchups won. Ingredient, bar and multipack fields are 0/1 flags.

Percentile ranks are neither physical sugar percentages nor currency prices. Ratios of percentiles do not measure economic value. Chocolate and fruit flags can overlap; non-chocolate is not synonymous with fruit. The ML class target is defined by winpercent ≥ 50.

Input schema
ColumnStored type
competitornamestr
chocolateint64
fruityint64
caramelint64
peanutyalmondyint64
nougatint64
crispedricewaferint64
hardint64
barint64
pluribusint64
sugarpercentfloat64
pricepercentfloat64
winpercentfloat64
popularstr

Your task · Follow

Candy popular is derived from winpercent. Select sugarpercent and pricepercent only.

Hint 1 — Think

A source used to construct the label can reveal the answer even if its name differs.

Hint 2 — Tools

Allowed feature lists and target-derived leakage.

Hint 3 — Approach

Select only the two permitted pre-outcome measurements and inspect the columns.

Explained solution
answer=df[['sugarpercent','pricepercent']]

winpercent defines popular, so retaining it would let a classifier reconstruct the label instead of learning the intended relationship.

Helpful prior knowledge: Preserving class representation These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Candy popular is derived from winpercent. Select sugarpercent and pricepercent only.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.