Understand the idea
A feature must exist at the intended prediction time and must not encode the answer. A genuine predictor can still be a shortcut that will fail in another setting.
Python skill: A list of columns that would be known when making the prediction.
Meet the syntax
X = df[available_predictors]available_predictors- A list of columns that would be known when making the prediction.
df[available_predictors]- Selects only those legitimate inputs; a target-derived column must stay out.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
answer=df[['sugarpercent','pricepercent']]
This practice: Read and run the Python. Next: Change · Could we know this at prediction time?
Given data · candy_class
85 observations. One candy product in the survey. The dataframe df is supplied afresh for each Run.
Download source CSV · Source and original dictionary
| competitorname | chocolate | fruity | caramel | peanutyalmondy | nougat | crispedricewafer | hard | bar | pluribus | sugarpercent | pricepercent | winpercent | popular |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 100 Grand | 1 | 0 | 1 | 0 | 0 | 1 | 0 | 1 | 0 | 0.732 | 0.86 | 66.9717 | 50% or above |
| 3 Musketeers | 1 | 0 | 0 | 0 | 1 | 0 | 0 | 1 | 0 | 0.604 | 0.511 | 67.6029 | 50% or above |
| One dime | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.011 | 0.116 | 32.2611 | below 50% |
| One quarter | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.011 | 0.511 | 46.1165 | below 50% |
| Air Heads | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.906 | 0.511 | 52.3415 | 50% or above |
| Almond Joy | 1 | 0 | 0 | 1 | 0 | 0 | 0 | 1 | 0 | 0.465 | 0.767 | 50.3475 | 50% or above |
| Baby Ruth | 1 | 0 | 1 | 1 | 1 | 0 | 0 | 1 | 0 | 0.604 | 0.767 | 56.9145 | 50% or above |
| Boston Baked Beans | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 1 | 0.313 | 0.511 | 23.4178 | below 50% |
Column meanings and units
sugarpercent: sugar percentile. pricepercent: price percentile. winpercent: percentage of survey matchups won. Ingredient, bar and multipack fields are 0/1 flags.
Percentile ranks are neither physical sugar percentages nor currency prices. Ratios of percentiles do not measure economic value. Chocolate and fruit flags can overlap; non-chocolate is not synonymous with fruit. The ML class target is defined by winpercent ≥ 50.
| Column | Stored type |
|---|---|
| competitorname | str |
| chocolate | int64 |
| fruity | int64 |
| caramel | int64 |
| peanutyalmondy | int64 |
| nougat | int64 |
| crispedricewafer | int64 |
| hard | int64 |
| bar | int64 |
| pluribus | int64 |
| sugarpercent | float64 |
| pricepercent | float64 |
| winpercent | float64 |
| popular | str |
Your task · Follow
Candy popular is derived from winpercent. Select sugarpercent and pricepercent only.
Hint 1 — Think
A source used to construct the label can reveal the answer even if its name differs.
Hint 2 — Tools
Allowed feature lists and target-derived leakage.
Hint 3 — Approach
Select only the two permitted pre-outcome measurements and inspect the columns.
Explained solution
answer=df[['sugarpercent','pricepercent']]
winpercent defines popular, so retaining it would let a classifier reconstruct the label instead of learning the intended relationship.
Helpful prior knowledge: Preserving class representation These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.