Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← ML Foundations lessonsQUESTIONS · MODELS · EVIDENCE
Honest evaluation · ML-F-R2 · 15–20 MIN

Honest evaluation retrieval

Retrieval 1

Exercises within this concept

  1. Retrieval 1Retrieve and applyCurrent exercise
  2. Retrieval 2Retrieve and apply
  3. Retrieval 3Retrieve and apply
Retrieval 1 · candy_class_REVIEW

Retrieve without the worked example: Candy popular is derived from winpercent.

Select sugarpercent and pricepercent only. Use the new retrieval population shown here.

Retrieve earlier concepts before combining them.

This practice: Retrieve earlier concepts before combining them. Next: Retrieval 2 · Honest evaluation retrieval.

Given data · candy_class_REVIEW

68 observations. One candy product in the survey. The dataframe df is supplied afresh for each Run.

Download source CSV · Source and original dictionary

candy_class_REVIEW · first 8 prepared rows
competitornamechocolatefruitycaramelpeanutyalmondynougatcrispedricewaferhardbarpluribussugarpercentpricepercentwinpercentpopular
Skittles wildberry0100000010.9410.2255.103750% or above
Hershey's Special Dark1000000100.430.91859.236150% or above
Dum Dums0100001000.7320.03439.4606below 50%
Sugar Babies0010000010.9650.76733.4376below 50%
Rolo1010000010.860.8665.716350% or above
Haribo Twin Snakes0100000010.4650.46542.1788below 50%
Pixie Sticks0000000010.0930.02337.7223below 50%
Sour Patch Kids0100000010.0690.11659.86450% or above

Column meanings and units

sugarpercent: sugar percentile. pricepercent: price percentile. winpercent: percentage of survey matchups won. Ingredient, bar and multipack fields are 0/1 flags.

Percentile ranks are neither physical sugar percentages nor currency prices. Ratios of percentiles do not measure economic value. Chocolate and fruit flags can overlap; non-chocolate is not synonymous with fruit. The ML class target is defined by winpercent ≥ 50.

Input schema
ColumnStored type
competitornamestr
chocolateint64
fruityint64
caramelint64
peanutyalmondyint64
nougatint64
crispedricewaferint64
hardint64
barint64
pluribusint64
sugarpercentfloat64
pricepercentfloat64
winpercentfloat64
popularstr

Supporting concepts: Could we know this at prediction time? →

Remember the idea

Use the inputs and evidence to recover the method. Hints and explained solutions remain collapsed; exact phrasing is not graded.

Retrieve earlier concepts before combining them.Candidate ACandidate BFitValidateCostCompare matching evidence; smaller error can cost more.
Schematic · Retrieve earlier concepts before combining them.Scroll the diagram horizontally if needed.
Hint 1 — Think

A renamed target source can still reveal the answer.

Hint 2 — Tools

Feature availability and target-derived leakage.

Hint 3 — Approach

Identify the source used to define popular and exclude it from the permitted measurement table.

Explained solution
answer=df[['sugarpercent','pricepercent']]

The two permitted measurements preserve the intended prediction task; winpercent would reconstruct the target directly.

Helpful prior knowledge: Could we know this at prediction time? These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Retrieval 1

Retrieve without the worked example: Candy popular is derived from winpercent. Select sugarpercent and pricepercent only. Use the new retrieval population shown here.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.