Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Regression lessonsQUESTIONS · MODELS · EVIDENCE
Curves and trees · ML-R-K1 · 25–35 MIN

Regression checkpoint

Retrieval 1

Exercises within this concept

  1. Retrieval 1Retrieve and applyCurrent exercise
Retrieval 1 · candy

Compare mixed-input linear regression and a tree on Candy with common training folds, a reference, diagnosis and final evidence.

This practice: Compare mixed-input linear regression and a tree on Candy with common training folds, a reference, diagnosis and final evidence. Next: explore Classification or choose another model family.

Given data · candy

85 observations. One candy product in the survey. The dataframe df is supplied afresh for each Run.

Download source CSV · Source and original dictionary

candy · first 8 prepared rows
competitornamechocolatefruitycaramelpeanutyalmondynougatcrispedricewaferhardbarpluribussugarpercentpricepercentwinpercent
100 Grand1010010100.7320.8666.9717
3 Musketeers1000100100.6040.51167.6029
One dime0000000000.0110.11632.2611
One quarter0000000000.0110.51146.1165
Air Heads0100000000.9060.51152.3415
Almond Joy1001000100.4650.76750.3475
Baby Ruth1011100100.6040.76756.9145
Boston Baked Beans0001000010.3130.51123.4178

Column meanings and units

sugarpercent: sugar percentile. pricepercent: price percentile. winpercent: percentage of survey matchups won. Ingredient, bar and multipack fields are 0/1 flags.

Percentile ranks are neither physical sugar percentages nor currency prices. Ratios of percentiles do not measure economic value. Chocolate and fruit flags can overlap; non-chocolate is not synonymous with fruit. The ML class target is defined by winpercent ≥ 50.

Input schema
ColumnStored type
competitornamestr
chocolateint64
fruityint64
caramelint64
peanutyalmondyint64
nougatint64
crispedricewaferint64
hardint64
barint64
pluribusint64
sugarpercentfloat64
pricepercentfloat64
winpercentfloat64

Declared validation design

Reserve 20% of the declared population for the final test using split seed 42. Use a random split.

Use the same five training folds, shuffled with seed 42, for reference comparison, candidates and selection. Use out-of-fold training predictions for diagnosis. Select with negative RMSE (larger is better) and compare with a training-mean reference. Report final RMSE in original target units. Open final-test evidence after selection and diagnosis.

Supporting concepts: Read residual patterns and limits → · Read coefficients after encoding → · Validate polynomial flexibility → · Control tree complexity →

Remember the idea

This checkpoint combines previously taught skills. Assemble the workflow; help remains available when needed.

Compare mixed-input linear regression and a tree on Candy with common training folds, a reference, diagnosis and final evidence.Training XScale numbersEncode categoriesFit estimatorEach fold learns its own preparation.
Schematic · Compare mixed-input linear regression and a tree on Candy with common training folds, a reference, diagnosis and final evidence.Scroll the diagram horizontally if needed.

Required Python variables and evidence

Use these names so Check can inspect your workflow. Each meaning is shown beside its name.

Workflow evidence contract
VariableMeaning
XFeature dataframe for the declared population, preserving row indices.
X_testFinal-test feature rows from the declared split.
X_trainTraining feature rows from the declared split.
chosen_nameModel ID nominated from training evidence.
cv_resultsCandidate validation evidence; comparison workflows use a dataframe indexed by model ID.
dfLoaded and prepared input dataframe.
final_modelChosen pipeline fitted on training rows, after selection and diagnosis.
final_predictionsUnaltered predictions of final_model on X_test.
final_rmseRoot mean squared error in original target units.
reference_resultscross_validate result for the dummy reference; test_score contains five scores.
residualsTraining-only actual, predicted and actual-minus-predicted residual evidence.
selectedDictionary mapping each requested model ID to its chosen pipeline.
yTarget series, aligned with X.
y_testFinal-test targets, aligned with X_test.
y_trainTraining targets, aligned with X_train.
Hint 1 — Think

Reconstruct a fair comparison of two different flexibility assumptions.

Hint 2 — Tools

Common KFold, mixed pipelines, tree-depth search, OOF residuals and final RMSE.

Hint 3 — Approach

Evaluate both candidates and the reference, tune only the declared tree grid, nominate from training evidence, diagnose and finish.

Explained solution
from sklearn.base import clone
from sklearn.model_selection import train_test_split, KFold, cross_validate, cross_val_predict, GridSearchCV
from sklearn.dummy import DummyRegressor
from sklearn.metrics import root_mean_squared_error

import matplotlib.pyplot as plt
X=df[['sugarpercent', 'pricepercent', 'chocolate', 'fruity', 'caramel', 'peanutyalmondy', 'nougat', 'crispedricewafer', 'hard', 'bar', 'pluribus']]
y=df['winpercent']
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42)
folds=KFold(n_splits=5,shuffle=True,random_state=42)
candidates={}
grids={}
# Prepare inputs and construct the candidate pipeline.
preprocessor = 'passthrough'
from sklearn.pipeline import Pipeline
from sklearn.linear_model import LinearRegression
model = LinearRegression()
pipeline = Pipeline([('prepare', preprocessor), ('model', model)])
model = pipeline
candidates['multiple_linear']=model
grids['multiple_linear']=None
# Prepare inputs and construct the candidate pipeline.
preprocessor = 'passthrough'
from sklearn.pipeline import Pipeline
from sklearn.tree import DecisionTreeRegressor
model = DecisionTreeRegressor(random_state=42)
pipeline = Pipeline([('prepare', preprocessor), ('model', model)])
model = pipeline
candidates['regression_tree']=model
grids['regression_tree']={'model__max_depth': [3, 5, None]}
# Validate initial candidates on matching training folds.
initial_results={name:cross_validate(candidate,X_train,y_train,cv=folds,scoring='neg_root_mean_squared_error') for name,candidate in candidates.items()}
# Compare with a simple reference before tuning.
reference_results=cross_validate(DummyRegressor(strategy='mean'),X_train,y_train,cv=folds,scoring='neg_root_mean_squared_error')
rows=[]
selected={}
for name,candidate in candidates.items():
    initial=initial_results[name]
    grid=grids[name]
    if grid:
        search=GridSearchCV(candidate,grid,cv=folds,scoring='neg_root_mean_squared_error',n_jobs=1)
        search.fit(X_train,y_train)
        selected[name]=search.best_estimator_
        validation_score=float(search.best_score_)
        settings=search.best_params_
    else:
        selected[name]=candidate
        validation_score=float(np.mean(initial['test_score']))
        settings='Keep defaults: no production search'
    rows.append({'model':name,'initial_score':float(np.mean(initial['test_score'])),'selected_score':validation_score,'settings':str(settings)})
cv_results=pd.DataFrame(rows).set_index('model')
# This solution nominates by mean training-fold score. Other evidence-based choices can be defensible.
chosen_name=cv_results.selected_score.idxmax()
chosen=selected[chosen_name]
oof_predictions=cross_val_predict(chosen,X_train,y_train,cv=folds)
diagnostic_y=y_train
residuals=pd.DataFrame({'actual':diagnostic_y,'predicted':oof_predictions,'residual':diagnostic_y-oof_predictions})
fig,ax=plt.subplots(figsize=(6,4))
ax.scatter(residuals.predicted,residuals.residual,s=12)
ax.axhline(0,color='black')
ax.set(xlabel='Training-validation prediction',ylabel='Residual',title='Training-only residuals')
fig.tight_layout()
fig.savefig('diagnostic.png',dpi=150,bbox_inches='tight')
final_model=clone(chosen).fit(X_train,y_train)
final_predictions=final_model.predict(X_test)
final_rmse=root_mean_squared_error(y_test,final_predictions)
print('Final RMSE:',final_rmse)

The matched design isolates candidate differences; original-unit residuals qualify the nomination before a single final evaluation of the selected workflow.

Helpful prior knowledge: Regression evidence retrieval · Flexibility retrieval These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Retrieval 1

Compare mixed-input linear regression and a tree on Candy with common training folds, a reference, diagnosis and final evidence.

cv_resultsreference_resultschosen_nameresidualsfinal_rmse

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.

Which Candy model earned selection from common-fold evidence?

Use your Run output as evidence. This response is optional, not machine-graded or saved.