Compare mixed-input linear regression and a tree on Candy with common training folds, a reference, diagnosis and final evidence.
This practice: Compare mixed-input linear regression and a tree on Candy with common training folds, a reference, diagnosis and final evidence. Next: explore Classification or choose another model family.
Given data · candy
85 observations. One candy product in the survey. The dataframe df is supplied afresh for each Run.
Download source CSV · Source and original dictionary
| competitorname | chocolate | fruity | caramel | peanutyalmondy | nougat | crispedricewafer | hard | bar | pluribus | sugarpercent | pricepercent | winpercent |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 100 Grand | 1 | 0 | 1 | 0 | 0 | 1 | 0 | 1 | 0 | 0.732 | 0.86 | 66.9717 |
| 3 Musketeers | 1 | 0 | 0 | 0 | 1 | 0 | 0 | 1 | 0 | 0.604 | 0.511 | 67.6029 |
| One dime | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.011 | 0.116 | 32.2611 |
| One quarter | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.011 | 0.511 | 46.1165 |
| Air Heads | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0.906 | 0.511 | 52.3415 |
| Almond Joy | 1 | 0 | 0 | 1 | 0 | 0 | 0 | 1 | 0 | 0.465 | 0.767 | 50.3475 |
| Baby Ruth | 1 | 0 | 1 | 1 | 1 | 0 | 0 | 1 | 0 | 0.604 | 0.767 | 56.9145 |
| Boston Baked Beans | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 1 | 0.313 | 0.511 | 23.4178 |
Column meanings and units
sugarpercent: sugar percentile. pricepercent: price percentile. winpercent: percentage of survey matchups won. Ingredient, bar and multipack fields are 0/1 flags.
Percentile ranks are neither physical sugar percentages nor currency prices. Ratios of percentiles do not measure economic value. Chocolate and fruit flags can overlap; non-chocolate is not synonymous with fruit. The ML class target is defined by winpercent ≥ 50.
| Column | Stored type |
|---|---|
| competitorname | str |
| chocolate | int64 |
| fruity | int64 |
| caramel | int64 |
| peanutyalmondy | int64 |
| nougat | int64 |
| crispedricewafer | int64 |
| hard | int64 |
| bar | int64 |
| pluribus | int64 |
| sugarpercent | float64 |
| pricepercent | float64 |
| winpercent | float64 |
Declared validation design
Reserve 20% of the declared population for the final test using split seed 42. Use a random split.
Use the same five training folds, shuffled with seed 42, for reference comparison, candidates and selection. Use out-of-fold training predictions for diagnosis. Select with negative RMSE (larger is better) and compare with a training-mean reference. Report final RMSE in original target units. Open final-test evidence after selection and diagnosis.
Supporting concepts: Read residual patterns and limits → · Read coefficients after encoding → · Validate polynomial flexibility → · Control tree complexity →
Remember the idea
This checkpoint combines previously taught skills. Assemble the workflow; help remains available when needed.
Required Python variables and evidence
Use these names so Check can inspect your workflow. Each meaning is shown beside its name.
| Variable | Meaning |
|---|---|
| X | Feature dataframe for the declared population, preserving row indices. |
| X_test | Final-test feature rows from the declared split. |
| X_train | Training feature rows from the declared split. |
| chosen_name | Model ID nominated from training evidence. |
| cv_results | Candidate validation evidence; comparison workflows use a dataframe indexed by model ID. |
| df | Loaded and prepared input dataframe. |
| final_model | Chosen pipeline fitted on training rows, after selection and diagnosis. |
| final_predictions | Unaltered predictions of final_model on X_test. |
| final_rmse | Root mean squared error in original target units. |
| reference_results | cross_validate result for the dummy reference; test_score contains five scores. |
| residuals | Training-only actual, predicted and actual-minus-predicted residual evidence. |
| selected | Dictionary mapping each requested model ID to its chosen pipeline. |
| y | Target series, aligned with X. |
| y_test | Final-test targets, aligned with X_test. |
| y_train | Training targets, aligned with X_train. |
Hint 1 — Think
Reconstruct a fair comparison of two different flexibility assumptions.
Hint 2 — Tools
Common KFold, mixed pipelines, tree-depth search, OOF residuals and final RMSE.
Hint 3 — Approach
Evaluate both candidates and the reference, tune only the declared tree grid, nominate from training evidence, diagnose and finish.
Explained solution
from sklearn.base import clone
from sklearn.model_selection import train_test_split, KFold, cross_validate, cross_val_predict, GridSearchCV
from sklearn.dummy import DummyRegressor
from sklearn.metrics import root_mean_squared_error
import matplotlib.pyplot as plt
X=df[['sugarpercent', 'pricepercent', 'chocolate', 'fruity', 'caramel', 'peanutyalmondy', 'nougat', 'crispedricewafer', 'hard', 'bar', 'pluribus']]
y=df['winpercent']
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42)
folds=KFold(n_splits=5,shuffle=True,random_state=42)
candidates={}
grids={}
# Prepare inputs and construct the candidate pipeline.
preprocessor = 'passthrough'
from sklearn.pipeline import Pipeline
from sklearn.linear_model import LinearRegression
model = LinearRegression()
pipeline = Pipeline([('prepare', preprocessor), ('model', model)])
model = pipeline
candidates['multiple_linear']=model
grids['multiple_linear']=None
# Prepare inputs and construct the candidate pipeline.
preprocessor = 'passthrough'
from sklearn.pipeline import Pipeline
from sklearn.tree import DecisionTreeRegressor
model = DecisionTreeRegressor(random_state=42)
pipeline = Pipeline([('prepare', preprocessor), ('model', model)])
model = pipeline
candidates['regression_tree']=model
grids['regression_tree']={'model__max_depth': [3, 5, None]}
# Validate initial candidates on matching training folds.
initial_results={name:cross_validate(candidate,X_train,y_train,cv=folds,scoring='neg_root_mean_squared_error') for name,candidate in candidates.items()}
# Compare with a simple reference before tuning.
reference_results=cross_validate(DummyRegressor(strategy='mean'),X_train,y_train,cv=folds,scoring='neg_root_mean_squared_error')
rows=[]
selected={}
for name,candidate in candidates.items():
initial=initial_results[name]
grid=grids[name]
if grid:
search=GridSearchCV(candidate,grid,cv=folds,scoring='neg_root_mean_squared_error',n_jobs=1)
search.fit(X_train,y_train)
selected[name]=search.best_estimator_
validation_score=float(search.best_score_)
settings=search.best_params_
else:
selected[name]=candidate
validation_score=float(np.mean(initial['test_score']))
settings='Keep defaults: no production search'
rows.append({'model':name,'initial_score':float(np.mean(initial['test_score'])),'selected_score':validation_score,'settings':str(settings)})
cv_results=pd.DataFrame(rows).set_index('model')
# This solution nominates by mean training-fold score. Other evidence-based choices can be defensible.
chosen_name=cv_results.selected_score.idxmax()
chosen=selected[chosen_name]
oof_predictions=cross_val_predict(chosen,X_train,y_train,cv=folds)
diagnostic_y=y_train
residuals=pd.DataFrame({'actual':diagnostic_y,'predicted':oof_predictions,'residual':diagnostic_y-oof_predictions})
fig,ax=plt.subplots(figsize=(6,4))
ax.scatter(residuals.predicted,residuals.residual,s=12)
ax.axhline(0,color='black')
ax.set(xlabel='Training-validation prediction',ylabel='Residual',title='Training-only residuals')
fig.tight_layout()
fig.savefig('diagnostic.png',dpi=150,bbox_inches='tight')
final_model=clone(chosen).fit(X_train,y_train)
final_predictions=final_model.predict(X_test)
final_rmse=root_mean_squared_error(y_test,final_predictions)
print('Final RMSE:',final_rmse)
The matched design isolates candidate differences; original-unit residuals qualify the nomination before a single final evaluation of the selected workflow.
Helpful prior knowledge: Regression evidence retrieval · Flexibility retrieval These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.