Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Supervised Workflow lessonsQUESTIONS · MODELS · EVIDENCE
Select, diagnose and finish · ML-W-K1 · 25–35 MIN

Supervised Workflow checkpoint

Retrieval 1

Exercises within this concept

  1. Retrieval 1Retrieve and applyCurrent exercise
Retrieval 1 · MIX60

Build a mixed LinearRegression pipeline, compare a mean reference with five-fold CV, create OOF residuals, keep defaults and report final RMSE.

No tree tuning is required.

Next: Readiness · unfamiliar regression.

Given data · MIX60

60 observations. One synthetic delivery observation generated for practice. The dataframe df is supplied afresh for each Run.

MIX60 · first 8 prepared rows
distanceweightserviceweekendduration
15.70526.75035standard061.7704
9.338694.81674express135.1694
17.31345.73931economy059.0766
14.257.69699standard156.5785
2.789376.42024express019.5577
19.53685.62508economy177.0289
15.46175.68023standard056.8109
15.93523.17871express152.08

Column meanings and units

Distance, weight and duration use the fixture’s numeric units; no kilometres, kilograms, minutes or other physical units are specified. RMSE is reported in the same synthetic duration units as the target.

These deterministic teaching observations do not describe real deliveries. Service effects and the alternating weekend flag are built into the generated response; they do not establish real-world causal effects.

distancefloat64
Numeric delivery-distance inputUnit / values: Synthetic distance units; physical unit unspecified
weightfloat64
Numeric parcel-weight inputUnit / values: Synthetic weight units; physical unit unspecified
servicestr
Delivery-service categoryUnit / values: standard / express / economy
weekendint64
Binary weekend input, alternating in the fixtureUnit / values: 0 / 1 indicator
durationfloat64
Numeric delivery-duration targetUnit / values: Synthetic duration units; physical unit unspecified

Declared validation design

Reserve 20% of the declared population for the final test using split seed 42. Use a random split.

Use the same five training folds, shuffled with seed 42, for reference comparison, candidates and selection. Use out-of-fold training predictions for diagnosis. Select with negative RMSE (larger is better) and compare with a training-mean reference. Report final RMSE in original target units. Open final-test evidence after selection and diagnosis.

Supporting concepts: Keep preparation with the estimator → · Read cross-validation evidence → · Finish once, then report →

Remember the idea

This checkpoint combines previously taught skills. Assemble the workflow; help remains available when needed.

Mixed inputs → matching training folds → final RMSE after diagnosis.Mixed inputsShared foldsFinal RMSECompare the mean reference; diagnose training-only residuals.Evaluate the fixed linear recipe on reserved rows once.
Schematic · Mixed inputs → matching training folds → final RMSE after diagnosis.Scroll the diagram horizontally if needed.

Required Python variables and evidence

Use these names so Check can inspect your workflow. Each meaning is shown beside its name.

Workflow evidence contract
VariableMeaning
XFeature dataframe for the declared population, preserving row indices.
X_testFinal-test feature rows from the declared split.
X_trainTraining feature rows from the declared split.
cv_resultsCandidate validation evidence; comparison workflows use a dataframe indexed by model ID.
final_modelChosen pipeline fitted on training rows, after selection and diagnosis.
final_predictionsUnaltered predictions of final_model on X_test.
final_rmseRoot mean squared error in original target units.
reference_resultscross_validate result for the dummy reference; test_score contains five scores.
residualsTraining-only actual, predicted and actual-minus-predicted residual evidence.
yTarget series, aligned with X.
y_testFinal-test targets, aligned with X_test.
y_trainTraining targets, aligned with X_train.
Hint 1 — Think

Reconstruct the workflow around information boundaries, not around a parameter search.

Hint 2 — Tools

Mixed ColumnTransformer/Pipeline, KFold, dummy CV, OOF predictions and RMSE.

Hint 3 — Approach

Prepare within fits, compare the default line and reference, diagnose training-only residuals, then fit and evaluate the fixed recipe.

Explained solution
from sklearn.base import clone
from sklearn.model_selection import train_test_split, KFold, StratifiedKFold, cross_validate, cross_val_predict, GridSearchCV
from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer, TransformedTargetRegressor
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.dummy import DummyRegressor, DummyClassifier
from sklearn.metrics import root_mean_squared_error, f1_score, accuracy_score, confusion_matrix

from sklearn.linear_model import LinearRegression
X = df[['distance', 'weight', 'service', 'weekend']]
y = df['duration']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
prepare = ColumnTransformer([
    ('numeric', 'passthrough', ['distance', 'weight']),
    ('category', OneHotEncoder(handle_unknown='ignore', sparse_output=False, drop='first'), ['service']),
    ('flags', 'passthrough', ['weekend'])
])
model = Pipeline([('prepare', prepare), ('model', LinearRegression())])
folds = KFold(n_splits=5, shuffle=True, random_state=42)
cv_results = cross_validate(model, X_train, y_train, cv=folds, scoring='neg_root_mean_squared_error')
reference_results = cross_validate(DummyRegressor(strategy='mean'), X_train, y_train, cv=folds, scoring='neg_root_mean_squared_error')
# LinearRegression has no parameter search here: keep its defaults.
oof_predictions = cross_val_predict(model, X_train, y_train, cv=folds)
residuals = pd.DataFrame({'actual': y_train, 'predicted': oof_predictions, 'residual': y_train-oof_predictions})
final_model = clone(model).fit(X_train, y_train)
final_predictions = final_model.predict(X_test)
final_rmse = root_mean_squared_error(y_test, final_predictions)
print('Final RMSE:', final_rmse)

The linear checkpoint exercises the full shared workflow without requiring tree tuning. Common folds and OOF diagnosis support development while final rows remain reserved until the recipe is fixed.

Helpful prior knowledge: Preparation retrieval · Workflow evidence retrieval These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Retrieval 1

Build a mixed LinearRegression pipeline, compare a mean reference with five-fold CV, create OOF residuals, keep defaults and report final RMSE. No tree tuning is required.

cv_resultsreference_resultsresidualsfinal_predictionsfinal_rmse

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.

Does the mixed-input line beat the mean reference on matching training folds?

Use your Run output as evidence. This response is optional, not machine-graded or saved.