Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Neural Networks lessonsQUESTIONS · MODELS · EVIDENCE
Go Further · ML-N-K1 · 25–35 MIN

Neural regression checkpoint

Retrieval 1

Exercises within this concept

  1. Retrieval 1Retrieve and applyCurrent exercise
Retrieval 1 · Wine600

Complete the Wine600 neural regression workflow with feature/target scaling, width selection, convergence evidence and original-unit errors.

This practice: Complete the Wine600 neural regression workflow with feature/target scaling, width selection, convergence evidence and original-unit errors. Next: Clustering and Discovery. Continue to Clustering and Discovery →

Given data · Wine600

600 observations. One red or white wine sample. The dataframe df is supplied afresh for each Run.

Download source CSV · Source and original dictionary

Wine600 · first 8 prepared rows
fixed acidityvolatile aciditycitric acidresidual sugarchloridesfree sulfur dioxidetotal sulfur dioxidedensitypHsulphatesalcoholqualitywine_type
5.30.320.126.60.043221410.99373.360.610.46white
6.40.30.274.40.055171350.99253.230.4412.26white
6.50.220.322.20.02836920.990763.270.5911.97white
80.240.486.80.047131340.996163.230.7105white
8.90.290.351.90.06725570.9973.181.3610.36red
6.60.210.392.30.041311020.992213.220.5810.97white
8.80.240.542.50.08325570.99833.390.549.25red
5.40.150.322.50.03710510.988783.040.5812.66white

Column meanings and units

quality: ordered sensory score (0–10). alcohol: volume percent. pH: acidity scale. Other chemistry units follow the linked source dictionary.

The score is ordinal but modelled as regression here. ML removes exact duplicate rows before splitting. Chemistry must be available at prediction time; predictive associations do not establish effects of changing an ingredient.

Input schema
ColumnStored type
fixed acidityfloat64
volatile acidityfloat64
citric acidfloat64
residual sugarfloat64
chloridesfloat64
free sulfur dioxidefloat64
total sulfur dioxidefloat64
densityfloat64
pHfloat64
sulphatesfloat64
alcoholfloat64
qualityint64
wine_typestr

Declared validation design

Reserve 20% of the declared population for the final test using split seed 42. Use a random split.

Use the same five training folds, shuffled with seed 42, for reference comparison, candidates and selection. Use out-of-fold training predictions for diagnosis. Select with negative RMSE (larger is better) and compare with a training-mean reference. Report final RMSE in original target units. Open final-test evidence after selection and diagnosis.

Supporting concepts: Neural regression workflow → · Convergence, early stopping and validation →

Remember the idea

This checkpoint combines previously taught skills. Assemble the workflow; help remains available when needed.

Complete the Wine600 neural regression workflow with feature/target scaling, width selection, convergence evidence and original-unit errors.Training XScale numbersEncode categoriesFit estimatorEach fold learns its own preparation.
Schematic · Complete the Wine600 neural regression workflow with feature/target scaling, width selection, convergence evidence and original-unit errors.Scroll the diagram horizontally if needed.

Required Python variables and evidence

Use these names so Check can inspect your workflow. Each meaning is shown beside its name.

Workflow evidence contract
VariableMeaning
XFeature dataframe for the declared population, preserving row indices.
X_testFinal-test feature rows from the declared split.
X_trainTraining feature rows from the declared split.
cv_resultsCandidate validation evidence; comparison workflows use a dataframe indexed by model ID.
final_modelChosen pipeline fitted on training rows, after selection and diagnosis.
final_predictionsUnaltered predictions of final_model on X_test.
final_rmseRoot mean squared error in original target units.
loss_curveTraining loss history of the selected neural regressor.
reference_resultscross_validate result for the dummy reference; test_score contains five scores.
residualsTraining-only actual, predicted and actual-minus-predicted residual evidence.
searchFitted GridSearchCV for the declared parameter comparison.
yTarget series, aligned with X.
y_testFinal-test targets, aligned with X_test.
y_trainTraining targets, aligned with X_train.
Hint 1 — Think

Reconstruct both feature and target preparation inside every neural fit.

Hint 2 — Tools

Pipeline, TransformedTargetRegressor, nested width search, loss curve and original-unit RMSE.

Hint 3 — Approach

Compare against the mean reference, select width on training folds, inspect convergence and residuals, then predict through the fitted wrapper.

Explained solution
from sklearn.base import clone
from sklearn.model_selection import train_test_split, KFold, StratifiedKFold, cross_validate, cross_val_predict, GridSearchCV
from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer, TransformedTargetRegressor
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.dummy import DummyRegressor, DummyClassifier
from sklearn.metrics import root_mean_squared_error, f1_score, accuracy_score, confusion_matrix

from sklearn.neural_network import MLPRegressor
X = df.drop(columns=['quality', 'wine_type'])
y = df['quality']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
model = Pipeline([('prepare', StandardScaler()), ('model', TransformedTargetRegressor(regressor=MLPRegressor(hidden_layer_sizes=(24,), max_iter=800, early_stopping=True, tol=1e-3, random_state=42), transformer=StandardScaler()))])
folds = KFold(n_splits=5, shuffle=True, random_state=42)
cv_results = cross_validate(model, X_train, y_train, cv=folds, scoring='neg_root_mean_squared_error')
reference_results = cross_validate(DummyRegressor(strategy='mean'), X_train, y_train, cv=folds, scoring='neg_root_mean_squared_error')
search = GridSearchCV(model, {'model__regressor__hidden_layer_sizes':[(16,), (24,)]}, cv=folds, scoring='neg_root_mean_squared_error')
search.fit(X_train, y_train)
oof_predictions = cross_val_predict(search.best_estimator_, X_train, y_train, cv=folds)
residuals = y_train - oof_predictions
loss_curve = search.best_estimator_.named_steps['model'].regressor_.loss_curve_
final_model = clone(search.best_estimator_).fit(X_train, y_train)
final_predictions = final_model.predict(X_test)
final_rmse = root_mean_squared_error(y_test, final_predictions)
print('Final RMSE in quality-score units:', final_rmse)

Fold-local target scaling supports optimisation without leakage, and the wrapper reverses it so final errors retain quality-score units.

Helpful prior knowledge: Neural regression workflow · Shared network retrieval These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Retrieval 1

Complete the Wine600 neural regression workflow with feature/target scaling, width selection, convergence evidence and original-unit errors.

searchcv_resultsreference_resultsloss_curveresidualsfinal_rmse

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.

Do original-unit errors and convergence evidence justify the chosen Wine network?

Use your Run output as evidence. This response is optional, not machine-graded or saved.