Complete the Wine600 neural regression workflow with feature/target scaling, width selection, convergence evidence and original-unit errors.
This practice: Complete the Wine600 neural regression workflow with feature/target scaling, width selection, convergence evidence and original-unit errors. Next: Clustering and Discovery. Continue to Clustering and Discovery →
Given data · Wine600
600 observations. One red or white wine sample. The dataframe df is supplied afresh for each Run.
Download source CSV · Source and original dictionary
| fixed acidity | volatile acidity | citric acid | residual sugar | chlorides | free sulfur dioxide | total sulfur dioxide | density | pH | sulphates | alcohol | quality | wine_type |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 5.3 | 0.32 | 0.12 | 6.6 | 0.043 | 22 | 141 | 0.9937 | 3.36 | 0.6 | 10.4 | 6 | white |
| 6.4 | 0.3 | 0.27 | 4.4 | 0.055 | 17 | 135 | 0.9925 | 3.23 | 0.44 | 12.2 | 6 | white |
| 6.5 | 0.22 | 0.32 | 2.2 | 0.028 | 36 | 92 | 0.99076 | 3.27 | 0.59 | 11.9 | 7 | white |
| 8 | 0.24 | 0.48 | 6.8 | 0.047 | 13 | 134 | 0.99616 | 3.23 | 0.7 | 10 | 5 | white |
| 8.9 | 0.29 | 0.35 | 1.9 | 0.067 | 25 | 57 | 0.997 | 3.18 | 1.36 | 10.3 | 6 | red |
| 6.6 | 0.21 | 0.39 | 2.3 | 0.041 | 31 | 102 | 0.99221 | 3.22 | 0.58 | 10.9 | 7 | white |
| 8.8 | 0.24 | 0.54 | 2.5 | 0.083 | 25 | 57 | 0.9983 | 3.39 | 0.54 | 9.2 | 5 | red |
| 5.4 | 0.15 | 0.32 | 2.5 | 0.037 | 10 | 51 | 0.98878 | 3.04 | 0.58 | 12.6 | 6 | white |
Column meanings and units
quality: ordered sensory score (0–10). alcohol: volume percent. pH: acidity scale. Other chemistry units follow the linked source dictionary.
The score is ordinal but modelled as regression here. ML removes exact duplicate rows before splitting. Chemistry must be available at prediction time; predictive associations do not establish effects of changing an ingredient.
| Column | Stored type |
|---|---|
| fixed acidity | float64 |
| volatile acidity | float64 |
| citric acid | float64 |
| residual sugar | float64 |
| chlorides | float64 |
| free sulfur dioxide | float64 |
| total sulfur dioxide | float64 |
| density | float64 |
| pH | float64 |
| sulphates | float64 |
| alcohol | float64 |
| quality | int64 |
| wine_type | str |
Declared validation design
Reserve 20% of the declared population for the final test using split seed 42. Use a random split.
Use the same five training folds, shuffled with seed 42, for reference comparison, candidates and selection. Use out-of-fold training predictions for diagnosis. Select with negative RMSE (larger is better) and compare with a training-mean reference. Report final RMSE in original target units. Open final-test evidence after selection and diagnosis.
Supporting concepts: Neural regression workflow → · Convergence, early stopping and validation →
Remember the idea
This checkpoint combines previously taught skills. Assemble the workflow; help remains available when needed.
Required Python variables and evidence
Use these names so Check can inspect your workflow. Each meaning is shown beside its name.
| Variable | Meaning |
|---|---|
| X | Feature dataframe for the declared population, preserving row indices. |
| X_test | Final-test feature rows from the declared split. |
| X_train | Training feature rows from the declared split. |
| cv_results | Candidate validation evidence; comparison workflows use a dataframe indexed by model ID. |
| final_model | Chosen pipeline fitted on training rows, after selection and diagnosis. |
| final_predictions | Unaltered predictions of final_model on X_test. |
| final_rmse | Root mean squared error in original target units. |
| loss_curve | Training loss history of the selected neural regressor. |
| reference_results | cross_validate result for the dummy reference; test_score contains five scores. |
| residuals | Training-only actual, predicted and actual-minus-predicted residual evidence. |
| search | Fitted GridSearchCV for the declared parameter comparison. |
| y | Target series, aligned with X. |
| y_test | Final-test targets, aligned with X_test. |
| y_train | Training targets, aligned with X_train. |
Hint 1 — Think
Reconstruct both feature and target preparation inside every neural fit.
Hint 2 — Tools
Pipeline, TransformedTargetRegressor, nested width search, loss curve and original-unit RMSE.
Hint 3 — Approach
Compare against the mean reference, select width on training folds, inspect convergence and residuals, then predict through the fitted wrapper.
Explained solution
from sklearn.base import clone
from sklearn.model_selection import train_test_split, KFold, StratifiedKFold, cross_validate, cross_val_predict, GridSearchCV
from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer, TransformedTargetRegressor
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.dummy import DummyRegressor, DummyClassifier
from sklearn.metrics import root_mean_squared_error, f1_score, accuracy_score, confusion_matrix
from sklearn.neural_network import MLPRegressor
X = df.drop(columns=['quality', 'wine_type'])
y = df['quality']
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
model = Pipeline([('prepare', StandardScaler()), ('model', TransformedTargetRegressor(regressor=MLPRegressor(hidden_layer_sizes=(24,), max_iter=800, early_stopping=True, tol=1e-3, random_state=42), transformer=StandardScaler()))])
folds = KFold(n_splits=5, shuffle=True, random_state=42)
cv_results = cross_validate(model, X_train, y_train, cv=folds, scoring='neg_root_mean_squared_error')
reference_results = cross_validate(DummyRegressor(strategy='mean'), X_train, y_train, cv=folds, scoring='neg_root_mean_squared_error')
search = GridSearchCV(model, {'model__regressor__hidden_layer_sizes':[(16,), (24,)]}, cv=folds, scoring='neg_root_mean_squared_error')
search.fit(X_train, y_train)
oof_predictions = cross_val_predict(search.best_estimator_, X_train, y_train, cv=folds)
residuals = y_train - oof_predictions
loss_curve = search.best_estimator_.named_steps['model'].regressor_.loss_curve_
final_model = clone(search.best_estimator_).fit(X_train, y_train)
final_predictions = final_model.predict(X_test)
final_rmse = root_mean_squared_error(y_test, final_predictions)
print('Final RMSE in quality-score units:', final_rmse)
Fold-local target scaling supports optimisation without leakage, and the wrapper reverses it so final errors retain quality-score units.
Helpful prior knowledge: Neural regression workflow · Shared network retrieval These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.