Understand the idea
Expansion, scaling and Ridge belong together inside CV. Scaling after expansion makes regularisation act on comparable terms. Compare a curve with a line on identical folds.
Python skill: Addresses degree inside the pipeline step named polynomial.
Meet the syntax
GridSearchCV(model, {'polynomial__degree':[2,3]}, cv=folds, scoring='neg_root_mean_squared_error')'polynomial__degree'- Addresses degree inside the pipeline step named polynomial.
[2,3]- Declares the candidate degrees; cross-validation chooses using training evidence.
cv=folds- Uses the same folds for both candidate feature expansions.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import PolynomialFeatures,StandardScaler
from sklearn.linear_model import Ridge
model=Pipeline([('polynomial',PolynomialFeatures(degree=2,include_bias=False)),('scale',StandardScaler()),('model',Ridge())])
answer=model.named_steps['polynomial'].fit_transform(X_train)
This practice: Read and run the Python. Next: Change · Validate polynomial flexibility.
Given data · CURVE48
48 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.
| x | y |
|---|---|
| -3 | 15.9571 |
| -2.87234 | 13.0709 |
| -2.74468 | 14.9362 |
| -2.61702 | 14.45 |
| -2.48936 | 9.39011 |
| -2.3617 | 9.68978 |
| -2.23404 | 11.2101 |
| -2.10638 | 9.96814 |
Column meanings and units
Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.
Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.
| Column | Stored type |
|---|---|
| x | float64 |
| y | float64 |
Supplied setup · available if you need to inspect it
This code runs before your editor on every Run. These are the objects your exercise uses.
from sklearn.model_selection import train_test_split
X=df[['x']]
y=df.y
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42)
Your task · Follow
Create the degree-two polynomial pipeline. Store the transformed training x and x² columns in answer, one row per training example.
Hint 1 — Think
Expansion belongs before scaling and the regularised estimator.
Hint 2 — Tools
Pipeline, PolynomialFeatures, StandardScaler and Ridge.
Hint 3 — Approach
Build the three named steps, then inspect the polynomial step’s training output.
Explained solution
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import PolynomialFeatures,StandardScaler
from sklearn.linear_model import Ridge
model=Pipeline([('polynomial',PolynomialFeatures(degree=2,include_bias=False)),('scale',StandardScaler()),('model',Ridge())])
answer=model.named_steps['polynomial'].fit_transform(X_train)
The dimensions expose the representation entering later preparation; they do not by themselves establish predictive quality.
Helpful prior knowledge: Make curved features · How a search makes a choice These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.