Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Regression lessonsQUESTIONS · MODELS · EVIDENCE
Curves and trees · ML-R07 · 12–18 MIN

Make curved features

Understand expansion before fitting.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeAdapt the Python
  3. TransferAdapt the Python

Understand the idea

PolynomialFeatures constructs powers and interactions. The subsequent regression remains linear in those expanded features, while predictions can curve in the original inputs.

Understand expansion before fitting.Curved featuresStraight-line candidateMore flexibility must earn its place in validation.
Schematic · Understand expansion before fitting.Scroll the diagram horizontally if needed.

Python skill: Builds powers and interactions from the original feature columns.

Meet the syntax

PolynomialFeatures(degree=2, include_bias=False)
PolynomialFeatures
Builds powers and interactions from the original feature columns.
degree=2
Includes terms up to degree two.
include_bias=False
Omits the all-ones column because the later estimator can fit its own intercept.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

from sklearn.preprocessing import PolynomialFeatures
expander=PolynomialFeatures(degree=2,include_bias=False)
answer=expander.fit_transform(df[['x']])

This practice: Read and run the Python. Next: Change · Make curved features.

Polynomial regression · from idea to workflow

Question: Does a smooth curve improve quantity prediction?

Mechanism: Expand powers of the inputs, then fit the regularised linear recipe.

Watch for: Higher degree can follow noise; extrapolated curves can diverge rapidly.

Interpret or debug · self-review: Compare degrees using training folds. A lower training error alone is not a reason to increase degree.

Independent full workflow: ML-X02. First practise complete regression and classification in F11–F12; check readiness in W-K2–W-K3.

Given data · CURVE48

48 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

CURVE48 · first 8 prepared rows
xy
-315.9571
-2.8723413.0709
-2.7446814.9362
-2.6170214.45
-2.489369.39011
-2.36179.68978
-2.2340411.2101
-2.106389.96814

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
xfloat64
yfloat64

Your task · Follow

Expand CURVE48 x into x and x squared.

Hint 1 — Think

A curved representation can be built before fitting a linear estimator.

Hint 2 — Tools

PolynomialFeatures, degree and include_bias.

Hint 3 — Approach

Fit the degree-two expansion to the single x column and inspect its transformed columns.

Explained solution
from sklearn.preprocessing import PolynomialFeatures
expander=PolynomialFeatures(degree=2,include_bias=False)
answer=expander.fit_transform(df[['x']])

Excluding the bias avoids a duplicate constant term; the two output columns carry the original input and its square.

Helpful prior knowledge: Read a fitted line · Keep preparation with the estimator These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Expand CURVE48 x into x and x squared.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.