Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Regression lessonsQUESTIONS · MODELS · EVIDENCE
Several predictors · ML-R06 · 12–18 MIN

Correlated predictors and unstable coefficients

Separate coefficient stability from predictive stability.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeExplain the Python result
  3. TransferExplain the Python result

Understand the idea

Highly correlated inputs can exchange coefficient weight with little change in predictions. This makes individual effects difficult to interpret.

Separate coefficient stability from predictive stability.x₁ ≈ x₂Similar combined signalβ₁ can rise while β₂ falls.β₁x₁ + β₂x₂ can remain nearly unchanged.Stable predictions do not imply stable coefficients.
Schematic · Separate coefficient stability from predictive stability.Scroll the diagram horizontally if needed.

Python skill: Creates a small reproducible perturbation for each observation.

Meet the syntax

rng.normal(0, .001, len(df))
pd.Series(model.coef_, index=X.columns)
rng.normal(0, .001, len(df))
Creates a small reproducible perturbation for each observation.
pd.Series(model.coef_, index=X.columns)
Labels the jointly fitted coefficients, including the near-duplicate input.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

from sklearn.linear_model import LinearRegression
rng=np.random.default_rng(42)
X=pd.DataFrame({'distance':df.distance,'near_copy':df.distance+rng.normal(0,.001,len(df))})
model=LinearRegression().fit(X,df.duration)
answer=pd.Series(model.coef_,index=X.columns)

This practice: Read and run the Python. Next: Change · Correlated predictors and unstable coefficients.

Given data · MIX60

60 observations. One synthetic delivery observation generated for practice. The dataframe df is supplied afresh for each Run.

MIX60 · first 8 prepared rows
distanceweightserviceweekendduration
15.70526.75035standard061.7704
9.338694.81674express135.1694
17.31345.73931economy059.0766
14.257.69699standard156.5785
2.789376.42024express019.5577
19.53685.62508economy177.0289
15.46175.68023standard056.8109
15.93523.17871express152.08

Column meanings and units

Distance, weight and duration use the fixture’s numeric units; no kilometres, kilograms, minutes or other physical units are specified. RMSE is reported in the same synthetic duration units as the target.

These deterministic teaching observations do not describe real deliveries. Service effects and the alternating weekend flag are built into the generated response; they do not establish real-world causal effects.

distancefloat64
Numeric delivery-distance inputUnit / values: Synthetic distance units; physical unit unspecified
weightfloat64
Numeric parcel-weight inputUnit / values: Synthetic weight units; physical unit unspecified
servicestr
Delivery-service categoryUnit / values: standard / express / economy
weekendint64
Binary weekend input, alternating in the fixtureUnit / values: 0 / 1 indicator
durationfloat64
Numeric delivery-duration targetUnit / values: Synthetic duration units; physical unit unspecified

Your task · Follow

Fit a line using distance and a near-copy of distance. Store the two coefficients in answer, labelled by their feature names.

Hint 1 — Think

Nearly duplicate predictors can share the same predictive contribution in unstable proportions.

Hint 2 — Tools

Seeded noise, dataframe construction and LinearRegression.coef_.

Hint 3 — Approach

Create the near-copy column, fit both columns together and label the resulting coefficients.

Explained solution
from sklearn.linear_model import LinearRegression
rng=np.random.default_rng(42)
X=pd.DataFrame({'distance':df.distance,'near_copy':df.distance+rng.normal(0,.001,len(df))})
model=LinearRegression().fit(X,df.duration)
answer=pd.Series(model.coef_,index=X.columns)

The tiny perturbation preserves strong correlation; inspecting both coefficients exposes instability that prediction quality alone can conceal.

Helpful prior knowledge: Read coefficients after encoding These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Fit a line using distance and a near-copy of distance. Store the two coefficients in answer, labelled by their feature names.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.