Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← ML Foundations lessonsQUESTIONS · MODELS · EVIDENCE
Learning from examples · ML-F04 · 12–18 MIN

Predicting new rows

Preserve feature shape and meaning at prediction time.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeAdapt the Python
  3. TransferAdapt the Python

Understand the idea

predict uses fitted parameters. New rows must provide the same feature meaning and order. A one-row dataframe remains two-dimensional.

Preserve feature shape and meaning at prediction time.ObservationFeature X₁Feature X₂Target yrow 12standardknownrow 25expressknownrow 38economyunknownKeep row identities aligned; new rows provide X.
Schematic · Preserve feature shape and meaning at prediction time.Scroll the diagram horizontally if needed.

Python skill: Uses the fitted model without learning new parameters.

Meet the syntax

model.predict(pd.DataFrame({'distance':[4,7]}))
model.predict
Uses the fitted model without learning new parameters.
pd.DataFrame({'distance':[4,7]})
Builds two new rows with the same feature name and two-dimensional shape used in fitting.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

answer=model.predict(pd.DataFrame({'distance':[4,7]}))

This practice: Read and run the Python. Next: Change · Predicting new rows.

Given data · LINE24

24 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

LINE24 · first 8 prepared rows
distanceduration
15.30501
1.4782611.2361
1.9565211.8488
2.4347814.809
2.9130410.7589
3.391319.4972
3.8695717.89
4.3478315.531

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
distancefloat64
durationfloat64
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.model_selection import train_test_split
X = df[['distance']]
y = df['duration']
X_train, X_test, y_train, y_test = train_test_split(X,y,test_size=.2,random_state=42)
from sklearn.linear_model import LinearRegression
model = LinearRegression().fit(X_train,y_train)
predictions = model.predict(X_test)

Your task · Follow

Predict durations for distance 4 and 7; store answer.

Hint 1 — Think

Prediction rows must use the feature name and shape the fitted line expects.

Hint 2 — Tools

pd.DataFrame and model.predict.

Hint 3 — Approach

Build the two requested distance rows, pass them through the supplied fitted model and retain their order.

Explained solution
answer=model.predict(pd.DataFrame({'distance':[4,7]}))

Prediction applies the existing learned line; it does not train a new model on the requested distances.

Helpful prior knowledge: What fitting does These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Predict durations for distance 4 and 7; store answer.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.