Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Supervised Workflow lessonsQUESTIONS · MODELS · EVIDENCE
Select, diagnose and finish · ML-W12 · 12–18 MIN

Diagnose without opening the final test

Use predictions from models that did not fit each diagnostic row.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeAdapt the Python
  3. TransferAdapt the Python

Understand the idea

Out-of-fold predictions allow training-only diagnostic plots. After parameter selection they remain development evidence, not an unbiased substitute for final evaluation.

Use predictions from models that did not fit each diagnostic row.actual −predictedResiduals measure signed vertical differences.
Schematic · Use predictions from models that did not fit each diagnostic row.Scroll the diagram horizontally if needed.

Python skill: Returns one held-out prediction for each training row, restoring the original row order.

Meet the syntax

cross_val_predict(model, X_train, y_train, cv=folds)
cross_val_predict
Returns one held-out prediction for each training row, restoring the original row order.
cv=folds
Ensures each prediction comes from a fit that excluded that row.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

from sklearn.model_selection import cross_val_predict
answer=cross_val_predict(model,X_train,y_train,cv=folds)

This practice: Read and run the Python. Next: Change · Diagnose without opening the final test.

Given data · LINE24

24 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

LINE24 · first 8 prepared rows
distanceduration
15.30501
1.4782611.2361
1.9565211.8488
2.4347814.809
2.9130410.7589
3.391319.4972
3.8695717.89
4.3478315.531

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
distancefloat64
durationfloat64
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.model_selection import train_test_split
X = df[['distance']]
y = df['duration']
X_train, X_test, y_train, y_test = train_test_split(X,y,test_size=.2,random_state=42)
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import KFold,cross_validate
folds=KFold(5,shuffle=True,random_state=42)
model=LinearRegression()

Your task · Follow

Create out-of-fold predictions for every training row.

Hint 1 — Think

Every diagnostic training prediction should come from a model that did not fit that row.

Hint 2 — Tools

cross_val_predict with the declared folds.

Hint 3 — Approach

Generate held-out predictions for the training population and preserve their returned order.

Explained solution
from sklearn.model_selection import cross_val_predict
answer=cross_val_predict(model,X_train,y_train,cv=folds)

Out-of-fold predictions support training-only diagnosis without reusing in-sample fitted predictions or opening the final test.

Helpful prior knowledge: How a search makes a choice These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Create out-of-fold predictions for every training row.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.