Understand the idea
Threshold changes trade false positives against false negatives. The default need not match a particular use case. Choose with training-only predictions, and assess uncertainty and consequences.
Python skill: Requests held-out probability vectors rather than class labels from cross_val_predict.
Meet the syntax
cross_val_predict(model, X_train, y_train, method='predict_proba', cv=folds)method='predict_proba'- Requests held-out probability vectors rather than class labels from cross_val_predict.
cv=folds- Keeps threshold investigation inside the training population.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
from sklearn.model_selection import cross_val_predict,StratifiedKFold
answer=cross_val_predict(model,X_train,y_train,cv=StratifiedKFold(5,shuffle=True,random_state=42),method='predict_proba')
This practice: Read and run the Python. Next: Change · Threshold choices depend on costs.
Given data · CLASS180
180 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.
| length | width | label |
|---|---|---|
| 236.033 | -0.805568 | A |
| 581.297 | 0.728558 | A |
| -1511.27 | -1.00866 | A |
| 99.0248 | -0.24496 | A |
| -13.0141 | -0.660765 | A |
| 681.179 | 0.602475 | A |
| 51.1472 | 0.873157 | A |
| 362.131 | -0.665605 | A |
Column meanings and units
Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.
Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.
| Column | Stored type |
|---|---|
| length | float64 |
| width | float64 |
| label | str |
Supplied setup · available if you need to inspect it
This code runs before your editor on every Run. These are the objects your exercise uses.
from sklearn.model_selection import train_test_split
X=df[['length','width']]
y=df.label
X_train,X_test,y_train,y_test=train_test_split(X,y,test_size=.2,random_state=42,stratify=y)
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
model=Pipeline([('scale',StandardScaler()),('model',LogisticRegression(max_iter=2000,random_state=42))]).fit(X_train,y_train)
Your task · Follow
Compute OOF probabilities for the supplied class model.
Hint 1 — Think
Threshold development needs probabilities for rows not used in their own fit.
Hint 2 — Tools
cross_val_predict, StratifiedKFold and method='predict_proba'.
Hint 3 — Approach
Generate class-probability vectors for every training row using the specified stratified folds.
Explained solution
from sklearn.model_selection import cross_val_predict,StratifiedKFold
answer=cross_val_predict(model,X_train,y_train,cv=StratifiedKFold(5,shuffle=True,random_state=42),method='predict_proba')
OOF probabilities permit training-only threshold analysis without opening the reserved final population.
Helpful prior knowledge: A logistic workflow These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.