Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← ML Foundations lessonsQUESTIONS · MODELS · EVIDENCE
First complete workflows · ML-F12 · 25–40 MIN

Your first complete classification workflow

Reuse the workflow to predict labels and inspect class errors.

Exercises within this concept

  1. FollowRun, predict and explainCurrent exercise
  2. ChangeRebuild key steps and justify
  3. TransferBuild from the eight-step outline

Understand the idea

Reuse the same eight steps for class labels. The scaler is inside the Pipeline, so each training fold learns its own scale. Macro F1 averages one precision-and-recall score per class; balanced accuracy averages each class’s recall. Compare the classifier with a majority-class dummy before claiming useful improvement.

Classify a new specimen from measurements available before its label is confirmed.Question / X, ySplit: reserve testTraining-only foldsDummy + candidateChoose + diagnoseFinal predict / explainFit preparation inside folds. Open final evidence once.
Schematic · Classify a new specimen from measurements available before its label is confirmed.Scroll the diagram horizontally if needed.

Python skill: Reuse a Pipeline with a classifier and choose an sklearn scoring string.

Meet the syntax

Pipeline([('prepare', preparation), ('model', estimator)])
cross_validate(candidate, X_train, y_train, cv=folds, scoring=metric)
cross_val_predict(candidate, X_train, y_train, cv=folds)
clone(candidate).fit(X_train, y_train)
validation['test_score']
f1_macro
balanced_accuracy
Pipeline
Keeps preparation with the model, so each fold learns fitted settings only from its training rows.
cross_validate
Fits fresh copies on training folds and returns their validation scores.
cross_val_predict
Makes one held-out training prediction per row for diagnosis; use cross_validate for the validation score.
clone(candidate).fit(X_train, y_train)
Starts an unfitted copy of the selected recipe and fits all development rows after selection.
test_score
Validation-fold scores inside cross_validate, despite the word test; the reserved final test is separate.
f1_macro
Computes F1 for each class, then averages them equally; balanced_accuracy instead averages class recall.
balanced_accuracy
Averages recall across classes, giving uncommon classes the same weight as common ones.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

from sklearn.model_selection import train_test_split, StratifiedKFold, cross_validate, cross_val_predict
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.dummy import DummyClassifier
from sklearn.base import clone
from sklearn.metrics import f1_score
# 1. Question, features available now, and later outcome
feature_names = ['length', 'width']
target = 'label'
X, y = df[feature_names], df[target]
# 2. Protect the final test; split X and y together
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=.2, random_state=42, stratify=y)
# 3. Inspect training rows only
print(X_train.describe())
# 4. Define preparation and the candidate; no fitting yet
candidates = {'simple': Pipeline([('prepare', StandardScaler()), ('model', LogisticRegression(max_iter=1000))])}
folds = StratifiedKFold(n_splits=3, shuffle=True, random_state=42)
metric = 'f1_macro'
# 5. Compare a reference and candidates on matching training folds
reference = DummyClassifier(strategy="most_frequent")
reference_results = cross_validate(reference, X_train, y_train, cv=folds, scoring=metric)
cv_results = {'simple': cross_validate(candidates['simple'], X_train, y_train, cv=folds, scoring=metric, return_train_score=True)}
print('Reference validation:', reference_results['test_score'])
print('Candidate validation:', cv_results['simple']['test_score'])
print('Candidate minus dummy mean:', cv_results['simple']['test_score'].mean() - reference_results['test_score'].mean())
# 6. Assess the candidate with training evidence; inspect training-only errors
chosen_name = 'simple'  # Sole candidate; report if it does not beat the dummy
diagnostic_predictions = cross_val_predict(candidates[chosen_name], X_train, y_train, cv=folds)
print(pd.crosstab(y_train, diagnostic_predictions))
# 7. Fit the fixed recipe; predict the final rows once
final_model = clone(candidates[chosen_name]).fit(X_train, y_train)
final_predictions = final_model.predict(X_test)
final_score = f1_score(y_test, final_predictions, average="macro")
print('Final metric:', final_score)
# 8. Explain the evidence and its limits in the self-review field

This practice: Run, predict and explain. Next: Change · Your first complete classification workflow.

The question and data dictionary

Classify a new specimen from measurements available before its label is confirmed.

Independent specimens from a stable measurement process; the three classes have equal importance.

Decide what can be known at prediction time
ColumnMeaning and availability
lengthMeasured length, in micrometres.
widthMeasured width, in millimetres.
labelConfirmed specimen class A, B or C.

The playground workflow

Use the same data boundary as ML → Workflow: frame, split, explore, prepare, validate against a reference, diagnose, then make the final evaluation. Each step answers one question.

  1. Question → X / yName the outcome and the inputs available when the prediction is needed. Later outcomes, identifiers and outcome-derived fields are not predictors.
  2. Split and protectReserve final rows before inspecting distributions or fitting. Independent rows permit a shuffled holdout; repeated entities or future forecasting need different designs.
  3. Explore training rowsInspect only development rows to identify types, missingness and class balance.
  4. Prepare → modelPut learned preparation inside a Pipeline so each fold learns it afresh. A numeric linear model can pass the original units through.
  5. Baseline → validateA dummy establishes what ignoring X achieves. Fit the candidate on matching training folds; these validation rows are not the final test.
  6. Choose → diagnoseChoose using training evidence, then inspect held-out training predictions. Simplify a model that fails to generalise; do not open the final test to choose.
  7. Final fit → predict → metricFix the recipe, fit all training rows, predict reserved rows once and calculate the declared metric from those saved predictions.
  8. Interpret and limitCompare validation with the baseline and final evidence. State units or class error costs, uncertainty from the small sample, and the population to which the claim applies.

Explain your decisions · self-review

Before Run, write your prediction. After Run, explain what the evidence supports and what it cannot establish. Code checks cannot award these reasoning scores.

  1. Frame and boundaryName prediction time, observation, target and unavailable inputs. Choose a split that matches intended use.
  2. Preparation and baselineFit learned preparation inside each training fold. Compare the dummy on those same folds.
  3. Validation and selectionChoose from training-fold scores and error patterns. Explain any gap between training and validation.
  4. Metric and final testJustify the metric and interpret the reserved-test result. Do not revise the recipe using that result.
  5. Interpretation and limitsQuote baseline, validation and final results. Explain error units or costs, one failure case and one limit; avoid causal claims.

Use the interpretation field beside your output. After an attempt, Check identifies code evidence and shows reasoning guidance. The optional explained solution is one defensible approach, not the only acceptable answer.

Logistic classification · from idea to workflow

Question: Which class is plausible from available inputs?

Mechanism: A linear score becomes class probabilities and then a decision.

Watch for: A linear boundary can miss curved separation; probability and decision cost are different.

Interpret or debug · self-review: Explain which error changes when the decision threshold moves, and why threshold selection belongs inside training validation.

Independent full workflow: ML-X05 · ML-X16. First practise complete regression and classification in F11–F12; check readiness in W-K2–W-K3.

Your inputs · CLASS180

180 synthetic independent observations. The dataframe df is supplied afresh on each Run. Column availability is described in the question’s dictionary above.

CLASS180 · first 8 rows
lengthwidthlabel
236.033-0.805568A
581.2970.728558A
-1511.27-1.00866A
99.0248-0.24496A
-13.0141-0.660765A
681.1790.602475A
51.14720.873157A
362.131-0.665605A

Your task · Follow

Predict the performance of an always-majority classifier on three balanced classes. Run the full workflow. Explain why scaling is learned inside each fold, and interpret the training-only confusion table before reading the final macro F1.

Required Python variables and evidence

Use these names so Check can inspect your workflow. Each meaning is shown beside its name.

Workflow evidence contract
VariableMeaning
target / feature_namesYour outcome column name and list of legitimate inputs. Choose from the data dictionary.
X / y / X_train / X_test / y_train / y_testOriginal indexed feature/target data and aligned partitions. A 15–30% holdout is supported; choose and justify its seed/design.
folds / metricA 3–5-fold shuffled KFold or StratifiedKFold object; metric is an sklearn scoring name. Regression: neg_root_mean_squared_error or neg_mean_absolute_error. Classification: f1_macro or balanced_accuracy.
reference / reference_resultsDummy estimator and its actual cross_validate result.
candidates / cv_resultsNamed Pipelines and matching cross_validate result dictionaries. Regression supports LinearRegression, Ridge and DecisionTreeRegressor; classification supports LogisticRegression, DecisionTreeClassifier, KNeighborsClassifier, GaussianNB, LinearDiscriminantAnalysis, SVC and MLPClassifier.
diagnostic_predictionsActual cross_val_predict outputs for the chosen candidate on training folds; inspect residuals or a confusion table.
chosen_name / final_modelName of your chosen validated candidate and a clone fitted on all training rows. Defend the choice, including any simplicity trade-off.
final_predictions / final_scoreOne saved final prediction array and its final metric. Positive error units for regression. No further selection after this call.
Hint 1 — Think

Predict the performance of an always-majority classifier on three balanced classes. Run the full workflow. Explain why scaling is learned inside each fold, and interpret the training-only confusion table before reading the final macro F1. Classify a new specimen from measurements available before its label is confirmed. Which fields exist at that moment?

Hint 2 — Tools

Use a dataframe/series pair, train_test_split, Pipeline, cross_validate and a dummy suited to class labels.

Hint 3 — Approach

Keep final rows outside every fit. Compare matching training-fold evidence before predicting reserved rows in CLASS180.

Explained solution
from sklearn.model_selection import train_test_split, StratifiedKFold, cross_validate, cross_val_predict
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.dummy import DummyClassifier
from sklearn.base import clone
from sklearn.metrics import f1_score
# 1. Question, features available now, and later outcome
feature_names = ['length', 'width']
target = 'label'
X, y = df[feature_names], df[target]
# 2. Protect the final test; split X and y together
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=.2, random_state=42, stratify=y)
# 3. Inspect training rows only
print(X_train.describe())
# 4. Define preparation and the candidate; no fitting yet
candidates = {'simple': Pipeline([('prepare', StandardScaler()), ('model', LogisticRegression(max_iter=1000))])}
folds = StratifiedKFold(n_splits=3, shuffle=True, random_state=42)
metric = 'f1_macro'
# 5. Compare a reference and candidates on matching training folds
reference = DummyClassifier(strategy="most_frequent")
reference_results = cross_validate(reference, X_train, y_train, cv=folds, scoring=metric)
cv_results = {'simple': cross_validate(candidates['simple'], X_train, y_train, cv=folds, scoring=metric, return_train_score=True)}
print('Reference validation:', reference_results['test_score'])
print('Candidate validation:', cv_results['simple']['test_score'])
print('Candidate minus dummy mean:', cv_results['simple']['test_score'].mean() - reference_results['test_score'].mean())
# 6. Assess the candidate with training evidence; inspect training-only errors
chosen_name = 'simple'  # Sole candidate; report if it does not beat the dummy
diagnostic_predictions = cross_val_predict(candidates[chosen_name], X_train, y_train, cv=folds)
print(pd.crosstab(y_train, diagnostic_predictions))
# 7. Fit the fixed recipe; predict the final rows once
final_model = clone(candidates[chosen_name]).fit(X_train, y_train)
final_predictions = final_model.predict(X_test)
final_score = f1_score(y_test, final_predictions, average="macro")
print('Final metric:', final_score)
# 8. Explain the evidence and its limits in the self-review field

Classify a new specimen from measurements available before its label is confirmed. The reference ignores features. Training-fold evidence informs the choice, and the final metric describes only the reserved sample. Excluding later outcomes prevents answering the question with information unavailable at prediction time. The reasoning rubric needs human review.

Helpful prior knowledge: Your first complete regression workflow · Preserving class representation These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Predict the performance of an always-majority classifier on three balanced classes. Run the full workflow. Explain why scaling is learned inside each fold, and interpret the training-only confusion table before reading the final macro F1.

reference_resultscv_resultschosen_namefinal_score
Keep beside your code

Classify a new specimen from measurements available before its label is confirmed.

lengthMeasured length, in micrometres.
widthMeasured width, in millimetres.
labelConfirmed specimen class A, B or C.

Pipeline Keeps preparation with the model, so each fold learns fitted settings only from its training rows.cross_validate Fits fresh copies on training folds and returns their validation scores.cross_val_predict Makes one held-out training prediction per row for diagnosis; use cross_validate for the validation score.

Frame → reserve final rows → compare dummy and candidate on training folds → diagnose → fit the fixed recipe → evaluate once.

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.

Which specimen class is most often missed in the out-of-fold predictions?

Use your Run output as evidence. This response is optional, not machine-graded or saved.