At inspection, identify units likely to develop a fault during the following week so a technician can prioritise checks.
Independently prioritise units that may develop a fault next week. Choose and justify the target, available features, split, metric and simple candidate(s). Build the workflow without viewing the solution first. Use actual numbers to explain baseline, validation, final result and limitations. The rubric is visible; reasoning requires self-review or lecturer review.
This practice: Build and explain independently. Next: choose Regression or Classification. Both routes build on the shared supervised workflow.
The question and data dictionary
At inspection, identify units likely to develop a fault during the following week so a technician can prioritise checks.
Independent units inspected under one stable operating regime. Faults are uncommon; missing faults and unnecessary inspections both matter. This synthetic exercise does not establish deployment safety.
| Column | Meaning and availability |
|---|---|
| unit_id | Administrative unit identifier with no predictive ordering. Each unit appears once. |
| vibration_mm_s | Vibration at inspection in mm/s. |
| temperature_c | Temperature at inspection in degrees Celsius. |
| next_week_state | State confirmed one week later: fault or normal. |
| replacement_authorized | Decision entered after the later fault assessment. |
The playground workflow
Use the same data boundary as ML → Workflow: frame, split, explore, prepare, validate against a reference, diagnose, then make the final evaluation. Each step answers one question.
- Question → X / yName the outcome and the inputs available when the prediction is needed. Later outcomes, identifiers and outcome-derived fields are not predictors.
- Split and protectReserve final rows before inspecting distributions or fitting. Independent rows permit a shuffled holdout; repeated entities or future forecasting need different designs.
- Explore training rowsInspect only development rows to identify types, missingness and class balance.
- Prepare → modelPut learned preparation inside a Pipeline so each fold learns it afresh. A numeric linear model can pass the original units through.
- Baseline → validateA dummy establishes what ignoring X achieves. Fit the candidate on matching training folds; these validation rows are not the final test.
- Choose → diagnoseChoose using training evidence, then inspect held-out training predictions. Simplify a model that fails to generalise; do not open the final test to choose.
- Final fit → predict → metricFix the recipe, fit all training rows, predict reserved rows once and calculate the declared metric from those saved predictions.
- Interpret and limitCompare validation with the baseline and final evidence. State units or class error costs, uncertainty from the small sample, and the population to which the claim applies.
Readiness rubric · self-review
For each criterion: 0 = missing or unsafe; 1 = plausible but unsupported; 2 = correct and supported by your actual evidence. Aim for 2 on every criterion. A target leak, test-driven selection or test-row fit means the boundary must be repaired before a readiness claim. Code checks cannot award these reasoning scores.
- Frame and boundaryName prediction time, observation, target and unavailable inputs. Choose a split that matches intended use.
- Preparation and baselineFit learned preparation inside each training fold. Compare the dummy on those same folds.
- Validation and selectionChoose from training-fold scores and error patterns. Explain any gap between training and validation.
- Metric and final testJustify the metric and interpret the reserved-test result. Do not revise the recipe using that result.
- Interpretation and limitsQuote baseline, validation and final results. Explain error units or costs, one failure case and one limit; avoid causal claims.
Use the interpretation field beside your output. After an attempt, Check identifies code evidence and shows reasoning guidance. The optional explained solution is one defensible approach, not the only acceptable answer.
Your inputs · SENSOR150
150 synthetic independent observations. The dataframe df is supplied afresh on each Run. Column availability is described in the question’s dictionary above.
| unit_id | vibration_mm_s | temperature_c | next_week_state | replacement_authorized |
|---|---|---|---|---|
| 3000 | 0.903212 | 38.6784 | normal | 0 |
| 3001 | 3.95001 | 62.9738 | normal | 0 |
| 3002 | 2.2144 | 62.5315 | normal | 0 |
| 3003 | 1.14971 | 53.1204 | normal | 0 |
| 3004 | 1.49546 | 64.8484 | normal | 0 |
| 3005 | 1.4053 | 27.1735 | normal | 0 |
| 3006 | 5.66259 | 63.254 | fault | 1 |
| 3007 | 1.58558 | 56.2962 | normal | 0 |
Supporting concepts: When a good score is misleading → · Learn missing-value replacements safely → · Respect time →
Remember the idea
At inspection, identify units likely to develop a fault during the following week so a technician can prioritise checks. Use the dictionary to frame a prediction claim. This dataset has not appeared in the teaching path. A working script is only one part of readiness: each decision also needs an evidence-based explanation.
Required Python variables and evidence
Use these names so Check can inspect your workflow. Each meaning is shown beside its name.
| Variable | Meaning |
|---|---|
| target / feature_names | Your outcome column name and list of legitimate inputs. Choose from the data dictionary. |
| X / y / X_train / X_test / y_train / y_test | Original indexed feature/target data and aligned partitions. A 15–30% holdout is supported; choose and justify its seed/design. |
| folds / metric | A 3–5-fold shuffled KFold or StratifiedKFold object; metric is an sklearn scoring name. Regression: neg_root_mean_squared_error or neg_mean_absolute_error. Classification: f1_macro or balanced_accuracy. |
| reference / reference_results | Dummy estimator and its actual cross_validate result. |
| candidates / cv_results | Named Pipelines and matching cross_validate result dictionaries. Regression supports LinearRegression, Ridge and DecisionTreeRegressor; classification supports LogisticRegression, DecisionTreeClassifier, KNeighborsClassifier, GaussianNB, LinearDiscriminantAnalysis, SVC and MLPClassifier. |
| diagnostic_predictions | Actual cross_val_predict outputs for the chosen candidate on training folds; inspect residuals or a confusion table. |
| chosen_name / final_model | Name of your chosen validated candidate and a clone fitted on all training rows. Defend the choice, including any simplicity trade-off. |
| final_predictions / final_score | One saved final prediction array and its final metric. Positive error units for regression. No further selection after this call. |
Hint 1 — Think
Independently prioritise units that may develop a fault next week. Choose and justify the target, available features, split, metric and simple candidate(s). Build the workflow without viewing the solution first. Use actual numbers to explain baseline, validation, final result and limitations. The rubric is visible; reasoning requires self-review or lecturer review. At inspection, identify units likely to develop a fault during the following week so a technician can prioritise checks. Which fields exist at that moment?
Hint 2 — Tools
Use a dataframe/series pair, train_test_split, Pipeline, cross_validate and a dummy suited to class labels.
Hint 3 — Approach
Keep final rows outside every fit. Compare matching training-fold evidence before predicting reserved rows in SENSOR150.
Explained solution
from sklearn.model_selection import train_test_split, StratifiedKFold, cross_validate, cross_val_predict
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.dummy import DummyClassifier
from sklearn.base import clone
from sklearn.metrics import f1_score
# 1. Question, features available now, and later outcome
feature_names = ['vibration_mm_s', 'temperature_c']
target = 'next_week_state'
X, y = df[feature_names], df[target]
# 2. Protect the final test; split X and y together
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=.2, random_state=42, stratify=y)
# 3. Inspect training rows only
print(X_train.describe())
# 4. Define preparation and the candidate; no fitting yet
candidates = {'simple': Pipeline([('prepare', StandardScaler()), ('model', LogisticRegression(max_iter=1000))])}
folds = StratifiedKFold(n_splits=3, shuffle=True, random_state=42)
metric = 'f1_macro'
# 5. Compare a reference and candidates on matching training folds
reference = DummyClassifier(strategy="most_frequent")
reference_results = cross_validate(reference, X_train, y_train, cv=folds, scoring=metric)
cv_results = {'simple': cross_validate(candidates['simple'], X_train, y_train, cv=folds, scoring=metric, return_train_score=True)}
print('Reference validation:', reference_results['test_score'])
print('Candidate validation:', cv_results['simple']['test_score'])
print('Candidate minus dummy mean:', cv_results['simple']['test_score'].mean() - reference_results['test_score'].mean())
# 6. Assess the candidate with training evidence; inspect training-only errors
chosen_name = 'simple' # Sole candidate; report if it does not beat the dummy
diagnostic_predictions = cross_val_predict(candidates[chosen_name], X_train, y_train, cv=folds)
print(pd.crosstab(y_train, diagnostic_predictions))
# 7. Fit the fixed recipe; predict the final rows once
final_model = clone(candidates[chosen_name]).fit(X_train, y_train)
final_predictions = final_model.predict(X_test)
final_score = f1_score(y_test, final_predictions, average="macro")
print('Final metric:', final_score)
# 8. Explain the evidence and its limits in the self-review field
At inspection, identify units likely to develop a fault during the following week so a technician can prioritise checks. The reference ignores features. Training-fold evidence informs the choice, and the final metric describes only the reserved sample. Excluding later outcomes prevents answering the question with information unavailable at prediction time. The reasoning rubric needs human review.
Helpful prior knowledge: When a good score is misleading · Supervised Workflow checkpoint These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.