Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Classification lessonsQUESTIONS · MODELS · EVIDENCE
Margins and rules · ML-C13 · 12–18 MIN

One-R with numeric inputs

Keep numeric bin learning inside each fit.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeAdapt the Python
  3. TransferAdapt the Python

Understand the idea

Production uses fold-local quantile bins for continuous values. Binary flags stay discrete. This is the Playground’s implementation, which differs in discretisation details from the original One-R paper.

Keep numeric bin learning inside each fit.Numeric intervalRule predictionx < aClass Aa ≤ x < bClass Bx ≥ bClass ACut points a and b are learned inside each fit.
Schematic · Keep numeric bin learning inside each fit.Scroll the diagram horizontally if needed.

Python skill: Addresses the One-R learner’s numeric-bin setting through the pipeline.

Meet the syntax

GridSearchCV(model, {'model__bins':[3,5,8]}, ...)
model.named_steps["model"].categorical_mask_
'model__bins'
Addresses the One-R learner’s numeric-bin setting through the pipeline.
[3,5,8]
Compares discretisation choices inside training-fold fits, not on the entire population.
model.named_steps["model"].categorical_mask_
Identifies which prepared feature positions are treated as categorical rather than numeric.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

model.fit(X,y)
answer=model.predict(X)

This practice: Read and run the Python. Next: Change · One-R with numeric inputs.

Given data · RULE24

24 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

RULE24 · first 8 prepared rows
distanceservicefragilelabel
1standard0on time
2express1on time
3economy0late
4standard1on time
5express0on time
6economy1on time
7standard0on time
8express1on time

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
distanceint64
servicestr
fragileint64
labelstr
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from ml_helpers import OneRClassifier,OneRPreprocessor
from sklearn.pipeline import Pipeline
from sklearn.model_selection import GridSearchCV,StratifiedKFold
X=df[['distance','fragile']]
y=df.label
model=Pipeline([('prepare',OneRPreprocessor(numeric_features=['distance'],categorical_features=['fragile'])),('model',OneRClassifier(bins=5))])

Your task · Follow

Fit the numeric rule and inspect its predictions.

Hint 1 — Think

Numeric One-R learns intervals before it can map new values to labels.

Hint 2 — Tools

Pipeline.fit and predict with numeric-rule preparation.

Hint 3 — Approach

Fit the supplied recipe on the aligned features and labels, then inspect predictions.

Explained solution
model.fit(X,y)
answer=model.predict(X)

The fit learns the interval/rule structure; these in-sample outputs illustrate behaviour rather than held-away performance.

Helpful prior knowledge: One feature, one rule These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Fit the numeric rule and inspect its predictions.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.