Understand the idea
Production uses fold-local quantile bins for continuous values. Binary flags stay discrete. This is the Playground’s implementation, which differs in discretisation details from the original One-R paper.
Python skill: Addresses the One-R learner’s numeric-bin setting through the pipeline.
Meet the syntax
GridSearchCV(model, {'model__bins':[3,5,8]}, ...)
model.named_steps["model"].categorical_mask_'model__bins'- Addresses the One-R learner’s numeric-bin setting through the pipeline.
[3,5,8]- Compares discretisation choices inside training-fold fits, not on the entire population.
model.named_steps["model"].categorical_mask_- Identifies which prepared feature positions are treated as categorical rather than numeric.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
model.fit(X,y)
answer=model.predict(X)
This practice: Read and run the Python. Next: Change · One-R with numeric inputs.
Given data · RULE24
24 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.
| distance | service | fragile | label |
|---|---|---|---|
| 1 | standard | 0 | on time |
| 2 | express | 1 | on time |
| 3 | economy | 0 | late |
| 4 | standard | 1 | on time |
| 5 | express | 0 | on time |
| 6 | economy | 1 | on time |
| 7 | standard | 0 | on time |
| 8 | express | 1 | on time |
Column meanings and units
Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.
Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.
| Column | Stored type |
|---|---|
| distance | int64 |
| service | str |
| fragile | int64 |
| label | str |
Supplied setup · available if you need to inspect it
This code runs before your editor on every Run. These are the objects your exercise uses.
from ml_helpers import OneRClassifier,OneRPreprocessor
from sklearn.pipeline import Pipeline
from sklearn.model_selection import GridSearchCV,StratifiedKFold
X=df[['distance','fragile']]
y=df.label
model=Pipeline([('prepare',OneRPreprocessor(numeric_features=['distance'],categorical_features=['fragile'])),('model',OneRClassifier(bins=5))])
Your task · Follow
Fit the numeric rule and inspect its predictions.
Hint 1 — Think
Numeric One-R learns intervals before it can map new values to labels.
Hint 2 — Tools
Pipeline.fit and predict with numeric-rule preparation.
Hint 3 — Approach
Fit the supplied recipe on the aligned features and labels, then inspect predictions.
Explained solution
model.fit(X,y)
answer=model.predict(X)
The fit learns the interval/rule structure; these in-sample outputs illustrate behaviour rather than held-away performance.
Helpful prior knowledge: One feature, one rule These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.