Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Classification lessonsQUESTIONS · MODELS · EVIDENCE
Class distributions · ML-C16 · 18–25 MIN

Match Naive Bayes to feature types

Use each production Naive Bayes preparation path.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeReason about the Python
  3. PractiseAdapt the Python
  4. TransferReason about the Python

Understand the idea

Production selects GaussianNB for continuous inputs, BernoulliNB for binary flags and OHE–BernoulliNB for categorical inputs. It does not expose arbitrary mixed inputs or CategoricalNB.

Use each production Naive Bayes preparation path.ObservationFeature X₁Feature X₂Target yrow 12standardknownrow 25expressknownrow 38economyunknownKeep row identities aligned; new rows provide X.
Schematic · Use each production Naive Bayes preparation path.Scroll the diagram horizontally if needed.

Python skill: Uses continuous measurement likelihoods.

Meet the syntax

GaussianNB()
BernoulliNB()
Pipeline([("encode", OneHotEncoder(...)), ("model", BernoulliNB())])
GaussianNB()
Uses continuous measurement likelihoods.
BernoulliNB()
Uses presence and absence of binary indicators.
OneHotEncoder(...)
Turns named categories into indicators before the production Bernoulli path; learn its schema inside fitting.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

from sklearn.naive_bayes import BernoulliNB
flags=['chocolate','fruity','caramel','peanutyalmondy','nougat','crispedricewafer','hard','bar','pluribus']
model=BernoulliNB().fit(train[flags],train.popular)
answer=model.predict(test[flags])

This practice: Read and run the Python. Next: Change · Match Naive Bayes to feature types.

Given data · candy_class

85 observations. One candy product in the survey. The dataframe df is supplied afresh for each Run.

Download source CSV · Source and original dictionary

candy_class · first 8 prepared rows
competitornamechocolatefruitycaramelpeanutyalmondynougatcrispedricewaferhardbarpluribussugarpercentpricepercentwinpercentpopular
100 Grand1010010100.7320.8666.971750% or above
3 Musketeers1000100100.6040.51167.602950% or above
One dime0000000000.0110.11632.2611below 50%
One quarter0000000000.0110.51146.1165below 50%
Air Heads0100000000.9060.51152.341550% or above
Almond Joy1001000100.4650.76750.347550% or above
Baby Ruth1011100100.6040.76756.914550% or above
Boston Baked Beans0001000010.3130.51123.4178below 50%

Column meanings and units

sugarpercent: sugar percentile. pricepercent: price percentile. winpercent: percentage of survey matchups won. Ingredient, bar and multipack fields are 0/1 flags.

Percentile ranks are neither physical sugar percentages nor currency prices. Ratios of percentiles do not measure economic value. Chocolate and fruit flags can overlap; non-chocolate is not synonymous with fruit. The ML class target is defined by winpercent ≥ 50.

Input schema
ColumnStored type
competitornamestr
chocolateint64
fruityint64
caramelint64
peanutyalmondyint64
nougatint64
crispedricewaferint64
hardint64
barint64
pluribusint64
sugarpercentfloat64
pricepercentfloat64
winpercentfloat64
popularstr
Supplied setup · available if you need to inspect it

This code runs before your editor on every Run. These are the objects your exercise uses.

from sklearn.model_selection import train_test_split
train,test=train_test_split(df,test_size=.2,random_state=42,stratify=df.popular)

Your task · Follow

Fit BernoulliNB to Candy’s binary flags and predict held-away rows.

Hint 1 — Think

The declared target must stay separate from the presence/absence indicators.

Hint 2 — Tools

BernoulliNB and an explicit flags list.

Hint 3 — Approach

Fit only the nine permitted flags to the training popular labels and predict the same test schema.

Explained solution
from sklearn.naive_bayes import BernoulliNB
flags=['chocolate','fruity','caramel','peanutyalmondy','nougat','crispedricewafer','hard','bar','pluribus']
model=BernoulliNB().fit(train[flags],train.popular)
answer=model.predict(test[flags])

The binary likelihood matches the feature meanings and excludes the target-derived winpercent source.

Helpful prior knowledge: Gaussian Naive Bayes · Encode categories without inventing order These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

Fit BernoulliNB to Candy’s binary flags and predict held-away rows.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.