Understand the idea
value_counts counts rows for each distinct value. normalize=True divides each count by the number of non-missing values being counted.
A small example
For A, A, B: counts are A: 2 and B: 1; proportions are 2/3 and 1/3; percentages are about 66.7 and 33.3.
Follow the code
Apply the idea to the supplied table. Read from top to bottom; the final line displays the result.
df["flavour"].value_counts(normalize=True)What each part does
value_counts()- counts per category
normalize=True- fractions rather than counts
Other choices for later exercises
dropna=False- also count missing values
Your inputs
The editable setup on the right creates df. Run executes the setup and your work from top to bottom.
| candy | flavour | price | rating | shelf |
|---|---|---|---|---|
| Gummy Bear | fruity | 1.2 | 4.1 | A |
| Choco Pop | chocolate | 2.1 | 4.6 | B |
| Mint Bite | mint | 1.5 | 3.8 | A |
| Berry Loop | fruity | 2.8 | 4.4 | B |
| Cocoa Cube | chocolate | 3.4 | 4.9 | A |
| Lemon Drop | fruity | 1.8 | 4 | B |
Your task · Follow
- Using df, return the proportion of observations in each flavour category.
- Use: value_counts().
Hint
normalize=True changes counts to proportions; missing values are excluded by default.
Reveal solution
One way to do it. Keep any supplied setup in the editor and use this in the Your work section.
df["flavour"].value_counts(normalize=True)