Understand the idea
cut assigns values to fixed intervals whose boundaries you choose. Add a label for each interval to make the result readable.
A small example
For bins [0, 3, 10], right=True and include_lowest=True: 0 and 3 are in the first bin; values above 3 through 10 are in the second.
Follow the code
Apply the idea to the supplied table. Read from top to bottom; the final line displays the result.
df["half"] = pd.cut(
df["price"], bins=[0, 2.1, 100], labels=["lower", "upper"], include_lowest=True
)
dfWhat each part does
pd.cut(df["price"], bins=[0, 2.1, 100], ...)- Place each price into an interval bounded by the listed edges.
labels=["lower", "upper"]- Name the first and second intervals in order.
include_lowest=True- Include the lowest edge, 0, in the first interval.
Your inputs
The editable setup on the right creates df. Run executes the setup and your work from top to bottom.
| candy | flavour | price | rating | shelf |
|---|---|---|---|---|
| Gummy Bear | fruity | 1.2 | 4.1 | A |
| Choco Pop | chocolate | 2.1 | 4.6 | B |
| Mint Bite | mint | 1.5 | 3.8 | A |
| Berry Loop | fruity | 2.8 | 4.4 | B |
| Cocoa Cube | chocolate | 3.4 | 4.9 | A |
| Lemon Drop | fruity | 1.8 | 4 | B |
Your task · Follow
- Using df, add half using cut with boundaries 0, 2.1, and 100; labels "lower" and "upper".
- Include zero in the first bin.
- Keep the changes in df and display it.
- Each interval includes its right boundary.
Hint
Fixed boundaries and sample quantiles answer different questions; include the lowest boundary explicitly.
Reveal solution
One way to do it. Keep any supplied setup in the editor and use this in the Your work section.
df["half"] = pd.cut(
df["price"], bins=[0, 2.1, 100], labels=["lower", "upper"], include_lowest=True
)
dfOptional stretch
Try a price exactly equal to the boundary. Which interval gets it? What happens to a price above the final edge?