Understand the idea
Rows with the same flavour form a group. .agg(...) means aggregate: it makes one result row per group by running the calculations inside. Each new_name=("input", "operation") defines one output column: the left name is its heading; the pair chooses the source field and calculation.
A small example
Fruity has prices 2 and 4, so its mean_amount is (2 + 4) / 2 = 3 and records is 2. Mint has one price, 8, so its mean_amount is 8 and records is 1.
| flavour | mean_amount | records |
|---|---|---|
| fruity | 3 | 2 |
| mint | 8 | 1 |
Follow the code
Apply the idea to the supplied table. Read from top to bottom; the final line displays the result.
df.groupby("flavour", as_index=False).agg(
mean_amount=("price", "mean"),
records=("price", "size")
)What each part does
df.groupby("flavour", as_index=False)- group rows by flavour; keep flavour as an output column
.agg(...)- Aggregate the rows in each flavour group. The calculations named inside these parentheses produce the result columns.
mean_amount=- Inside .agg(...), this chosen name becomes a new result column heading. There is no earlier mean_amount column to find.
("price", "mean")- The parentheses and comma make a two-item tuple: price is the input column; mean averages its known values within each group.
records=- Inside .agg(...), this chosen name becomes the second result column heading: the count of records in the group.
("price", "size")- The pair names price as the input column and size as the pandas operation used for records.
Your inputs
The editable setup on the right creates df. Run executes the setup and your work from top to bottom.
| candy | flavour | price | rating | shelf |
|---|---|---|---|---|
| Gummy Bear | fruity | 1.2 | 4.1 | A |
| Choco Pop | chocolate | 2.1 | 4.6 | B |
| Mint Bite | mint | 1.5 | 3.8 | A |
| Berry Loop | fruity | 2.8 | 4.4 | B |
| Cocoa Cube | chocolate | 3.4 | 4.9 | A |
| Lemon Drop | fruity | 1.8 | 4 | B |
Your task · Follow
- Using df, group by flavour.
- Return a DataFrame with flavour, mean_amount (mean price) and records (row count), in that order.
- Use: groupby(), agg().
Hint
Inside .agg(...), mean_amount=("price", "mean") averages price for each flavour. records=("price", "size") counts all rows in that flavour group.
Reveal solution
One way to do it. Keep any supplied setup in the editor and use this in the Your work section.
df.groupby("flavour", as_index=False).agg(
mean_amount=("price", "mean"), records=("price", "size")
)