For df records with age at least 2, display average weight and row count by species.
The editable setup on the right creates df. Run executes the setup and your work from top to bottom.
Your requirements
- Use columns species, average, n, in that order.
Your inputs
| name | species | age | weight | room |
|---|---|---|---|---|
| Milo | Cat | 3 | 4.2 | A |
| Pepper | Dog | 7 | 18.5 | B |
| Luna | Cat | 2 | 3.6 | A |
| Bean | Rabbit | 4 | 2.4 | B |
| Rex | Dog | 5 | 22 | A |
| Nori | Rabbit | 1 | 1.8 | B |
Remember the idea
Rows with the same flavour form a group. .agg(...) means aggregate: it makes one result row per group by running the calculations inside. Each new_name=("input", "operation") defines one output column: the left name is its heading; the pair chooses the source field and calculation.
A small example
Fruity has prices 2 and 4, so its mean_amount is (2 + 4) / 2 = 3 and records is 2. Mint has one price, 8, so its mean_amount is 8 and records is 1.
| flavour | mean_amount | records |
|---|---|---|
| fruity | 3 | 2 |
| mint | 8 | 1 |
| Code or choice | Meaning |
|---|---|
df.groupby("flavour", as_index=False) | Collect rows by flavour and keep flavour as a regular output column. |
.agg(...) | Aggregate: reduce each group to one result row using the named calculations inside. |
mean_amount=("price", "mean") | Create an output column called mean_amount from the mean of price in each group. |
records=("price", "size") | Create an output column called records from the row count in each group. |
Hint
size counts records; count counts known measurements.
Reveal solution
One way to do it. Keep any supplied setup in the editor and use this in the Your work section.
df[df["age"] >= 2].groupby("species", as_index=False).agg(
average=("weight", "mean"), n=("weight", "size")
)