Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Data Foundations · A little practice goes a long way.

← Wrangle / Preprocess lessonsTINY TABLES · REAL PYTHON · YOUR PACE
Groups & reshaping · W19 · 8 MIN

Group and aggregate

Create readable, named group summaries.

Exercises within this concept

  1. FollowFollow the techniqueCurrent exercise
  2. ChangeAdapt a requirement
  3. TransferChoose and combine

Understand the idea

Rows with the same flavour form a group. .agg(...) means aggregate: it makes one result row per group by running the calculations inside. Each new_name=("input", "operation") defines one output column: the left name is its heading; the pair chooses the source field and calculation.

Concept sketch: one row per group, multiple summariesgroupvalueA2A4B8groupmeannA32B81one row per group, multiple summaries
Illustration · not the exercise output

A small example

Fruity has prices 2 and 4, so its mean_amount is (2 + 4) / 2 = 3 and records is 2. Mint has one price, 8, so its mean_amount is 8 and records is 1.

Result after .agg(...)
flavourmean_amountrecords
fruity32
mint81

Follow the code

Apply the idea to the supplied table. Read from top to bottom; the final line displays the result.

df.groupby("flavour", as_index=False).agg(
    mean_amount=("price", "mean"),
    records=("price", "size")
)

What each part does

df.groupby("flavour", as_index=False)
group rows by flavour; keep flavour as an output column
.agg(...)
Aggregate the rows in each flavour group. The calculations named inside these parentheses produce the result columns.
mean_amount=
Inside .agg(...), this chosen name becomes a new result column heading. There is no earlier mean_amount column to find.
("price", "mean")
The parentheses and comma make a two-item tuple: price is the input column; mean averages its known values within each group.
records=
Inside .agg(...), this chosen name becomes the second result column heading: the count of records in the group.
("price", "size")
The pair names price as the input column and size as the pandas operation used for records.

Your inputs

The editable setup on the right creates df. Run executes the setup and your work from top to bottom.

Candy shop · 6 synthetic rows
candyflavourpriceratingshelf
Gummy Bearfruity1.24.1A
Choco Popchocolate2.14.6B
Mint Bitemint1.53.8A
Berry Loopfruity2.84.4B
Cocoa Cubechocolate3.44.9A
Lemon Dropfruity1.84B

Your task · Follow

  1. Using df, group by flavour.
  2. Return a DataFrame with flavour, mean_amount (mean price) and records (row count), in that order.
  3. Use: groupby(), agg().
Hint

Inside .agg(...), mean_amount=("price", "mean") averages price for each flavour. records=("price", "size") counts all rows in that flavour group.

Reveal solution

One way to do it. Keep any supplied setup in the editor and use this in the Your work section.

df.groupby("flavour", as_index=False).agg(
    mean_amount=("price", "mean"), records=("price", "size")
)
Your task · Follow
  1. Using df, group by flavour.
  2. Return a DataFrame with flavour, mean_amount (mean price) and records (row count), in that order.
  3. Use: groupby(), agg().

Tab: indent · Shift+Tab: outdent · Esc, then Tab: leave editor

Edit Python. Control or Command plus Enter runs it. Tab indents by four spaces. Shift plus Tab outdents. Press Escape, then Tab or Shift plus Tab to leave the editor.

Each run executes all editor code in a fresh Python session. Display a value by leaving it on the final line.

Python starts when you open a lesson.

Output

Run your code to see what Python returns.