Understand the idea
duplicated returns one True/False flag per row. By default it compares every column’s values, ignoring the index; it does not remove rows.
A small example
Suppose the complete rows are A, B, A, A. A occurs three times; B occurs once. Read the flags in the same row order.
| Row | "first" | "last" | False |
|---|---|---|---|
| A · first | False | True | True |
| B · unique | False | False | False |
| A · second | True | True | True |
| A · third | True | False | True |
Follow the code
Apply the idea to the supplied table. Read from top to bottom; the final line displays the result.
df.duplicated().sum()What each part does
df.duplicated()- Mark later copies of complete rows True; the first copy stays False.
.sum()- Count those True flags to get the number of extra copies.
Other choices for later exercises
keep="last"- Exempt the last copy instead of the first.
keep=False- Mark every copy of a repeated row, including the first.
df[mask]- Show the rows whose duplicate flags are True.
Your inputs
The editable setup on the right creates df. Run executes the setup and your work from top to bottom.
| order | drink | size | price | tip | date |
|---|---|---|---|---|---|
| 101 | " latte " | LARGE | 6.20 | 1 | 2026-06-01 |
| 102 | TEA | small | 3.10 | 0.5 | 2026-06-02 |
| 103 | " mocha " | LARGE | oops | None | not a date |
| 104 | Latte | Small | None | 0.8 | 2026-06-04 |
| 104 | Latte | Small | None | 0.8 | 2026-06-04 |
| 105 | "tea " | SMALL | 4.20 | 0.6 | 2026-06-05 |
| 106 | ESPRESSO | small | 2.50 | 0.2 | 2026-06-06 |
| 107 | " mocha" | large | 6.80 | 1.5 | 2026-06-07 |
Your task · Follow
- Using df, count extra copies of exact duplicate rows, excluding the first occurrence of each repeated row.
- Display one number.
- Use duplicated().
Hint
duplicated marks the extra copies only by default, not the first occurrence.
Reveal solution
One way to do it. Keep any supplied setup in the editor and use this in the Your work section.
df.duplicated().sum()