Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Data Foundations · A little practice goes a long way.

← Inspect lessonsTINY TABLES · REAL PYTHON · YOUR PACE
Understanding values · I17 · 8 MIN

Find duplicate rows

Count repeated records without removing anything.

Exercises within this concept

  1. FollowFollow the techniqueCurrent exercise
  2. ChangeAdapt a requirement
  3. TransferChoose and combine

Understand the idea

duplicated returns one True/False flag per row. By default it compares every column’s values, ignoring the index; it does not remove rows.

Concept sketch: one EXTRA duplicate; keep all rowsnamepriceextra?A2noB4noA2yesone EXTRA duplicate; keep all rows
Illustration · not the exercise output

A small example

Suppose the complete rows are A, B, A, A. A occurs three times; B occurs once. Read the flags in the same row order.

Same rows, different keep choices
Row"first""last"False
A · firstFalseTrueTrue
B · uniqueFalseFalseFalse
A · secondTrueTrueTrue
A · thirdTrueFalseTrue

Follow the code

Apply the idea to the supplied table. Read from top to bottom; the final line displays the result.

df.duplicated().sum()

What each part does

df.duplicated()
Mark later copies of complete rows True; the first copy stays False.
.sum()
Count those True flags to get the number of extra copies.
Other choices for later exercises
keep="last"
Exempt the last copy instead of the first.
keep=False
Mark every copy of a repeated row, including the first.
df[mask]
Show the rows whose duplicate flags are True.

Your inputs

The editable setup on the right creates df. Run executes the setup and your work from top to bottom.

Messy café orders · 8 synthetic rows
orderdrinksizepricetipdate
101" latte "LARGE6.2012026-06-01
102TEAsmall3.100.52026-06-02
103" mocha "LARGEoopsNonenot a date
104LatteSmallNone0.82026-06-04
104LatteSmallNone0.82026-06-04
105"tea "SMALL4.200.62026-06-05
106ESPRESSOsmall2.500.22026-06-06
107" mocha"large6.801.52026-06-07

Your task · Follow

  1. Using df, count extra copies of exact duplicate rows, excluding the first occurrence of each repeated row.
  2. Display one number.
  3. Use duplicated().
Hint

duplicated marks the extra copies only by default, not the first occurrence.

Reveal solution

One way to do it. Keep any supplied setup in the editor and use this in the Your work section.

df.duplicated().sum()
Your task · Follow
  1. Using df, count extra copies of exact duplicate rows, excluding the first occurrence of each repeated row.
  2. Display one number.
  3. Use duplicated().

Tab: indent · Shift+Tab: outdent · Esc, then Tab: leave editor

Edit Python. Control or Command plus Enter runs it. Tab indents by four spaces. Shift plus Tab outdents. Press Escape, then Tab or Shift plus Tab to leave the editor.

Each run executes all editor code in a fresh Python session. Display a value by leaving it on the final line.

Python starts when you open a lesson.

Output

Run your code to see what Python returns.