Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Data Foundations · A little practice goes a long way.

← Wrangle / Preprocess lessonsTINY TABLES · REAL PYTHON · YOUR PACE
Types, dates & missingness · W18 · 8 MIN

Remove duplicates

Remove verified extra copies.

Exercises within this concept

  1. FollowFollow the techniqueCurrent exercise
  2. ChangeAdapt a requirement
  3. TransferChoose and combine

Understand the idea

drop_duplicates removes repeated rows. Decide which copy survives; rows that occur only once are always retained.

Concept sketch: remove extra copies, keep first rownamepriceA2B4A2namepriceA2B4remove extra copies, keep first row
Illustration · not the exercise output

A small example

For full rows A, B, A, A at indices 0, 1, 2, 3: keeping first retains indices 0, 1; keeping last retains 1, 3.

Follow the code

Apply the idea to the supplied table. Read from top to bottom; the final line displays the result.

df = df.drop_duplicates()
df

What each part does

df.drop_duplicates()
remove extra identical rows
keep="first"
retain first occurrence
Other choices for later exercises
subset=["order"]
compare only this key if it must be unique

Your inputs

The editable setup on the right creates df. Run executes the setup and your work from top to bottom.

Messy café orders · 8 synthetic rows
orderdrinksizepricetipdate
101" latte "LARGE6.2012026-06-01
102TEAsmall3.100.52026-06-02
103" mocha "LARGEoopsNonenot a date
104LatteSmallNone0.82026-06-04
104LatteSmallNone0.82026-06-04
105"tea "SMALL4.200.62026-06-05
106ESPRESSOsmall2.500.22026-06-06
107" mocha"large6.801.52026-06-07

Your task · Follow

  1. Using df, remove extra exact duplicate rows and preserve the first occurrence and its index.
  2. Keep the changes in df and display it.
  3. Use: drop_duplicates().
Hint

These identical rows are confirmed copies; the default keeps the first one.

Reveal solution

One way to do it. Keep any supplied setup in the editor and use this in the Your work section.

df = df.drop_duplicates()
df
Your task · Follow
  1. Using df, remove extra exact duplicate rows and preserve the first occurrence and its index.
  2. Keep the changes in df and display it.
  3. Use: drop_duplicates().

Tab: indent · Shift+Tab: outdent · Esc, then Tab: leave editor

Edit Python. Control or Command plus Enter runs it. Tab indents by four spaces. Shift plus Tab outdents. Press Escape, then Tab or Shift plus Tab to leave the editor.

Each run executes all editor code in a fresh Python session. Display a value by leaving it on the final line.

Python starts when you open a lesson.

Output

Run your code to see what Python returns.