Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Data Foundations · A little practice goes a long way.

← Wrangle / Preprocess lessonsTINY TABLES · REAL PYTHON · YOUR PACE
A responsible finish · W31 · 20 MIN

Wrangle checkpoint

Task 1 · Retrieve and combine

Exercises within this concept

  1. Task 1Retrieve and combineCurrent exercise
  2. Task 2Adapt the workflow
  3. Task 3Combine earlier skills
Task 1 · Messy café orders

Create clean as a separate copy of df.

The editable setup on the right creates df. Run executes the setup and your work from top to bottom.

Your requirements

  1. Remove confirmed extra identical rows, keeping the first occurrence.
  2. Strip and lowercase drink; strip and title-case size.
  3. Convert price to numeric and date to datetime with year-month-day format, making invalid values missing.
  4. Fill missing tip with the median after deduplication.
  5. Sort by order ascending and reset to consecutive row labels without adding an index column.
  6. Keep all other values and retain rows with unknown price.
  7. Display clean.

Your inputs

Messy café orders · 8 synthetic rows
orderdrinksizepricetipdate
101" latte "LARGE6.2012026-06-01
102TEAsmall3.100.52026-06-02
103" mocha "LARGEoopsNonenot a date
104LatteSmallNone0.82026-06-04
104LatteSmallNone0.82026-06-04
105"tea "SMALL4.200.62026-06-05
106ESPRESSOsmall2.500.22026-06-06
107" mocha"large6.801.52026-06-07

Revisit the concept lesson →

Remember the idea

Concept sketch: copy → deduplicate → clean → sortraw" TEA "" TEA ""Mint "clean"mint""tea"original stays intactcopy → deduplicate → clean → sort

A cleaning policy specifies which records and values may change, and why.

A small example

Parse a numeric field before deciding whether it meets a report’s required-field policy.

What each choice does
Code or choiceMeaning
Copy and deduplicateKeep the source; remove only confirmed accidental copies.
Normalize and parseClean labels, convert types and keep invalid values visible as gaps.
Prepare the handoffApply the stated fill and eligibility rules, then select, sort and reset as requested.
Hint

Copy first. Remove confirmed duplicate records before calculating the imputation median; preserve unknown prices.

Reveal solution

One way to do it. Keep any supplied setup in the editor and use this in the Your work section.

clean = df.copy()
clean = clean.drop_duplicates()
clean["drink"] = clean["drink"].str.strip().str.lower()
clean["size"] = clean["size"].str.strip().str.title()
clean["price"] = pd.to_numeric(clean["price"], errors="coerce")
clean["date"] = pd.to_datetime(clean["date"], format="%Y-%m-%d", errors="coerce")
clean["tip"] = clean["tip"].fillna(clean["tip"].median())
clean = clean.sort_values("order").reset_index(drop=True)
clean
Your task · Task 1
  1. Create clean as a separate copy of df.
  2. Remove confirmed extra identical rows, keeping the first occurrence.
  3. Strip and lowercase drink; strip and title-case size.
  4. Convert price to numeric and date to datetime with year-month-day format, making invalid values missing.
  5. Fill missing tip with the median after deduplication.
  6. Sort by order ascending and reset to consecutive row labels without adding an index column.
  7. Keep all other values and retain rows with unknown price.
  8. Display clean.

Tab: indent · Shift+Tab: outdent · Esc, then Tab: leave editor

Edit Python. Control or Command plus Enter runs it. Tab indents by four spaces. Shift plus Tab outdents. Press Escape, then Tab or Shift plus Tab to leave the editor.

Each run executes all editor code in a fresh Python session. Display a value by leaving it on the final line.

Python starts when you open a lesson.

Output

Run your code to see what Python returns.