Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Data Foundations · A little practice goes a long way.

← Wrangle / Preprocess lessonsTINY TABLES · REAL PYTHON · YOUR PACE
New columns & clean text · W11 · 8 MIN

Search / extract text

Select text matches and split structured names.

Exercises within this concept

  1. FollowFollow the techniqueCurrent exercise
  2. ChangeAdapt a requirement
  3. TransferChoose and combine

Understand the idea

str.contains returns a True/False mask telling you whether each string contains the requested text.

Concept sketch: contains “a”: select matching textnameMintApplePearnameApplePearcontains “a”: select matching text
Illustration · not the exercise output

A small example

Searching for "a" with case=False matches "Ada" and "TEA"; a missing value gets False with na=False.

Follow the code

Apply the idea to the supplied table. Read from top to bottom; the final line displays the result.

df[df["candy"].str.contains("a", case=False, na=False, regex=False)]

What each part does

df["candy"].str.contains("a", ...)
Check each candy name for the text a; return one True/False flag per row.
case=False
Match A and a alike.
na=False
Treat a missing name as no match so the mask can filter rows.
regex=False
Treat the search text literally rather than as a pattern.
df[...]
Keep the rows whose flags are True.
Other choices for later exercises
~mask
Reverse the flags to select non-matches instead.

Your inputs

The editable setup on the right creates df. Run executes the setup and your work from top to bottom.

Candy shop · 6 synthetic rows
candyflavourpriceratingshelf
Gummy Bearfruity1.24.1A
Choco Popchocolate2.14.6B
Mint Bitemint1.53.8A
Berry Loopfruity2.84.4B
Cocoa Cubechocolate3.44.9A
Lemon Dropfruity1.84B

Your task · Follow

  1. Using df, keep rows whose candy contains the letter "a", ignoring case.
  2. Return the filtered DataFrame.
  3. Use: str.contains().
Hint

regex=False makes the pattern literal; na=False keeps missing text out of the matches.

Reveal solution

One way to do it. Keep any supplied setup in the editor and use this in the Your work section.

df[df["candy"].str.contains("a", case=False, na=False, regex=False)]
Optional stretch

Create a Series with ["A-12", "B-35", "C-8", "A-20"]. Try str.split("-").str[0] and str.extract(r"(\d+)").

Your task · Follow
  1. Using df, keep rows whose candy contains the letter "a", ignoring case.
  2. Return the filtered DataFrame.
  3. Use: str.contains().

Tab: indent · Shift+Tab: outdent · Esc, then Tab: leave editor

Edit Python. Control or Command plus Enter runs it. Tab indents by four spaces. Shift plus Tab outdents. Press Escape, then Tab or Shift plus Tab to leave the editor.

Each run executes all editor code in a fresh Python session. Display a value by leaving it on the final line.

Python starts when you open a lesson.

Output

Run your code to see what Python returns.