Create a separate clean copy of df with parsed dates.
The editable setup on the right creates df. Run executes the setup and your work from top to bottom.
Your requirements
- Keep only records dated on or after 2026-08-04 and sort them by date ascending.
- Display clean; preserve df.
Your inputs
| order | item | package | price | discount | date |
|---|---|---|---|---|---|
| 301 | " hay-bale " | BULK | 8.50 | 1 | 2026-08-01 |
| 302 | FOOD | small | 12 | 2 | 2026-08-02 |
| 303 | " toy-ball " | SINGLE | bad | None | invalid |
| 304 | Bowl | Small | None | 0.5 | 2026-08-04 |
| 304 | Bowl | Small | None | 0.5 | 2026-08-04 |
| 305 | " food " | BULK | 6 | 1.5 | 2026-08-06 |
Remember the idea
Parse date strings before comparing calendar dates. An explicit format tells pandas which part is the year, month and day.
A small example
"2026-08-04" with "%Y-%m-%d" means 4 August 2026. Invalid text becomes NaT with errors="coerce".
| Code or choice | Meaning |
|---|---|
%Y / %m / %d | Four-digit year / month number / day number; literal hyphens match the input separators. |
errors="coerce" | Make invalid dates missing (NaT); the default raises a parsing error. |
parsed >= "2026-08-04" | Keep that date and later dates. Missing dates do not pass this comparison. |
Hint
Parse before chronological comparison; invalid dates will not pass.
Reveal solution
One way to do it. Keep any supplied setup in the editor and use this in the Your work section.
clean = df.copy()
clean["date"] = pd.to_datetime(clean["date"], format="%Y-%m-%d", errors="coerce")
clean = clean[clean["date"] >= "2026-08-04"].sort_values("date")
clean