Skip to learning content

DATA SCIENCE PYTHON PLAYGROUND

Machine Learning · Learn / Refresh

← Choose and Explain Models lessonsQUESTIONS · MODELS · EVIDENCE
Communicate evidence · ML-M04 · 12–18 MIN

What changes outside this dataset?

Identify evidence needed for a changed use context.

Exercises within this concept

  1. FollowRead and run the PythonCurrent exercise
  2. ChangeExplain the Python result
  3. TransferExplain the Python result

Understand the idea

Population shift, group structure and changed feature availability can break apparent validation success. State the intended use and the new evidence needed rather than promising transportability.

Identify evidence needed for a changed use context.Original settingProposed new settingSame variable definitions and availability?Comparable population, horizon and costs?Seek validation evidence for the intended use.
Schematic · Identify evidence needed for a changed use context.Scroll the diagram horizontally if needed.

Python skill: Use sets to check an incoming schema before calling predict().

Meet the syntax

set(incoming.columns)
required - set(incoming.columns)
list
set(incoming.columns)
Collects the available feature names into a set.
required - set(incoming.columns)
Finds required names missing from the incoming table.
list
Turns the missing-name set into a list; its order does not affect this check.

Follow the code

Use the numbered comments to connect each Python block to the workflow above.

required = {'distance', 'weight'}
incoming = pd.DataFrame({'distance': [4, 7]})
answer = list(required - set(incoming.columns))

This practice: Read and run the Python. Next: Change · What changes outside this dataset?

Given data · LINE24

24 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.

LINE24 · first 8 prepared rows
distanceduration
15.30501
1.4782611.2361
1.9565211.8488
2.4347814.809
2.9130410.7589
3.391319.4972
3.8695717.89
4.3478315.531

Column meanings and units

Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.

Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.

Input schema
ColumnStored type
distancefloat64
durationfloat64

Your task · Follow

A fitted workflow needs distance and weight. The incoming table has distance only. Store the missing feature names in a list named answer.

Hint 1 — Think

Use sets to check an incoming schema before calling predict().

Hint 2 — Tools

Use set(incoming.columns), required - set(incoming.columns), list. Read the visible syntax meanings before editing.

Hint 3 — Approach

A fitted workflow needs distance and weight. The incoming table has distance only. Store the missing feature names in a list named answer. Keep the supplied row order and inspect the named output after running.

Explained solution
required = {'distance', 'weight'}
incoming = pd.DataFrame({'distance': [4, 7]})
answer = list(required - set(incoming.columns))

The incoming table is missing weight, so it cannot support this two-feature workflow as specified. Matching names is only the first check: units, availability at prediction time and the population must also match the intended use.

Helpful prior knowledge: Explain a result responsibly These links are guidance, not locks.

Sources and API context

Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.

Your task · Follow

A fitted workflow needs distance and weight. The incoming table has distance only. Store the missing feature names in a list named answer.

answer

Ctrl/⌘+Enter: Run · Tab: indent · Esc then Tab: leave editor

Python loads when you run. Code and results stay in this activity only.

Run your code to inspect its output. Check uses that same run.