Understand the idea
Population shift, group structure and changed feature availability can break apparent validation success. State the intended use and the new evidence needed rather than promising transportability.
Python skill: Use sets to check an incoming schema before calling predict().
Meet the syntax
set(incoming.columns)
required - set(incoming.columns)
listset(incoming.columns)- Collects the available feature names into a set.
required - set(incoming.columns)- Finds required names missing from the incoming table.
list- Turns the missing-name set into a list; its order does not affect this check.
Follow the code
Use the numbered comments to connect each Python block to the workflow above.
required = {'distance', 'weight'}
incoming = pd.DataFrame({'distance': [4, 7]})
answer = list(required - set(incoming.columns))
This practice: Read and run the Python. Next: Change · What changes outside this dataset?
Given data · LINE24
24 observations. Deterministic teaching observations; values illustrate the concept rather than a real population claim. The dataframe df is supplied afresh for each Run.
| distance | duration |
|---|---|
| 1 | 5.30501 |
| 1.47826 | 11.2361 |
| 1.95652 | 11.8488 |
| 2.43478 | 14.809 |
| 2.91304 | 10.7589 |
| 3.3913 | 19.4972 |
| 3.86957 | 17.89 |
| 4.34783 | 15.531 |
Column meanings and units
Column names describe the supplied features and target. Keep the stated units and row identities when making comparisons.
Synthetic data are deliberately small and reproducible. Their patterns illustrate an idea; they are not evidence about a real population.
| Column | Stored type |
|---|---|
| distance | float64 |
| duration | float64 |
Your task · Follow
A fitted workflow needs distance and weight. The incoming table has distance only. Store the missing feature names in a list named answer.
Hint 1 — Think
Use sets to check an incoming schema before calling predict().
Hint 2 — Tools
Use set(incoming.columns), required - set(incoming.columns), list. Read the visible syntax meanings before editing.
Hint 3 — Approach
A fitted workflow needs distance and weight. The incoming table has distance only. Store the missing feature names in a list named answer. Keep the supplied row order and inspect the named output after running.
Explained solution
required = {'distance', 'weight'}
incoming = pd.DataFrame({'distance': [4, 7]})
answer = list(required - set(incoming.columns))
The incoming table is missing weight, so it cannot support this two-feature workflow as specified. Matching names is only the first check: units, availability at prediction time and the population must also match the intended use.
Helpful prior knowledge: Explain a result responsibly These links are guidance, not locks.
Sources and API context
Examples run with this Playground’s scikit-learn 1.4.2 / Pyodide 0.26.4 runtime.