Skip to content

Migrate DropConstantFeatures to narwhals, add polars support - #1080

Open
solegalli wants to merge 3 commits into
narwhals-migrationfrom
narwhals-drop-constant-features
Open

solegalli wants to merge 3 commits into
narwhals-migrationfrom
narwhals-drop-constant-features

Conversation

@solegalli

Copy link
Copy Markdown
Collaborator

Stacked on #1070 (selection base classes): its commit shows in the diff until #1070 is merged.

Summary

DropConstantFeatures now works with pandas, polars and other narwhals dataframes, and returns the same library it receives (transform is inherited from BaseSelector).

  • fit follows the nw_X = check_X(X) pattern and passes the native X to _select_all_variables, _check_contains_na and _get_feature_names_in.
  • pandas: X[var].nunique(dropna=...) == 1 when tol=1, and X[var].value_counts(dropna=..., sort=False).max() / n_rows >= tol otherwise.
  • polars: tol=1 uses narwhals n_unique() for all variables in one select. tol<1 uses polars' unique_counts().max() for all variables in one select, through nw.get_native_namespace.
  • Other narwhals backends: narwhals n_unique() and Series.value_counts() for each variable.
  • NaN and null are both treated as missing values, like _check_contains_na does. So polars NaN values count as missing too. A small helper, _missing_values, applies fill_nan(None) to float variables and drop_nulls() when missing values are ignored. It works on narwhals expressions, polars expressions and series.
  • With missing_values="include", missing values are counted as their own value (dropna=False). The dataframe is no longer copied and filled with the string "missing_values".
  • Init error messages now end with Got {param} instead., and missing_values is type-checked before the membership test.
  • Docstring: tol is described as "equal to or greater than", which is what the code does. The ignore semantics are spelled out, and there is a polars example. User guide: fixed the default of tol (it said zero, it is 1), updated the value_counts outputs (Name: proportion), and added "Missing values" and "With polars" sections. All examples were run.

Benchmarks

Median of repeats, alternating candidates. Data: a quarter each of quasi-constant ints, floats with 10% NaN, strings with 5% None, and ints with 1000 levels. The numbers are the per-variable counting step for all variables (ms). dropna=True is ignore/raise and dropna=False is include.

pandas, most-frequent count (tol<1)

rows cols dropna value_counts(sort=False).max() value_counts().iloc[0] original (fillna + sorted vc) narwhals len().over() narwhals group_by loop
10k 50 T 11.2 16.5 18.4 101.7 37.9
10k 50 F 19.0 26.4 75.4 61.6 89.4
10k 200 T 132 149 165 2358 322
10k 200 F 56 110 338 445 444
100k 10 F 20.8 49.6 200 56.6 41.6
100k 50 T 286 268 385 564 329
100k 200 T 855 960 1178 3806 1437
100k 200 F 814 905 4913 1628 1178
500k 10 T 467 485 611 382 341
500k 50 T 2566 2134 2456 2247 1924
500k 50 F 1893 2987 12632 2000 2147

value_counts(sort=False).max() wins in almost every case at 10k–100k rows, often by 2–20x. At 500k rows the narwhals group_by loop wins 3 of 4 runs by 15–30%, and loses by 2–10x at 10k–100k with many columns. numpy (np.sort(axis=0) on numeric columns, then the longest run) only works for numeric variables. It lost at 10k–100k (e.g. 572 vs 1101 ms at 100k×200) and won at 500k by about 1.2–1.5x, so I did not use it.

pandas, n_unique (tol=1): the X[var].nunique() loop was fastest or tied in most cases (e.g. 100k×50: 65 ms, vs 93 for X[vars].nunique() and 197 for narwhals).

polars, most-frequent count (tol<1), NaN handled as missing in every candidate

rows cols dropna pl unique_counts().max() pl value_counts() narwhals len()/count().over() narwhals Series.value_counts loop
500k 10 T 40.1 36.1 119.2 97.6
500k 10 F 32.3 39.6 84.7 103.9
500k 50 T 106.9 78.0 208.8 542.0
500k 50 F 123.4 103.5 232.7 635.8
500k 200 T 631.6 718.4 1094.5 3147.3
500k 200 F 505.5 528.5 1107.8 3671.6
2M 10 T 190.0 170.7 192.9 400.2
2M 10 F 197.7 233.2 210.2 648.0
2M 50 T 1087.6 1071.1 1223.2 2835.3
2M 50 F 589.8 675.3 822.9 3021.3

polars-native counting is 1.1–3x faster than the best narwhals option. unique_counts and value_counts are tied, and I used unique_counts because the expression is simpler.

polars, n_unique (tol=1): narwhals n_unique() and polars n_unique() are tied (e.g. 500k×50: 68 vs 74 ms; 2M×50: 308 vs 344 ms), so this stays on narwhals.

Whole fit(), pandas, old (origin/main) vs new (ms)

rows × cols tol=1 ignore tol=1 include tol=0.9 ignore tol=0.9 include
10k × 10 1.7 → 0.8 4.9 → 0.7 3.3 → 4.3 11.0 → 3.5
10k × 100 15.7 → 17.5 54.3 → 17.2 117.4 → 29.5 148.1 → 20.5
100k × 10 13.9 → 10.3 97.8 → 9.5 35.8 → 23.4 75.2 → 15.2
100k × 100 156 → 103 890 → 90 437 → 250 882 → 149
500k × 10 64 → 47 295 → 49 206 → 168 1468 → 140
500k × 50 696 → 261 1440 → 288 1283 → 804 4443 → 620

Whole fit(), polars (new, ms): 500k×10: 16 / 16 / 37 / 41; 500k×50: 51 / 66 / 103 / 86; 2M×10: 54 / 68 / 180 / 165; 2M×50: 316 / 277 / 550 / 644 (same column order as the pandas table).

Behaviour

Before migrating, I recorded outputs on origin/main for about 60 cases: every tol and missing_values branch, NaN/None/NaT, variables with only missing values, a variable list, a single variable, confirm_variables, tol=0, all features dropped, integer column names, category, Int64, pandas index, and whether the input is modified. pandas output is identical except for the bug fixes below. polars gives the same features_to_drop_, variables_ and output columns as pandas in every case. The generic narwhals path, forced with polars input, gives the same outputs as the polars path.

Bugs fixed (each has a test that fails on the old code):

  1. missing_values="ignore" with tol<1 raised IndexError: index 0 is out of bounds when a variable had only missing values. Such variables are now kept, like they already were with tol=1: nunique() is 0.
  2. missing_values="include" raised TypeError for pandas category variables ("Cannot setitem on a Categorical with a new category") and nullable integer variables (Int64), because of the fillna("missing_values").
  3. missing_values="include" merged missing values with an existing "missing_values" value. For example, ["missing_values", "missing_values", None, None, "x"] was dropped with tol=0.8: it counted 4/5 of the same value. Missing values are now a value of their own (2/5).

Tests

tests/test_selection/test_drop_constant_features.py is rewritten to the conventions: init error tests with the full messages, test_init_param_assignment, then fit/transform tests over both backends with make_df and frame_to_dict. The file also covers NaN handled as missing on both backends and the bug fixes. pandas-only tests cover integer column names, the pandas index, category dtype and Int64.

tests/test_selection, base origin/narwhals-selection-base vs this branch: 137 failed / 406 passed → 128 failed / 489 passed. There are no new failures. The 9 fixed failures are the 5 old DropConstantFeatures tests and 4 test_check_estimator_selectors.py checks for DropConstantFeatures. tests/parametrize_with_checks_selection_v16.py: 289 failed before and after. It does not collect a DropConstantFeatures check. flake8 feature_engine tests is clean. mypy feature_engine shows the same 2 pre-existing errors as the base (datetime_subtraction.py, log.py).

Needs decision

  1. Three code paths. The polars-native unique_counts() is 1.1–3x faster than the fastest narwhals alternative, so polars gets its own branch for tol<1 (nwd.is_polars_dataframe + nw.get_native_namespace). Other backends fall back to a narwhals Series.value_counts() loop. The fallback is not exercised by the test suite, because pyarrow is not installed; I checked it by forcing the path with polars input. If you prefer two paths, the else branch could use narwhals len()/count().over(var).max() for polars too. That costs 1.1–3x on that step.
  2. missing_values="ignore" is not consistent between tol=1 and tol<1. I kept the old behaviour. With tol=1, a variable with one value besides missing values is dropped (nunique(dropna=True) == 1). With tol<1, the share of the most frequent value is still divided by all rows, missing ones included. For example, [1, 1, 1, None, None] is dropped with tol=1 but not with tol=0.8 (3/5). This is now documented in the docstring and the user guide.
  3. _check_variable_number() is not called. DropConstantFeatures accepted a single variable before, and still does. This is kept to avoid a behaviour change.

Pre-existing issues, not fixed

  • None found beyond the bugs fixed above.

solegalli and others added 3 commits September 19, 2026 11:44
…s support

BaseSelector.transform() returns the retained features in the train set
order, in the same library as the input (pandas X[features], narwhals
select otherwise). BaseRecursiveSelector.fit() trains the estimators on
native frames and returns (nw_X, y). The helpers in
base_selection_functions no longer import pandas: correlations are
computed with numpy (np.corrcoef, or matrix products for pairwise
complete observations when there are missing values), and feature
importances are pandas Series for pandas input and dicts otherwise.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
fit() finds constant and quasi-constant features in pandas, polars and
other narwhals dataframes. pandas uses nunique() and
value_counts(sort=False).max() per variable; polars counts the values of
all variables in one select with unique_counts(); other backends use
narwhals. NaN and null are both treated as missing values.

With missing_values="include", missing values are counted as their own
value instead of being replaced by the string "missing_values". This
fixes category and nullable integer variables, which raised a
TypeError, and stops merging missing values with a "missing_values"
category. With missing_values="ignore" and tol<1, variables with only
missing values no longer raise an IndexError: they are kept.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant