Conversation
…s support BaseSelector.transform() returns the retained features in the train set order, in the same library as the input (pandas X[features], narwhals select otherwise). BaseRecursiveSelector.fit() trains the estimators on native frames and returns (nw_X, y). The helpers in base_selection_functions no longer import pandas: correlations are computed with numpy (np.corrcoef, or matrix products for pairwise complete observations when there are missing values), and feature importances are pandas Series for pandas input and dicts otherwise. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
fit() finds constant and quasi-constant features in pandas, polars and other narwhals dataframes. pandas uses nunique() and value_counts(sort=False).max() per variable; polars counts the values of all variables in one select with unique_counts(); other backends use narwhals. NaN and null are both treated as missing values. With missing_values="include", missing values are counted as their own value instead of being replaced by the string "missing_values". This fixes category and nullable integer variables, which raised a TypeError, and stops merging missing values with a "missing_values" category. With missing_values="ignore" and tol<1, variables with only missing values no longer raise an IndexError: they are kept. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #1070 (selection base classes): its commit shows in the diff until #1070 is merged.
Summary
DropConstantFeaturesnow works with pandas, polars and other narwhals dataframes, and returns the same library it receives (transformis inherited fromBaseSelector).fitfollows thenw_X = check_X(X)pattern and passes the nativeXto_select_all_variables,_check_contains_naand_get_feature_names_in.X[var].nunique(dropna=...) == 1whentol=1, andX[var].value_counts(dropna=..., sort=False).max() / n_rows >= tolotherwise.tol=1uses narwhalsn_unique()for all variables in oneselect.tol<1uses polars'unique_counts().max()for all variables in oneselect, throughnw.get_native_namespace.n_unique()andSeries.value_counts()for each variable._check_contains_nadoes. So polars NaN values count as missing too. A small helper,_missing_values, appliesfill_nan(None)to float variables anddrop_nulls()when missing values are ignored. It works on narwhals expressions, polars expressions and series.missing_values="include", missing values are counted as their own value (dropna=False). The dataframe is no longer copied and filled with the string"missing_values".Got {param} instead., andmissing_valuesis type-checked before the membership test.tolis described as "equal to or greater than", which is what the code does. Theignoresemantics are spelled out, and there is a polars example. User guide: fixed the default oftol(it said zero, it is 1), updated thevalue_countsoutputs (Name: proportion), and added "Missing values" and "With polars" sections. All examples were run.Benchmarks
Median of repeats, alternating candidates. Data: a quarter each of quasi-constant ints, floats with 10% NaN, strings with 5% None, and ints with 1000 levels. The numbers are the per-variable counting step for all variables (ms).
dropna=Trueis ignore/raise anddropna=Falseis include.pandas, most-frequent count (
tol<1)len().over()value_counts(sort=False).max()wins in almost every case at 10k–100k rows, often by 2–20x. At 500k rows the narwhals group_by loop wins 3 of 4 runs by 15–30%, and loses by 2–10x at 10k–100k with many columns. numpy (np.sort(axis=0)on numeric columns, then the longest run) only works for numeric variables. It lost at 10k–100k (e.g. 572 vs 1101 ms at 100k×200) and won at 500k by about 1.2–1.5x, so I did not use it.pandas, n_unique (
tol=1): theX[var].nunique()loop was fastest or tied in most cases (e.g. 100k×50: 65 ms, vs 93 forX[vars].nunique()and 197 for narwhals).polars, most-frequent count (
tol<1), NaN handled as missing in every candidateunique_counts().max()value_counts()len()/count().over()Series.value_countslooppolars-native counting is 1.1–3x faster than the best narwhals option.
unique_countsandvalue_countsare tied, and I usedunique_countsbecause the expression is simpler.polars, n_unique (
tol=1): narwhalsn_unique()and polarsn_unique()are tied (e.g. 500k×50: 68 vs 74 ms; 2M×50: 308 vs 344 ms), so this stays on narwhals.Whole
fit(), pandas, old (origin/main) vs new (ms)Whole
fit(), polars (new, ms): 500k×10: 16 / 16 / 37 / 41; 500k×50: 51 / 66 / 103 / 86; 2M×10: 54 / 68 / 180 / 165; 2M×50: 316 / 277 / 550 / 644 (same column order as the pandas table).Behaviour
Before migrating, I recorded outputs on
origin/mainfor about 60 cases: everytolandmissing_valuesbranch, NaN/None/NaT, variables with only missing values, a variable list, a single variable,confirm_variables,tol=0, all features dropped, integer column names, category,Int64, pandas index, and whether the input is modified. pandas output is identical except for the bug fixes below. polars gives the samefeatures_to_drop_,variables_and output columns as pandas in every case. The generic narwhals path, forced with polars input, gives the same outputs as the polars path.Bugs fixed (each has a test that fails on the old code):
missing_values="ignore"withtol<1raisedIndexError: index 0 is out of boundswhen a variable had only missing values. Such variables are now kept, like they already were withtol=1:nunique()is 0.missing_values="include"raisedTypeErrorfor pandas category variables ("Cannot setitem on a Categorical with a new category") and nullable integer variables (Int64), because of thefillna("missing_values").missing_values="include"merged missing values with an existing"missing_values"value. For example,["missing_values", "missing_values", None, None, "x"]was dropped withtol=0.8: it counted 4/5 of the same value. Missing values are now a value of their own (2/5).Tests
tests/test_selection/test_drop_constant_features.pyis rewritten to the conventions: init error tests with the full messages,test_init_param_assignment, then fit/transform tests over both backends withmake_dfandframe_to_dict. The file also covers NaN handled as missing on both backends and the bug fixes. pandas-only tests cover integer column names, the pandas index, category dtype andInt64.tests/test_selection, baseorigin/narwhals-selection-basevs this branch: 137 failed / 406 passed → 128 failed / 489 passed. There are no new failures. The 9 fixed failures are the 5 old DropConstantFeatures tests and 4test_check_estimator_selectors.pychecks forDropConstantFeatures.tests/parametrize_with_checks_selection_v16.py: 289 failed before and after. It does not collect a DropConstantFeatures check.flake8 feature_engine testsis clean.mypy feature_engineshows the same 2 pre-existing errors as the base (datetime_subtraction.py, log.py).Needs decision
unique_counts()is 1.1–3x faster than the fastest narwhals alternative, so polars gets its own branch fortol<1(nwd.is_polars_dataframe+nw.get_native_namespace). Other backends fall back to a narwhalsSeries.value_counts()loop. The fallback is not exercised by the test suite, because pyarrow is not installed; I checked it by forcing the path with polars input. If you prefer two paths, the else branch could use narwhalslen()/count().over(var).max()for polars too. That costs 1.1–3x on that step.missing_values="ignore"is not consistent betweentol=1andtol<1. I kept the old behaviour. Withtol=1, a variable with one value besides missing values is dropped (nunique(dropna=True) == 1). Withtol<1, the share of the most frequent value is still divided by all rows, missing ones included. For example,[1, 1, 1, None, None]is dropped withtol=1but not withtol=0.8(3/5). This is now documented in the docstring and the user guide._check_variable_number()is not called.DropConstantFeaturesaccepted a single variable before, and still does. This is kept to avoid a behaviour change.Pre-existing issues, not fixed