Skip to content

Migrate SmartCorrelatedSelection to narwhals, add polars support - #1086

Open
solegalli wants to merge 3 commits into
narwhals-migrationfrom
narwhals-smart-correlated-selection
Open

solegalli wants to merge 3 commits into
narwhals-migrationfrom
narwhals-smart-correlated-selection

Conversation

@solegalli

Copy link
Copy Markdown
Collaborator

Stacked on #1070 (selection base classes): its commit shows in the diff until #1070 is merged.

Summary

Migrates SmartCorrelatedSelection to narwhals: it now accepts pandas and polars dataframes (and other narwhals-supported libraries) and transform() (inherited from BaseSelector) returns the same library as the input.

  • fit() no longer uses pandas. nw_X = check_X(X), or nw_X, y = check_X_y(X, y) when selection_method is "model_performance" or "corr_with_target".
  • The statistics used to order the features before the correlation search:
    • missing values, standard deviation, cardinality: pandas methods on pandas input (isna().sum(), std(), nunique()), narwhals expressions otherwise (null_count(), std(), drop_nulls().n_unique()).
    • correlation with the target: numpy/scipy per feature for pandas (and for kendall / callables and other libraries), with the same functions pandas' corrwith() uses (np.corrcoef, spearmanr, kendalltau, on the rows where the feature is not NaN). For polars with pearson/spearman, polars expressions via nw.get_native_namespace(X) (corr() on the non-NaN rows).
    • sorting: a stable np.argsort (_sort_features), which gives the same order as pandas' sort_values(kind="mergesort"), with ties kept in the input order and NaN last. I checked this on 2,000 random vectors with ties and NaN.
  • model_performance picks the best feature with _sort_features instead of pd.Series(...).sort_values().
  • __init__: missing_values and selection_method are checked with isinstance(..., str) before the membership test. The messages don't change.
  • Docstring: the first example showed the wrong output (it returns x1, x3, not x2, x3), fixed. Removed references to pandas.corr() and added a polars example.
  • User guide: the text no longer points to pandas.corr(). Added a "With polars" section, with its output pasted from a real run. I re-ran all existing examples. The values match, except that sets and the std() output are printed in set order, which varies from run to run.

Benchmarks

The machine was shared with other jobs (load average 15-45), so single numbers are noisy. Medians of 5-7 repeats, with 5% missing values in every column.

Statistic by statistic, same data, candidates timed alternately (ms):

pandas

statistic rows cols pandas narwhals numpy
missing 100k 100 14.9 24.5 4.2
missing 500k 50 33.0 29.8 24.7
missing 500k 100 41.5 60.3 50.5
std 100k 100 97.6 117.3 80.6
std 500k 10 35.2 54.9 52.9
std 500k 100 842.9 634.5 546.1
nunique 100k 100 501.9 (per column: 383) 516.2 -
nunique 500k 50 1336 (per column: 1185) 1569 -
nunique 500k 100 2292 (per column: 2516) 3575 -
corr target (pearson) 100k 100 98.9 935.7 65.2
corr target (pearson) 500k 100 441.7 3377 366.6
corr target (spearman) 500k 100 11746 - 9813

For pandas, missing, std and nunique are close between candidates and swap places from size to size. I kept pandas' own methods: they are as fast as the others at 500k rows and give exactly pandas' results. For the correlation with the target, the numpy/scipy loop won at every size (it is what corrwith() runs inside, without the per-column overhead). A vectorised numpy version was not faster. Narwhals on pandas was 8-10x slower.

polars

statistic rows cols narwhals polars native numpy
null count 500k 100 0.9 11.8 487
null count 1M 100 4.4 48.2 6851
std 500k 100 63 - 1135
std 1M 100 127 - 4049
n_unique (drop_nulls) 1M 100 614 (minus-null-flag: 773) - -
corr target (pearson) 500k 100 461 84 1533
corr target (pearson) 1M 100 908 347 2587
corr target (spearman) 1M 100 - 4959 31220

For the polars correlation with the target, native polars corr() was 3-5x faster than the same expression written in narwhals, and 1.1-2x faster over the whole fit(). Timings for the whole fit(), pearson: 500k x 100 went from 767 to 527 ms, and 1M x 100 from 2848 to 1706 ms. That's why this is the one place that uses polars through nw.get_native_namespace.

Whole fit(), pandas, main (pandas-only code) vs this PR (ms, pearson, best of 2 runs):

selection_method rows cols main this PR speed-up
missing_values 100k 100 816 49 16.6x
missing_values 500k 100 4830 360 13.4x
variance 500k 50 1287 216 6.0x
variance 500k 100 4385 454 9.6x
cardinality 500k 100 4473 745 6.0x
corr_with_target 100k 100 924 128 7.2x
corr_with_target 500k 100 4649 710 6.6x
any 10k-500k 10 - - 1.0-2.0x

Most of this gain comes from the faster correlation matrix in #1070. With 10 columns, the new code is as fast or faster at every size.

Whole fit(), polars, this PR (ms): missing_values 912 / 2252, variance 553 / 587, cardinality 398 / 893, corr_with_target 390 / 712, at 500k / 1M rows x 100 columns.

model_performance is dominated by cross-validation, and the selector's own work there is unchanged.

Behaviour

I recorded the outputs of 127 cases on main (pandas) and compared them with this branch. The cases cover every selection method × pearson/spearman/kendall, NaN, inf, constant columns, thresholds, int and nullable Int64 columns, integer column names, non-default index, list/array targets, cv generators, groups, variables/confirm_variables, reordered columns in transform, callables and all error paths. I compared features_to_drop_, correlated_feature_sets_, correlated_feature_dict_ (including key order), variables_, feature_names_in_, get_support() and the transformed data.

  • pandas: identical in all cases except the two listed under "Bugs fixed".
  • polars: same values as pandas in all 115 shared cases. The one difference is an invalid method string, see "Needs decision".
  • With polars, missing values are nulls: missing_values counts nulls, and cardinality does not count null as a value, like nunique() in pandas. A float NaN in a polars column is a value for those two methods, like in the imputers. The correlations skip it, like the base correlation matrix does. The user guide says this in one sentence.
  • corr_with_target with an invalid method string: pandas raised pandas' "method must be either ..." error. It now raises TypeError: 'str' object is not callable, on both backends. The other selection methods still raise pandas' error on pandas input, because that error comes from the base helper.
  • Order of checks: when y is required and missing, the "y is needed" error is now raised before the variable checks. It used to come after them. The message is unchanged.

Bugs fixed

  1. Mismatched indexes of X and y. With corr_with_target, pandas' corrwith() aligned y to X by index. When the indexes differed, it silently correlated the wrong rows and returned a different selection (main drops var_0 instead of var_8 in the user-guide data). With model_performance, y was not checked at all. Both now go through check_X_y, which raises "The indexes of X and y do not match.", as in the other migrated transformers. Test: test_error_if_index_of_X_and_y_differ.
  2. Ties in model_performance were not deterministic. Each correlated group (a set) was passed to single_feature_performance in set iteration order, which changes with PYTHONHASHSEED. When two features trained equally good models (e.g. duplicated columns), the retained feature changed between runs. On main, 4 of 6 seeds kept var_0_duplicated and 2 kept var_0. The group is now evaluated in alphabetical order, so the tie goes to the alphabetically first feature. That's the same rule the existing "duplicated features" test describes for corr_with_target. Test: test_model_performance_ties_keep_first_feature_alphabetically, which fails without the fix (checked with 4 hash seeds). This changes results only for exact ties.
  3. Docstring example output (see summary).

Tests

tests/test_selection/test_smart_correlation_selection.py is rewritten to the conventions. It has 93 tests:

  • Init: one error test per message, parametrized with wrong values and types (threshold, missing_values, selection_method, estimator missing, missing_values/selection_method clash, confirm_variables), then test_init_param_assignment covering every init parameter.
  • Fit and transform: every selection method, and the correlation methods (pearson, spearman and kendall each give a different selection on the same data). Also: ties, cv generator, groups, list/array targets, NaN/inf errors, callables, variables/confirm_variables, and that the input is not modified. All of these run on both backends via make_df, with data as dicts.
  • pandas-only: integer column names, pandas index, the index-mismatch error, and the pandas-specific invalid-method error.

Expected values are explicit. New ones were computed on main.

tests/test_selection, base (origin/narwhals-selection-base) vs this branch:

failed passed
before 137 406
after 125 480

There are no new failures. The 12 tests that now pass are the 8 old SmartCorrelatedSelection tests and 4 test_check_estimator_selectors.py checks for this class. tests/parametrize_with_checks_selection_v16.py fails on the same tests before and after (289 failures). flake8 feature_engine tests is clean, and mypy feature_engine still shows the 2 errors that are already on the base.

Needs decision

  • Validate method in __init__. Neither this selector nor DropCorrelatedFeatures validates method. The error comes from pandas at fit() for pandas input, and is a TypeError for polars or corr_with_target. I suggest validating it in __init__ ('pearson', 'spearman', 'kendall' or a callable, ending "Got {method} instead."), for both selectors, or in the base. I left it as is because it changes when the error is raised, and the error for most paths comes from the base helper.
  • Polars through the native namespace. The polars correlation with the target uses polars' own corr() through nw.get_native_namespace(X) (see benchmarks). If you prefer narwhals-only code, the narwhals expression version is about 3-5x slower for that step. It is still 3x faster than numpy.

Pre-existing issues, not fixed

  • test_error_if_method_not_permitted stays pandas-only, because the message comes from pandas (see "Needs decision").

solegalli and others added 3 commits September 19, 2026 11:44
…s support

BaseSelector.transform() returns the retained features in the train set
order, in the same library as the input (pandas X[features], narwhals
select otherwise). BaseRecursiveSelector.fit() trains the estimators on
native frames and returns (nw_X, y). The helpers in
base_selection_functions no longer import pandas: correlations are
computed with numpy (np.corrcoef, or matrix products for pairwise
complete observations when there are missing values), and feature
importances are pandas Series for pandas input and dicts otherwise.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
fit() no longer imports pandas. The statistics used to order the features
(missing values, standard deviation, cardinality) use pandas methods for
pandas input and narwhals expressions otherwise. The correlation with the
target is computed with numpy/scipy per feature (pandas and other
backends) or with polars expressions for polars input. Features are sorted
with a stable numpy argsort, which reproduces pandas' ordering.

X and y are checked with check_X_y when the selection method needs the
target, so mismatched pandas indexes now raise an error instead of
silently aligning. In model_performance, each correlated group is
evaluated in alphabetical order so that ties are resolved deterministically.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant